Method and device for updating three-dimensional Gaussian model
By optimizing camera extrinsic parameters during the 3D Gaussian model update process, the problem of poor reconstruction quality caused by inaccurate initial camera pose was solved, and the accuracy of the reconstructed image was improved.
Patent Information
- Application Number
- CN202511629122.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-02-24
AI Technical Summary
In existing technologies, the reconstruction quality of 3D Gaussian models is poor due to inaccurate initial camera pose sets during the reconstruction process. This makes it impossible to effectively optimize camera extrinsic parameters and affects the accuracy of the reconstructed image.
By acquiring the reconstructed image and the original image of the 3D Gaussian model from the target viewpoint, the camera extrinsic parameters of the 3D Gaussian model are updated based on the reconstruction loss, and the camera pose is optimized to improve the reconstruction quality.
By optimizing camera extrinsic parameters through feedback, the initial camera pose matching deviation is reduced, thereby improving the reconstruction quality and image accuracy of the 3D Gaussian model.
Smart Images

Figure CN121564196A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of three-dimensional Gaussian model technology, and particularly to a method and apparatus for updating three-dimensional Gaussian models. Background Technology
[0002] The production and application of 3D Gaussian models have long been considered a major technological innovation following the widespread adoption of images and videos. With the continuous evolution of bandwidth and terminal computing power, the dimensions of information expression and interaction are constantly expanding. The development of 3D reconstruction technology has enabled free-viewpoint and interactive content to break through the limitations of traditional videos that can only be viewed from a fixed angle, allowing users to freely switch their focus from any angle. This not only provides new presentation formats for complex scene displays, product detail demonstrations, and live entertainment content, but also brings explosive innovation opportunities to commercial scenarios such as customized advertising embedding, multi-view live streaming, and immersive social interaction.
[0003] In the process of reconstructing a 3D Gaussian model, the 3D Gaussian model is often updated based on an initially determined set of camera poses, which can easily lead to poor reconstruction quality of the 3D Gaussian model. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a method for updating a three-dimensional Gaussian model. One or more embodiments of this specification also relate to an apparatus for updating a three-dimensional Gaussian model, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a method for updating a three-dimensional Gaussian model is provided, comprising: acquiring a three-dimensional Gaussian model constructed based on a three-dimensional scene, a reconstructed image of the three-dimensional Gaussian model from a target viewpoint, and an original image, wherein the three-dimensional Gaussian model includes multiple three-dimensional Gaussian points, the reconstructed image is determined based on the multiple three-dimensional Gaussian points and updated camera extrinsic parameters from the target viewpoint, and the updated camera extrinsic parameters are determined based on the reconstruction loss of the target viewpoint in the previous update cycle; determining the reconstruction loss of the target viewpoint in the current update cycle based on the original image and the reconstructed image; and updating the three-dimensional Gaussian model by updating the attribute information of the multiple three-dimensional Gaussian points based on the reconstruction loss.
[0006] According to a second aspect of the embodiments of this specification, a device for updating a three-dimensional Gaussian model is provided, comprising: an acquisition module configured to acquire a three-dimensional Gaussian model constructed based on a three-dimensional scene, a reconstructed image of the three-dimensional Gaussian model from a target viewpoint, and an original image, wherein the three-dimensional Gaussian model includes multiple three-dimensional Gaussian points, the reconstructed image is determined based on the multiple three-dimensional Gaussian points and updated camera extrinsic parameters from the target viewpoint, and the updated camera extrinsic parameters are determined based on the reconstruction loss of the target viewpoint in the previous update cycle; a determination module configured to determine the reconstruction loss of the target viewpoint in the current update cycle based on the original image and the reconstructed image; and an update module configured to update the three-dimensional Gaussian model by updating the attribute information of the multiple three-dimensional Gaussian points based on the reconstruction loss.
[0007] According to a third aspect of the embodiments of this specification, a computing device is provided, including: a memory and a processor; the memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, wherein the computer programs / instructions, when executed by the processor, implement the steps of the above-described method for updating the three-dimensional Gaussian model.
[0008] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method for updating a three-dimensional Gaussian model.
[0009] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method for updating a three-dimensional Gaussian model.
[0010] This specification provides a method for updating a 3D Gaussian model through one or more embodiments. The method acquires a 3D Gaussian model constructed based on a 3D scene, a reconstructed image of the 3D Gaussian model from a target viewpoint, and the original image. The 3D Gaussian model includes multiple 3D Gaussian points. The reconstructed image is determined based on the multiple 3D Gaussian points and updated camera extrinsic parameters from the target viewpoint. The updated camera extrinsic parameters are determined based on the reconstruction loss from the target viewpoint in the previous update cycle. The reconstruction loss from the target viewpoint in the current update cycle is determined based on the original image and the reconstructed image. Based on the reconstruction loss, the 3D Gaussian model is updated by updating the attribute information of the multiple 3D Gaussian points. Thus, determining the updated camera extrinsic parameters based on the reconstruction loss from the previous update cycle achieves adjustment of the camera extrinsic parameters. Determining the reconstructed image based on the updated camera extrinsic parameters and multiple 3D Gaussian points improves the accuracy of the reconstructed image. By introducing feedback optimization of the camera extrinsic parameters during the 3D Gaussian model update process, information mutual exclusion caused by initial camera pose matching deviations can be reduced, improving the reconstruction quality of the 3D Gaussian model. Attached Figure Description
[0011] Figure 1This is a 3DGS reconstruction flowchart;
[0012] Figure 2 This is an architecture diagram of a three-dimensional Gaussian model update system provided in one embodiment of this specification;
[0013] Figure 3 This is a flowchart of a three-dimensional Gaussian model update method provided in one embodiment of this specification;
[0014] Figure 4 This is a flowchart illustrating a two-terminal interaction provided in one embodiment of this specification;
[0015] Figure 5 This is a schematic diagram of the structure of a three-dimensional Gaussian model updating device provided in one embodiment of this specification;
[0016] Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0017] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0018] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to any or all possible combinations comprising one or more of the associated listed items. The term “at least one” as used in one or more embodiments of this specification means “one or more,” and “a plurality of” means “two or more.” The term “comprising” is an open-ended description and should be understood as “including but not limiting,” and may include other content in addition to what has been described.
[0019] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determining...".
[0020] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0021] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0022] 3D Gaussian Splatting (3DGS) is an emerging technology for high-quality 3D scene reconstruction and rendering. It is an explicit 3D scene representation and real-time rendering technology that achieves high-quality new perspective synthesis through millions of optimizable 3D Gaussian primitives and their radiation properties.
[0023] Neural Radiance Fields (NeRF) is a deep learning technique used for 3D reconstruction and new perspective synthesis. It is an implicit 3D scene representation method based on neural networks. Through deep learning, the scene is encoded into a continuous 5D radiance field function to achieve realistic new perspective synthesis.
[0024] Traditional NeRF technology is costly to produce and has poor reconstruction accuracy, making it difficult to use in industrial settings. However, 3DGS models solve the core training challenges of 3D scenes / objects, making low-cost, high-quality 3D content pipeline production possible. 3DGS may become the core technology foundation for countless streaming media applications based on 3D interactive content in the future. The embodiments in this specification introduce the key technologies in the 3DGS content production technology architecture, which can effectively improve the effect of 3DGS in reconstructing 3D objects.
[0025] In the embodiments of this specification, the 3DGS model or 3DGS training process can be considered as a standardized 3DGS training model. All known optimization techniques based on 3DGS model training can be regarded as a variant of this process, and are basically applicable to the optimization techniques proposed in the embodiments of this specification.
[0026] The technology for object segmentation (subject segmentation, foreground segmentation) is now relatively mature. The segmentation process mentioned in the embodiments of the specification does not limit the segmentation technology actually used. You can refer to the introduction of any technology based on traditional methods or neural network methods.
[0027] Camera pose generally refers to the position and orientation of a camera in three-dimensional space when an image is captured. It typically includes camera pose (R, t): R is a 3x3 rotation matrix (or quaternion / Euler angles) describing the camera's orientation; t is a 3x1 translation vector describing the camera's spatial position, defining the accurate spatial position of that viewpoint in the world coordinate system. Only with the pose of each image can the true "viewpoint" of each frame be reconstructed within a unified 3D scene. Based on multiple viewpoint images, algorithms automatically infer and optimize the position and orientation of each camera in the world coordinate system, establishing an accurate spatial correspondence between them in 3D space. This process is called "camera pose matching."
[0028] Camera pose matching technology can generally be divided into two categories:
[0029] (1) The method based on feature points and geometric constraints in traditional computer vision (Structure-from-Motion (SfM)). A representative example of this method is COLMAP, an open-source tool for SfM and Multi-View Stereo (MVS), and one of the most widely used 3D reconstruction and camera pose restoration (registration) tools in academia and industry. Its main workflow is as follows:
[0030] a) Feature extraction: Extract local image features (such as scale-invariant feature transform (SIFT)) from all images and perform feature description.
[0031] b) Feature matching: Using methods such as nearest neighbor, we pair up images with the same features.
[0032] c) Geometric verification: Use robust algorithms such as Random Sample Consensus (RANSAC) to filter out false matches.
[0033] d) Initial two-view geometric estimation: Calculate the fundamental matrix and the essential matrix to determine the relative pose of the initial two images.
[0034] e) Incremental 3D reconstruction: Gradually add more photos and use algorithms such as the Perspective-n-Point (PnP) algorithm to recursively calculate their spatial pose.
[0035] f) Global Bundle Optimization: Optimize the rotation, translation parameters and spatial point coordinates of all cameras using Bundle Adjustment and other methods.
[0036] g) Output: Output the extrinsic parameters (pose matrix) of all observation cameras as the data basis for 3DGS and other 3D reconstruction technologies.
[0037] (2) Camera pose optimization method based on neural network. This method is an emerging method in recent years. The most applicable method is the Visual Geometry Grounded Transformer (VGGT). The core is to use a neural network model based on graph structure and gradient optimization to optimize the pose of each camera by using rendering error to make its reprojection and rendering image as aligned as possible. It is significantly better than COLMAP in most content. The biggest problem at present is that the maximum supported resolution is 512x512 due to the limitations of video memory and computing power.
[0038] Figure 1 A 3DGS reconstruction flowchart is shown, such as Figure 1 As shown, video data is acquired, and frame extraction and quality enhancement are performed to obtain an image dataset, which is then transmitted to the cloud. In the cloud, monocular depth estimation is performed on the image dataset to obtain depth information for each image. This depth information represents the depth of the scene and can be the depth value of each pixel in a 2D image. Camera pose estimation is then performed on the image dataset to obtain the camera pose for each image. The camera pose represents the shooting perspective corresponding to the image. Image texture enhancement is then performed on the image dataset to obtain 2D images, which consist of multiple images for which texture enhancement has been performed. Based on the depth information from multiple images, the camera pose, and the multiple 2D images, 3DGS reconstruction is performed to obtain a reconstructed 3DGS model. The 3DGS model includes multiple 3D Gaussian points, each following a three-dimensional Gaussian distribution and possessing parameters such as position, covariance matrix, transparency, and color. The 3DGS model is a model for reconstructing a 3D scene from video data.
[0039] In the data preparation phase, collect images from multiple perspectives, ideally at least 20 to 50 images from different angles, ensuring sufficient overlap and including camera parameters (focal length, position, etc.). The image dataset can be based on video data. Video frame extraction aims to extract keyframes from the video, ensuring data diversity and representativeness. Extraction can be done at fixed time intervals, suitable for videos with smooth motion and slow content changes; or it can be based on inter-frame differences (optical flow / pixel changes), suitable for videos with intense motion or rapid content changes. Quality enhancement aims to improve the visual quality and usability of the extracted images. Basic enhancement operations can include denoising, achieved through non-local mean denoising or deep learning denoising; or increasing resolution or performing color correction. Advanced restoration operations can include handling motion blur and addressing low-light conditions.
[0040] For camera pose estimation, 3D reconstruction tools such as COLMAP or SFM are used. 1) In the data preparation and feature extraction stage, multi-view images of the same scene are first collected, with an overlap rate of over 60% recommended. Scale- and rotation-invariant feature points are extracted from each image. Traditional methods commonly use SIFT features, while modern methods employ deep learning-based algorithms. Feature points need to contain both location information and feature descriptors. 2) In the feature matching and filtering stage, feature matching is performed between all image pairs. The K-nearest neighbor algorithm is used to find the best match for each feature point in another image. Then, a ratio test is applied to filter out unreliable matches, typically retaining pairs with a distance ratio less than 0.7. To further improve accuracy, the RANSAC algorithm can be used to estimate the fundamental matrix and remove outliers. 3) In the initial reconstruction stage, the two images with the most matching points are selected as the initial image pairs. The essential matrix is calculated using the eight-point algorithm, which then decomposes it to obtain the rotation matrix and translation vector of the two cameras. Triangulation is used to calculate the 3D spatial points corresponding to the matching feature points, establishing the initial sparse point cloud and camera pose. 4) In the incremental reconstruction phase, the remaining images are added to the reconstruction process sequentially. For each new image: the camera pose is estimated using the PnP algorithm, the newly observed feature points are triangulated into 3D points, and Bundle Adjustment (BA) is used to jointly optimize all camera parameters and 3D point positions. 5) In the global optimization phase, after all images have been added, global BA optimization is performed to minimize reprojection error. All camera poses and 3D point positions are optimized to minimize the reprojection error of each 3D point in all its visible images. This step is crucial for improving the overall reconstruction accuracy. 6) In the result output phase, the final output includes: the camera pose (rotation matrix R and translation vector t) for each image, a sparse 3D point cloud, and the association information between each 3D point and its corresponding 2D image feature points.
[0041] In monocular depth estimation, the process involves input image, feature extraction, depth prediction, post-processing optimization, and depth map output. In the feature extraction stage, pre-trained convolutional neural networks can be used to extract multi-scale features, or other models can be used to extract global context, combining shallow detail features with high-level semantic features through skip connections. In the depth prediction stage, spatial resolution is restored through progressive upsampling, and feature map magnification is performed using transposed convolution or interpolation. Depth predictions are output at each scale (multi-scale supervision), ultimately outputting a depth map with the same resolution as the input. Inverse depth representation is used to improve the accuracy of near-field details. In the post-processing optimization stage, bilateral filtering is used to maintain object boundary sharpness (edge enhancement), and unreliable prediction regions (such as transparent / reflective surfaces) are optimized using Conditional Random Field (CRF)-based methods (hole filling), or relative depth is converted to absolute depth using camera parameters or scene priors (scale restoration).
[0042] Image texture enhancement involves inputting the image, analyzing the texture, enhancing features, and fusing the output. Specifically, it can be achieved through wavelet transform enhancement, performing multi-level wavelet decomposition of the image, enhancing high-frequency subband coefficients, controlling the enhancement amplitude through a nonlinear gain function, and obtaining the enhancement result through wavelet reconstruction. Alternatively, it can be achieved through local contrast enhancement, calculating the local standard deviation map, enhancing high-variance regions through S-curve mapping, and maintaining edge sharpness through joint bilateral filtering.
[0043] During camera pose estimation, a sparse point cloud is generated, which is used to initialize the parameters of the 3DGS model. During 3DGS model training, given the viewpoint (camera pose), the 3DGS model is projected onto a 2D plane to obtain a 2D reconstructed image. The loss between the reconstructed image and the original 2D image is calculated, and the gradient is calculated based on the loss. The parameters of the 3DGS model are then optimized through backpropagation. During optimization, the density of the 3D Gaussian distribution is dynamically adjusted: 3D Gaussian distributions are added to under-reconstructed regions, and redundant 3D Gaussian distributions are removed from over-reconstructed regions, ultimately resulting in a reconstructed 3DGS model.
[0044] The process involves encoding the point cloud data of a 3DGS model using an encoder to obtain the corresponding encoded data. This encoded data is then transmitted to a mobile device via cloud transmission. On the mobile device, a decoder decodes the encoded data to obtain the point cloud data. The encoding and decoding methods can be any method capable of encoding or decoding. For example, the encoding method could be a traditional method based on geometric characteristics, using octree encoding. The encoding process involves recursively dividing the point cloud space into eight cubes, recording the occupancy state of each node, and storing the tree structure in breadth-first order. The decoding process involves reconstructing the spatial partitioning based on the hierarchical relationship and generating 3D points at the leaf node positions. Alternatively, the encoding method could be a projection-based method using depth map projection. The encoding steps are: projecting the point cloud onto six orthogonal views to generate an RGB-D image sequence, which is then compressed using a video encoder (H.265); the decoding and reconstruction process involves decoding the video stream to obtain a depth map and back-projecting it to generate 3D points.
[0045] Mobile point cloud reconstruction and rendering are performed based on the decoded point cloud data to output a 3DGS model. Six degrees of freedom (6DoF) interaction is then performed on the 3DGS model to obtain video data.
[0046] Point cloud reconstruction and rendering includes point cloud data loading and model rendering. Point cloud data loading: For common 3DGS data structures, the original data can be obtained optionally through local data reading or online resource loading, reading in the coordinates, covariance matrix, transparency, color, and spherical harmonic coefficients of each Gaussian point. b) Model rendering: For any projection viewpoint, the 3D Gaussian (ellipsoid) is projected onto a 2D image space (ellipse) for rendering. Given the view transformation W and the 3D covariance matrix O, the projection 2D covariance matrix O' = JWOW is calculated. T J T , where J is the Jacobian matrix of the affine approximation in the projection transformation. Based on the parameters of the projected 2D Gaussian distribution and the projection viewpoint, the rasterized image after projection is determined and rendered to display the 3DGS model.
[0047] 6DoF interaction allows users to freely move (forward, backward, left, right, up, down) and rotate (pitch, yaw, roll) 3DGS models in 3D space. After 6DoF interaction, the corresponding video data of the 3DGS model is obtained through projection rendering.
[0048] During the training process of 3DGS reconstruction, the training model randomly matches several initial camera poses for view reconstruction. Loss feedback is then calculated by comparing the reconstructed view with the original view. In this process, all pixels have equal weights; that is, every pixel in the initial view participates equally in the reconstruction of the Gaussian point cloud in terms of position, color, transparency, and ellipsoid equation. Given a fixed set of initial camera poses, the 3DGS model does not have methods for fine-tuning or error handling.
[0049] In the update process of the 3D Gaussian model, it is necessary to estimate the camera poses corresponding to multiple original images. Based on the estimated camera poses, the reconstructed image of the 3D Gaussian model at the target viewpoint (corresponding to the camera pose) is obtained. The reconstruction loss is determined based on the reconstructed image and the original images, and the 3D Gaussian model is updated based on the reconstruction loss. After the camera pose is determined at the initial time step, the 3D Gaussian model is continuously updated based on the initially determined set of camera poses, without adjusting the camera pose during the update process. If the initially determined camera pose is inaccurate, the 3D Gaussian model does not have the ability to adjust the camera pose, which will lead to inaccurate reconstructed images determined based on the camera pose, affecting the reconstruction quality of the 3D Gaussian model.
[0050] To improve the reconstruction quality of the 3D Gaussian model, the embodiments in this specification determine the updated camera extrinsic parameters based on the reconstruction loss of the target viewpoint from the previous update cycle during the 3D Gaussian model update process, and determine the reconstructed image based on the 3D Gaussian model and the updated camera extrinsic parameters. By introducing feedback optimization of the camera extrinsic parameters during the 3D Gaussian model update process, the information mutual exclusion caused by the initial camera extrinsic parameter matching deviation is reduced, thereby improving the reconstruction quality of the 3D Gaussian model.
[0051] This specification provides a method for updating a three-dimensional Gaussian model. It also relates to a device for updating a three-dimensional Gaussian model, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0052] Considering the large number of parameters in a 3D Gaussian model and the limited computing resources on the client side, the 3D Gaussian model update method proposed in the embodiments of this specification can be applied to, for example... Figure 2 The three-dimensional Gaussian model update system shown is not limited to this. See also Figure 2 , Figure 2 This specification illustrates an architecture diagram of a three-dimensional Gaussian model update system according to an embodiment of the present specification. The three-dimensional Gaussian model update system may include a client 100 and a server 200.
[0053] Client 100 is used to send a 3D Gaussian model based on the 3D scene to server 200;
[0054] Server 200 is used to acquire a 3D Gaussian model constructed based on a 3D scene, a reconstructed image of the 3D Gaussian model from the target viewpoint, and the original image from the target viewpoint. The 3D Gaussian model includes multiple 3D Gaussian points. The reconstructed image is determined based on the multiple 3D Gaussian points and the updated camera extrinsic parameters from the target viewpoint. The updated camera extrinsic parameters are determined based on the reconstruction loss of the target viewpoint in the previous update cycle. Based on the original image and the reconstructed image, the reconstruction loss of the target viewpoint in the current update cycle is determined. Based on the reconstruction loss, the 3D Gaussian model is updated by updating the attribute information of the multiple 3D Gaussian points. The 3D Gaussian model is then sent to client 100.
[0055] Client 100 is also used to receive the three-dimensional Gaussian model sent by server 200.
[0056] like Figure 2 As shown, the server 200 can connect to one or more clients 100 via a local area network (LAN), a wide area network (WAN), the Internet, or other types of data networks. Clients 100 may include, but are not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Clients 100 can also interact with users through a graphical user interface to implement the content display methods provided in the embodiments of this specification.
[0057] Server 200 may include servers providing various services, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0058] It is worth noting that the method for updating the 3D Gaussian model provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, if the client's runtime resources can meet the deployment and operation conditions of the 3D Gaussian model updating system, the client can also have similar functions to the server, thereby executing the method for updating the 3D Gaussian model provided in the embodiments of this specification. In other embodiments, the method for updating the 3D Gaussian model provided in the embodiments of this specification can also be executed jointly by the client and the server.
[0059] See Figure 3 , Figure 3 This specification shows a flowchart of a three-dimensional Gaussian model update method according to an embodiment, which specifically includes the following steps:
[0060] Step 302: Obtain a 3D Gaussian model constructed based on the 3D scene, a reconstructed image of the 3D Gaussian model from the target viewpoint, and the original image. The 3D Gaussian model includes multiple 3D Gaussian points. The reconstructed image is determined based on the multiple 3D Gaussian points and the updated camera extrinsic parameters from the target viewpoint. The updated camera extrinsic parameters are determined based on the reconstruction loss of the target viewpoint from the previous update cycle.
[0061] It should be noted that a 3D scene refers to a scene composed of one or more objects with a three-dimensional spatial structure. A 3D scene can be a real-world scene, such as a scene of many people sitting around eating in the real world, or it can be a scene in the virtual world, such as a scene of game characters talking in a game world.
[0062] A 3D Gaussian model refers to an interactive 3D content representation of a 3D scene; it can be a 3D reconstructed model corresponding to the 3D scene. In other words, a 3D Gaussian model can refer to any reconstructed model that can represent a 3D scene. For example, a 3D Gaussian model can be a 3DGS model.
[0063] A 3D Gaussian model is constructed based on point cloud data of a 3D scene. For example, a 3D Gaussian model is a 3DGS model, and the point cloud data includes data from multiple 3D Gaussian points (typically 1 million to 5 million 3D Gaussian points). Each 3D Gaussian point follows a 3D Gaussian distribution and has position, covariance matrix, color, and transparency. Position is the center location of the Gaussian point, which is a 3D data element; the covariance matrix is a 3*3 matrix that controls the shape and orientation of the Gaussian point; color is represented by spherical harmonic coefficients, which is also a 3D data element. Each Gaussian point stores a set of spherical harmonic coefficients to encode the color variation of that Gaussian point in different viewpoints; transparency determines the contribution of the Gaussian point to the final pixel color during rendering. Transparency represents the degree of opacity of the Gaussian point, and the transparency range is [0, 1]. The lower the transparency, the more transparent the Gaussian point; the higher the transparency, the less transparent the Gaussian point.
[0064] The target viewpoint refers to the viewpoint allowed for projection by the 3D Gaussian model. It can be understood that the target viewpoint can refer to any one or more of the viewpoints from which the 3D Gaussian model can be projected. The target viewpoint is determined based on the shooting viewpoints of multiple original images corresponding to the 3D Gaussian model. For example, if the original images include 20 images from viewpoint 1 to viewpoint 20, then the target viewpoint is any one or more of viewpoints from viewpoint 1 to viewpoint 20.
[0065] A reconstructed image is a rasterized image projected onto a two-dimensional plane from the target viewpoint. For subsequent loss calculations, the 3D Gaussian model needs to be projected into a two-dimensional image space to obtain the reconstructed image. The target viewpoint can refer to one or more viewpoints from which the 3D Gaussian model is projected.
[0066] The original images are the input data used to construct the 3D Gaussian model, and are typically real-world photographs taken from multiple perspectives (such as objects or scenes photographed from different angles). Multiple perspective coverage of the original images is required, usually ranging from dozens to hundreds of images. The original images provide the initial basis for the point cloud data and serve as supervisory signals to guide the alignment of the reconstructed image with the real image.
[0067] Updating camera extrinsics refers to obtaining updated camera extrinsics after adjusting the initial camera extrinsics. Camera extrinsics are the extrinsics of the camera pose, including the rotation matrix and translation vector. The updated camera extrinsics can be determined based on the reconstruction loss of the target viewpoint from the previous update cycle. There is a one-to-one correspondence between the camera extrinsics and the target viewpoint.
[0068] The previous update cycle refers to the update cycle immediately preceding the current update cycle. The 3D Gaussian model needs to be updated over multiple update cycles. In each update cycle, the reconstruction loss is calculated, and the 3D Gaussian model is updated.
[0069] The reconstruction loss is used to optimize the content parameters of the 3D Gaussian model, and is usually determined based on the difference between the reconstructed image and the original image from the target viewpoint.
[0070] In practical applications, the 3D Gaussian model is trained based on point cloud data of a 3D scene. In the initial update cycle, point cloud initialization is performed: for the 3D scene, a set of original static scene images are acquired, and the camera parameters (usually the camera extrinsic parameters corresponding to the viewpoint) of each original image are calibrated using SfM. This process generates a sparse point cloud, which is then initialized to obtain multiple 3D Gaussian points, forming the initial 3D Gaussian model. Each 3D Gaussian point follows a 3D Gaussian distribution, with parameters including position, covariance matrix, color, and transparency. When SFM data is unavailable, sparse point clouds can be obtained through uniform / random sampling, or through backprojection from multi-view depth maps. In other update cycles, the acquired 3D Gaussian model is the updated 3D Gaussian model from the previous update cycle.
[0071] When initializing the position of the 3D Gaussian point, the 3D coordinates of the SFM point cloud can be directly used as the center of the Gaussian. The covariance matrix is usually initialized as an isotropic spherical Gaussian. The initial scale can be set to a small value (such as 0.01 to 0.1) to avoid excessive blurring during rendering. The orientation can be initialized as an identity matrix (without rotation) or estimated from the normal of the SFM point cloud. The transparency is initialized to a medium value (such as 0.5) to allow for subsequent optimization and adjustment. The initial color is initialized from the RGB values of the three primary colors of the SFM point cloud or set to a neutral color.
[0072] When selecting the target viewpoint, it is necessary to ensure that the images from different viewpoints have sufficient overlap and that the images from multiple viewpoints can achieve complete coverage of the 3D scene.
[0073] The original images can be obtained by taking pictures of the scene from different angles with a camera. It is recommended to take at least 20 to 50 images from different angles to ensure that there is enough overlap between the images.
[0074] From the target viewpoint, the 3D Gaussian model is projected onto a 2D image plane for rendering to obtain the reconstructed image. When projecting the 3D Gaussian points onto the 2D image plane, the parameters (position, covariance matrix, color, and transparency) of the projected 2D Gaussian need to be determined. Each projected 2D Gaussian corresponds to a 3D Gaussian point before projection. The position of the projected 2D Gaussian can be determined by the position of the corresponding 3D Gaussian point and the camera extrinsic parameters corresponding to the projection viewpoint. The covariance matrix O′ of the projected 2D Gaussian is determined by formula (1):
[0075] O′=JROR T J T (1)
[0076] Where R is the camera rotation matrix (the rotation part of the extrinsic parameters), O is the covariance matrix of the 3D Gaussian points, and J is the approximate Jacobian matrix of the projection transformation. The color of the projected 2D Gaussian point can directly inherit the color of the 3D Gaussian point, or it can be calculated based on the viewpoint direction and spherical harmonic function. The transparency of the projected 2D Gaussian point can directly inherit the transparency of the 3D Gaussian point, or it can be adjusted according to depth or viewpoint (objects appear larger when closer and smaller when farther away).
[0077] Based on the position and covariance matrix of the projected 2D Gaussians, the probability density of any pixel position of the 2D Gaussians in the projection plane can be calculated. Given a pixel position, according to the camera parameters corresponding to the projection viewpoint, the distance to all overlapping 2D Gaussians can be calculated, i.e., the depth of the corresponding 3D Gaussian point. Then, based on the depth of the 3D Gaussian points, the order of the preceding and following positions of each 2D Gaussian under the projection viewpoint is formed. Based on the order and transparency of the 2D Gaussians, the cumulative transmittance of any projected 2D Gaussian corresponding to the preceding Gaussian point can be determined. The cumulative transmittance is the product of the transparency of all preceding Gaussian points, and the transparency is calculated by subtracting the transparency from 1. For any pixel position, based on the cumulative transmittance, transparency, probability density, and color of the multiple 2D Gaussians corresponding to that pixel position, the pixel at that pixel position is determined. Based on the pixels at any pixel position, a reconstructed image is constructed.
[0078] For example, when training a 3D Gaussian model, the Gaussian point set is initialized. For each Gaussian point g i initial position μ i ∈R 3 Initial covariance matrix M i ∈R 3×3 Initial color c i ∈R 3 Initial transparency α i ∈[0,1]. After training, the final three-dimensional Gaussian model is obtained. For a set of Gaussian points, its rasterized image on the two-dimensional plane is defined as Equation (2):
[0079]
[0080] Wherein, weight w i (u,v) is determined by the Gaussian point's projection onto the pixel and its transparency, and can be written as formula (3):
[0081]
[0082] Among them, T i α is the cumulative transmittance of all preceding Gaussian points. i Let P be the transparency of the i-th Gaussian point. i(u,v) represents the probability density (two-dimensional Gaussian projection) of the i-th Gaussian point at pixel (u,v). C(u,v) is the reconstructed image of the three-dimensional Gaussian model. Here, the parameters of the Gaussian points used in calculating the reconstructed image are the parameters of the projected 2D Gaussian.
[0083] In practical applications, the reconstructed image is determined based on a 3D Gaussian model and updated camera extrinsics from the target viewpoint. In one possible implementation of this specification, in each update cycle, updated camera extrinsics from the target viewpoint are obtained, and the reconstructed image from the target viewpoint is determined based on the 3D Gaussian model and these updated camera extrinsics. It is understood that in the initial update cycle (the first update cycle), the updated camera extrinsics are the initial camera extrinsics, and in subsequent update cycles, the updated camera extrinsics are the camera extrinsics updated in the previous update cycle.
[0084] In another possible implementation of this specification, obtaining the reconstructed image of the 3D Gaussian model from the target viewpoint includes: obtaining the updated camera extrinsic parameters from the target viewpoint in the current update cycle; and generating the reconstructed image of the 3D Gaussian model from the target viewpoint based on the 3D Gaussian model and the updated camera extrinsic parameters, provided that the update order of the previous update cycle is within a preset range.
[0085] It's important to note that the 3D Gaussian model needs to be updated across multiple update cycles. The current update cycle refers to the current update cycle in which the model is being trained. The previous update cycle refers to the update cycle immediately preceding the current one. In essence, during the training of the 3D Gaussian model, each update cycle can be considered an iteration; the current update cycle refers to the current iteration, and the previous update cycle refers to the iteration before that.
[0086] The preset range refers to the range within which the update order of the update cycle falls. This preset range can be set according to requirements. It's understandable that when updating camera extrinsics, it's not necessary to adjust them in every update cycle. This is because in the initial update phase, the reconstruction loss determined based on the 3D Gaussian model is relatively large. Adjusting the camera extrinsics based on this loss would result in a significant difference between the updated extrinsics and the actual camera extrinsics. Therefore, in the initial update phase, when the 3D Gaussian model changes significantly, it's unnecessary to update the target view's camera extrinsics. Similarly, in the final update phase, the 3D Gaussian model is relatively stable, and the updated camera extrinsics for the target view also tend to stabilize, so further adjustment is unnecessary. Therefore, when the update order of the update cycle falls within the preset range, the target view's camera extrinsics are adjusted. For example, if the preset range is [5000, 15000], meaning the update order is greater than or equal to 5000 and less than or equal to 15000, the camera extrinsics are adjusted to obtain the updated extrinsics.
[0087] In practical applications, updating camera extrinsic parameters can be divided into two categories: first camera extrinsic parameters and second camera extrinsic parameters. First camera extrinsic parameters refer to the camera extrinsic parameters obtained by adjusting the initial camera extrinsic parameters of the target viewpoint in the previous update cycle based on the reconstruction loss of the target viewpoint in the previous update cycle; second camera extrinsic parameters refer to the initial camera extrinsic parameters of the target viewpoint in the previous update cycle, that is, the initial camera extrinsic parameters of the target viewpoint in the previous update cycle have not been adjusted.
[0088] If the update order of the previous update cycle is outside the preset range, the initial camera extrinsic parameters of the target view in the previous update cycle are directly determined as the updated camera extrinsic parameters of the target view in the current update cycle; if the update order of the previous update cycle is within the preset range, the updated camera extrinsic parameters of the target view in the current update cycle are obtained, including: obtaining the reconstruction loss of the target view in the previous update cycle; and determining the updated camera extrinsic parameters of the target view in the current update cycle based on the initial camera extrinsic parameters and reconstruction loss determined in the previous update cycle.
[0089] It should be noted that the reconstruction loss of the target view in the previous update cycle is determined based on the reconstructed image and the original image of the target view in the previous update cycle.
[0090] The initial camera extrinsic parameters under the target view in the previous update cycle refer to the camera extrinsic parameters under the target view in the previous update cycle, that is, the camera extrinsic parameters used to determine the reconstructed image in the previous update cycle.
[0091] In practical applications, there are multiple ways to determine the updated camera extrinsics for the current update cycle based on the initial camera extrinsics and reconstruction loss determined in the previous update cycle's target viewpoint. The specific method chosen depends on the actual situation, and this specification does not impose any limitations on these methods in the embodiments. In one possible implementation, the gradient of the reconstruction loss with respect to the initial camera extrinsics is calculated based on the reconstruction loss and the initial camera extrinsics. The initial camera extrinsics are then adjusted based on this gradient to obtain the updated camera extrinsics.
[0092] For example, for the camera extrinsic parameter V used in this calculation k If M0 ≤ k ≤ M1 (M0 = 5000, M1 = 15000 can be taken, where k, M0, and M1 refer to the number of iterations), for the camera extrinsic parameter V k Loss feedback is performed, as shown in formula (4):
[0093]
[0094] in, This represents the updated camera extrinsic parameters for the j-th target viewpoint during the (k+1)-th update cycle. This represents the initial camera extrinsic parameters for the j-th target viewpoint in the k-th update cycle, where γ is the learning rate. To rebuild losses For the initial camera extrinsic parameters V j The gradient of V is obtained by taking its partial derivative. k The index k indicates the update order.
[0095] Specifically, the gradient of the reconstruction loss with respect to the initial camera extrinsic parameters is propagated through a gradient chain. First, the gradient of the reconstruction loss with respect to the reconstructed image, the gradient of the reconstructed image with respect to the 3D Gaussian parameters in the camera coordinate system, and the gradient of the 3D Gaussian parameters in the camera coordinate system with respect to the camera extrinsic parameters are determined. Then, these three gradients are multiplied to obtain the gradient of the reconstruction loss with respect to the initial camera extrinsic parameters. (Camera extrinsic parameters V) k Including the rotation matrix R and the translation vector t, calculating the gradient of the reconstruction loss with respect to the initial camera extrinsic parameters is actually determining the gradient of the reconstruction loss with respect to the rotation matrix and the translation vector.
[0096] By applying the scheme of the embodiments in this specification, the reconstruction loss of the target view in the previous update cycle is obtained; based on the initial camera extrinsic parameters and reconstruction loss determined in the previous update cycle under the target view, the updated camera extrinsic parameters under the target view in the current update cycle are determined. In this way, the initial camera extrinsic parameters are optimized based on the reconstruction loss, and the camera extrinsic parameters are adjusted based on the loss during the update process, so as to obtain accurate updated camera extrinsic parameters.
[0097] In another possible implementation of this specification, the updated camera extrinsic parameters under the target view determined in the previous update cycle are determined based on the initial camera extrinsic parameters and reconstruction loss determined in the previous update cycle. This includes: if the reconstruction loss is less than the loss threshold, determining the updated camera extrinsic parameters under the target view determined in the previous update cycle based on the initial camera extrinsic parameters and reconstruction loss determined in the previous update cycle; and if the reconstruction loss is greater than or equal to the loss threshold, discarding the initial camera extrinsic parameters.
[0098] It should be noted that the loss threshold is used to determine whether the initial camera pose needs to be removed. Understandably, the reconstruction loss of the previous update cycle is determined based on the initial camera extrinsic parameters from the target's viewpoint in the previous update cycle. If the reconstruction loss is greater than or equal to the loss threshold, it indicates that the difference between the reconstructed image and the original image is too large. This means that the initial camera extrinsic parameters corresponding to the reconstructed image differ significantly from the actual camera extrinsic parameters. Therefore, the initial camera extrinsic parameters can be removed. The loss threshold is determined according to requirements.
[0099] Different loss thresholds can be set for different update cycles. For example, the loss threshold can be gradually reduced as the update sequence increases.
[0100] In practical applications, when the reconstruction loss is less than the loss threshold, the initial camera extrinsic parameters are updated. That is, based on the initial camera extrinsic parameters and reconstruction loss determined in the previous update cycle under the target viewpoint, the updated camera extrinsic parameters under the target viewpoint in the current update cycle are determined. When the reconstruction loss is greater than or equal to the loss threshold, the initial camera extrinsic parameters are removed and no longer updated.
[0101] For example, cameras that exceed the threshold value τ (loss threshold) are eliminated, as shown in formula (5):
[0102] V′={j|L j <τ}(5)
[0103] Where V′ represents the set of updated camera extrinsic parameters after removing the initial camera extrinsic parameters in the previous update cycle. j represents the index of the camera extrinsic parameter (the index of the target viewpoint). L j The reconstruction loss is for the j-th target viewpoint.
[0104] Applying the scheme of the embodiments in this specification, when the reconstruction loss is less than the loss threshold, the initial camera extrinsic parameters are adjusted; when the reconstruction loss is greater than or equal to the loss threshold, the initial camera extrinsic parameters are discarded. In this way, by assessing the difference between the initial camera extrinsic parameters and the actual camera extrinsic parameters based on the loss, and directly discarding the initial camera extrinsic parameters when they are seriously flawed, the camera extrinsic parameters can be optimized more effectively, thus improving the optimization efficiency.
[0105] If the update order of the previous update cycle is within a preset range, a reconstructed image of the 3D Gaussian model under the target view is generated based on the 3D Gaussian model and the updated camera extrinsic parameters (first camera extrinsic parameters); if the update order of the previous update cycle is outside the preset range, a reconstructed image of the 3D Gaussian model under the target view is generated based on the 3D Gaussian model and the updated camera extrinsic parameters (second camera extrinsic parameters).
[0106] The position of the projected 2D Gaussian can be determined by the position of the corresponding 3D Gaussian point and the camera extrinsic parameters corresponding to the target viewpoint. The covariance matrix O′ of the projected 2D Gaussian is determined by the above formula (1). The color of the projected 2D Gaussian can be directly inherited from the color of the 3D Gaussian point, or it can be calculated based on the viewpoint direction and spherical harmonic function. The transparency of the projected 2D Gaussian can be directly inherited from the transparency of the 3D Gaussian point, or it can be adjusted according to the depth or viewpoint (larger for closer objects and smaller for farther objects). The pixel at any pixel position is determined by the above formulas (2) and (3) based on the position, covariance matrix, transparency, and color of the 2D Gaussian. The reconstructed image is constructed based on the pixel at any pixel position.
[0107] By applying the scheme of the embodiments in this specification, the updated camera extrinsic parameters under the target viewpoint in the current update cycle are obtained; if the update order of the previous update cycle is within a preset range, a reconstructed image of the 3D Gaussian model under the target viewpoint is generated based on the 3D Gaussian model and the updated camera extrinsic parameters. In this way, the update order being within the preset range can meet the conditions for effective updating of camera extrinsic parameters, resulting in accurate updated camera extrinsic parameters and accurate reconstructed images.
[0108] Step 304: Based on the original image and the reconstructed image, determine the reconstruction loss of the target viewpoint in the current update cycle.
[0109] It should be noted that reconstruction loss refers to the pixel loss between the original image and the reconstructed image, and is determined by the pixel differences between the original image and the reconstructed image.
[0110] In practical applications, there are various ways to determine the reconstruction loss of the target viewpoint in the current update cycle based on the original image and the reconstructed image. The specific method is selected according to the actual situation, and the embodiments in this specification do not limit this. In one possible implementation of this specification, after obtaining the original image and the reconstructed image, the loss between the two images is directly calculated.
[0111] The loss function can be the mean absolute error loss (L1 loss), which measures the absolute difference between the predicted and the true values. As shown in formula (6):
[0112]
[0113] Among them, y i It's a real pixel, y i i is the predicted pixel, and N is the number of pixels.
[0114] The loss function can also be the mean squared error loss (L2 loss), which measures the squared difference between the predicted and actual values. As shown in formula (7):
[0115]
[0116] Among them, y i It's a real pixel, y i i is the predicted pixel, and N is the number of pixels.
[0117] The loss function can also be the Structural Similarity Index Measure (SSIM) loss, which measures the differences between two images in terms of brightness, contrast, and structure.
[0118] In practical applications, the loss function can be determined according to the requirements and is not limited to any particular form. It can be one or a combination of the loss functions mentioned above, or other loss functions.
[0119] For example, for all initial perspectives Define each camera V k =(R k ,t k Each has an extrinsic parameter that is a rotation matrix R. k Translation vector t k M represents the number of initial views. The reconstruction loss for the current update cycle is calculated; if there are multiple views, the losses from all views are summed, as shown in formula (8).
[0120]
[0121] Where M is the number of target viewpoints, K refers to the k-th target viewpoint, and C k To reconstruct the image, I k The original image is used. Formula (8) determines the reconstruction loss by calculating the L2 loss of the two images.
[0122] In another possible implementation of this specification, the reconstruction loss of the target viewpoint in the current update cycle is determined based on the original image and the reconstructed image, including: replacing the background of the original image and the background of the reconstructed image with the target background to obtain the target original image and the target reconstructed image; and determining the reconstruction loss of the target viewpoint in the current update cycle based on the target original image and the target reconstructed image.
[0123] It's important to note that real-world scene backgrounds may contain complex, dynamic, or irrelevant noise (such as moving pedestrians or lighting changes). This information is detrimental to 3D reconstruction and may even interfere with optimization. By replacing the original and reconstructed images with simple backgrounds (such as solid colors or static backgrounds), the loss function only evaluates the rendering quality of foreground objects, avoiding background noise misleading gradient updates. This improves the efficiency of optimizing the subject's geometry and appearance.
[0124] Furthermore, 3D Gaussian models rely on sparse SfM point cloud initialization. Background regions (such as the sky and distant objects) may be incompletely reconstructed due to a lack of feature points, resulting in holes or artifacts in the background rendering. After replacing the background, there is no need to optimize the Gaussian distribution of the background region; only the foreground needs to be focused on. This reduces unnecessary computation and avoids background errors affecting foreground quality.
[0125] In the process of updating the 3D Gaussian model, in order to reduce the probability of misidentifying the background of the original image as the subject, the background of both the original and reconstructed images can be replaced during the update process. Different random background colors are used for background replacement in different update cycles. In this way, when calculating the reconstruction loss, the background of the original and reconstructed images is consistent and changes continuously during the update process. This can effectively reduce the probability of misidentifying the background of the original image as the subject for reconstruction, so that the 3D Gaussian model focuses on the subject for reconstruction.
[0126] The target background can be set according to requirements, or a random background color can be used. The target background can be the same or different for different update cycles.
[0127] In practical applications, the background of an image can only be replaced with the target background after the background of the image has been identified. The background of an image can be identified by segmenting the foreground and background. There are various methods for segmenting the foreground and background of an image to identify the background; the specific method chosen depends on the actual situation, and this specification does not limit this approach. In one possible implementation of this specification, a target segmentation model is called to perform foreground segmentation on the image, obtain the foreground of the image, and determine the region in the image other than the foreground as the background. The target segmentation model can be any model capable of foreground segmentation.
[0128] Target segmentation models can be any model capable of segmenting an image into foreground and background. For example, the U-Net model includes an encoder, decoder, skip connections, and an output layer. The encoder progressively extracts multi-scale features and captures semantic information from the image. The decoder progressively restores spatial resolution and accurately locates segmentation boundaries. Skip connections address the issue of information loss in deep networks, preserving low-level features (such as edges and textures). The output layer generates the final segmentation mask. Another example is the FCN model, which transforms the fully connected layers of a classification network into equivalent 1x1 convolutional layers. Upsampling restores spatial resolution through transposed convolutions, and skip connections combine shallow and deep features to improve detail accuracy.
[0129] In another possible implementation of this specification, a threshold is set based on pixel grayscale or color values to divide the image into foreground and background. This threshold can be set using a global threshold or an adaptive threshold. This method is suitable for simple images with uniform lighting and high foreground-background contrast.
[0130] By applying the scheme of the embodiments in this specification, the background of the original image and the background of the reconstructed image are replaced with the target background to obtain the target original image and the target reconstructed image. Based on the target original image and the target reconstructed image, the reconstruction loss of the target viewpoint in the current update cycle is determined. In this way, replacing the background of the original image and the reconstructed image during the update process can effectively reduce the probability of misidentifying the background part of the original image as the main part for reconstruction, so that the 3D Gaussian model focuses on the main part for reconstruction, avoids background noise misleading gradient updates, and improves the update efficiency of the 3D Gaussian model.
[0131] In another possible implementation of this specification, the background of the original image and the background of the reconstructed image are replaced with the target background to obtain the target original image and the target reconstructed image, including: obtaining the target background information of the target background; and, while keeping the foreground of the original image and the foreground of the reconstructed image unchanged, performing background replacement on the background of the original image and the background of the reconstructed image based on the target background information to obtain the target original image and the target reconstructed image.
[0132] It should be noted that the target background information refers to the pixel information corresponding to the target background. This background information can be the RGB values of the pixel's three primary colors, or it can be other color information. The target background information corresponding to different pixel positions at different viewpoints can be the same or different. For example, the target background information could be b... k , representing the target background information for the k-th update cycle.
[0133] In practical applications, the target background information can be set according to requirements. When background replacement is needed, the target background information can be obtained.
[0134] When performing background replacement on the original and reconstructed images, at the pixel level, for any pixel in the image, if the pixel corresponds to the foreground, the pixel is not replaced; if the pixel corresponds to the background, the pixel is replaced with the corresponding target background information.
[0135] For example, assuming a total of M training iterations are performed, during the k-th iteration, several test perspectives are randomly selected, and for each perspective V... k Take the original image I(·) and the reconstructed image C(·) from this viewpoint and perform background replacement, let b k The random background color for the k-th viewpoint is shown in formulas (9) and (10):
[0136]
[0137] in, Let f(k,u,v) represent the target reconstructed image, f(k,u,v) represent the segmentation function, and C represent the segmentation function. k(u,v) represents the reconstructed image, b k This represents the background information of the target, where the subscript k indicates the k-th viewpoint, and (u,v) represents the pixel position. I represents the original image of the target. k (u,v) represents the original image.
[0138] The segmentation function f(·) is defined on the image pixels. For the k-th viewpoint and pixel (u,v), it is as shown in formula (11):
[0139]
[0140] For the k-th viewpoint and pixel (u,v), f(k,u,v) = 1 when pixel (u,v) belongs to the foreground and f(k,u,v) = 0 when pixel (u,v) belongs to the background.
[0141] By applying the scheme of the embodiments in this specification, target background information of the target background is obtained. While keeping the foreground of the original image and the foreground of the reconstructed image unchanged, background replacement is performed on the background of the original image and the background of the reconstructed image based on the target background information to obtain the target reconstructed image and the target original image. Thus, by keeping the foreground of the image unchanged and replacing the background of the image based on the target background information, background replacement can be performed on images segmented by foreground and background, resulting in accurate target reconstructed images and target original images.
[0142] Step 306: Based on the reconstruction loss, update the 3D Gaussian model by updating the attribute information of multiple 3D Gaussian points.
[0143] It should be noted that attribute information is used to describe the characteristics of a 3D Gaussian point. Attribute information can include the position, covariance matrix, color, and transparency of the 3D Gaussian point.
[0144] In practical applications, there are various ways to update the 3D Gaussian model based on reconstruction loss. The specific method chosen depends on the actual situation, and the embodiments in this specification do not impose any limitations on this. In one possible implementation of this specification, the gradient of attribute information (parameters of 3D Gaussian points) is calculated based on the reconstruction loss. The gradient includes the gradient of the 3D Gaussian point parameters, and the gradient of the loss function is backpropagated to the parameters of the 3D Gaussian points using the chain rule. The gradient of attribute information includes the position gradient, covariance gradient, transparency gradient, and color (spherical harmonic coefficient) gradient. The attribute information of the 3D Gaussian points is updated based on the gradient of the attribute information.
[0145] For example, for the training model of 3DGS, backward gradient propagation is performed as shown in Equation (12):
[0146]
[0147] in, For the 3DGS model, L represents the reconstruction loss, and k refers to the k-th target viewpoint. To reconstruct the image, I k Let M be the original image and M be the number of target viewpoints. By minimizing the reconstruction loss, for The parameters are updated.
[0148] Applying the scheme of the embodiments in this specification, a 3D Gaussian model constructed based on a 3D scene, a reconstructed image of the 3D Gaussian model from the target viewpoint, and the original image are obtained. The 3D Gaussian model includes multiple 3D Gaussian points. The reconstructed image is determined based on the multiple 3D Gaussian points and the updated camera extrinsic parameters from the target viewpoint. The updated camera extrinsic parameters are determined based on the reconstruction loss of the target viewpoint in the previous update cycle. Based on the original image and the reconstructed image, the reconstruction loss of the target viewpoint in the current update cycle is determined. Based on the reconstruction loss, the 3D Gaussian model is updated by updating the attribute information of the multiple 3D Gaussian points. Thus, determining the updated camera extrinsic parameters based on the reconstruction loss of the previous update cycle achieves adjustment of the camera extrinsic parameters. Determining the reconstructed image based on the updated camera extrinsic parameters and the 3D Gaussian model can improve the accuracy of the reconstructed image. By introducing feedback optimization of the camera extrinsic parameters during the update process of the 3D Gaussian model, information mutual exclusion caused by initial camera pose matching deviations can be reduced, improving the reconstruction quality of the 3D Gaussian model.
[0149] In another possible implementation of this specification, the attribute information includes position; based on the reconstruction loss, the three-dimensional Gaussian model is updated by updating the attribute information of multiple three-dimensional Gaussian points, including: for any three-dimensional Gaussian point, determining the position gradient of any three-dimensional Gaussian point based on the reconstruction loss; if the position gradient is greater than the gradient threshold, performing density control on any three-dimensional Gaussian point; and updating the three-dimensional Gaussian model based on the density-controlled three-dimensional Gaussian points.
[0150] It should be noted that the 3D Gaussian model includes multiple point cloud data, and each point cloud data can be regarded as a 3D Gaussian point that satisfies the 3D Gaussian distribution.
[0151] The positional gradient is the partial derivative of the loss function with respect to the center position of a 3D Gaussian point. It represents the contribution of the current position to the rendering error and also indicates the direction in which the 3D Gaussian point should move to reduce the loss. In adaptive density control, the positional gradient is used to determine which regions need more Gaussian points (under-reconstruction) and which Gaussian points need to be split (over-reconstruction).
[0152] The gradient threshold is used to determine whether density control is needed for 3D Gaussian points. Density control is required for 3D Gaussian points when the position gradient is greater than the gradient threshold.
[0153] Density control is a key strategy for dynamically adjusting the 3D Gaussian distribution during the update process of a 3D Gaussian model. Its purpose is to more accurately represent the geometry and appearance of a scene by adding, deleting, or adjusting 3D Gaussian points. The core idea is to add Gaussian points in areas requiring detail and reduce them in redundant areas, thereby balancing reconstruction quality and computational efficiency.
[0154] In practical applications, starting with the sparse point cloud generated by SfM, the number and density of Gaussian points per unit volume are adaptively controlled to gradually transition from the initial sparse Gaussian set to a dense set that better represents the scene, while ensuring the correct parameters. After the optimization warm-up phase, density adjustment is performed every 100 iterations, and any essentially transparent Gaussian points (i.e., Gaussian points with transparency less than a threshold) are removed.
[0155] Adaptive density control requires filling empty regions, focusing on areas lacking geometric features ("under-reconstructed") or areas with excessive Gaussian point coverage (typically "over-reconstructed"). The observed large positional gradients in view space for these two cases may be because these regions are not yet properly optimized, and the optimization process attempts to correct this by shifting the Gaussian points.
[0156] For small Gaussian points in underreconstructed regions, cloning is needed to cover the new geometry. Specifically, this involves creating copies of the same size and moving them along the direction of the location gradient. For large Gaussian points in high-variance regions, they need to be split into two smaller Gaussian points, with the positions of the new Gaussian points determined by sampling the probability density function of the original Gaussian points.
[0157] In the first case, the total volume and number of Gaussian points are increased by cloning; in the second case, the total volume is kept constant while the number of Gaussian points is increased by splitting. Gaussian points may shrink or grow and significantly overlap with other Gaussian points; Gaussian points that are very large in world space or occupy a large area in view space are periodically removed. This strategy effectively controls the total number of Gaussian points.
[0158] Specifically, for any three-dimensional Gaussian point, based on the reconstruction loss, the partial derivative of the position parameters of any three-dimensional Gaussian point is calculated to obtain the position gradient of any three-dimensional Gaussian point.
[0159] Filter 3D Gaussian points: remove 3D Gaussian points with transparency below the threshold, and remove 3D Gaussian points that are too large or exceed the scene boundary.
[0160] If the positional gradient is greater than the gradient threshold, the scale (covariance matrix) of the 3D Gaussian point and the scale threshold are further determined. If the scale of the 3D Gaussian point is less than the scale threshold, the 3D Gaussian point is copied, and its new position is finely adjusted along the gradient direction while keeping other parameters (color, scale) unchanged. The purpose is to fill the geometrically accurate region. If the scale of the 3D Gaussian point is greater than or equal to the scale threshold, the 3D Gaussian point is split into two, its scale is reduced, and the new Gaussian point is offset along the principal direction of the covariance matrix. The purpose is to decompose the overly smoothed region.
[0161] After density control is applied to any 3D Gaussian point, the 3D Gaussian model is updated based on the density-controlled 3D Gaussian point.
[0162] For example, for Density enhancement and parameter optimization are performed, as shown in formulas (13) and (14):
[0163]
[0164] in, Let N be the set of Gaussian points, and N be the number of Gaussian points. Let N be the set of Gaussian points after the kth update cycle, and let N be the number of Gaussian points. (k) . Let D be the set of Gaussian points for the (k+1)th update cycle, and let D(·) be the function that optimizes the set of Gaussian points based on the parameter gradient and the updated camera extrinsic parameters. For the k-th update cycle, reconstruct the gradient of the loss with respect to the parameters of the Gaussian point set, where V′ is the updated set of camera extrinsic parameters. End training and output. As the final 3DGS model. N (M) This refers to the number of Gaussian points after M update cycles.
[0165] Applying the scheme of the embodiments in this specification, for any 3D Gaussian point, the position gradient of that point is determined based on the reconstruction loss; if the position gradient is greater than a gradient threshold, density control is applied to that point; and the 3D Gaussian model is updated based on the density-controlled Gaussian points. In this way, by controlling the density of the 3D Gaussian model, 3D Gaussian points can be added in areas requiring detail and reduced in redundant areas, thereby balancing reconstruction quality and computational efficiency.
[0166] Figure 4 A flowchart illustrating a two-terminal interactive process is shown in one embodiment of this specification. Figure 4 As shown, the two ends include mobile and cloud. The specific interaction process is as follows:
[0167] Step 402: Obtain video data.
[0168] The mobile device acquires video data and sends it to the cloud.
[0169] Step 404: Perform video frame extraction and quality enhancement on the video data to obtain an image dataset.
[0170] In this process, after receiving the video data, the cloud performs frame extraction and quality enhancement to obtain an image dataset.
[0171] Step 406: Based on the image dataset, perform monocular depth estimation, camera pose estimation, and image texture enhancement to obtain depth information, initial camera pose, and original image.
[0172] In this process, the cloud performs monocular depth estimation, camera pose estimation, and image texture enhancement on the image dataset to obtain multiple 2D images and the depth information and camera pose of each 2D image.
[0173] Step 408: Initialize the 3DGS model based on depth information, initial camera pose, and original image.
[0174] In this process, when determining the initial camera pose of the original image, the cloud generates a sparse point cloud, initializes the sparse point cloud, and obtains the initial 3DGS model. Depth information can constrain the initial position and scale of the Gaussian distribution, and loss functions can be designed based on depth information, as well as occlusion processing optimization based on depth information.
[0175] Step 410: Obtain the reconstruction loss of the target view in the previous update cycle; if the reconstruction loss is less than the loss threshold, determine the updated camera pose of the target view in the current update cycle based on the initial camera pose and reconstruction loss determined in the previous update cycle; if the reconstruction loss is greater than or equal to the loss threshold, discard the initial camera pose.
[0176] Specifically, the cloud determines the reconstruction loss of the target viewpoint in the previous update cycle based on the reconstructed image and the original image from the target viewpoint in the previous update cycle. If the reconstruction loss is less than the loss threshold, the updated camera pose in the target viewpoint in the current update cycle is determined based on the initial camera pose and reconstruction loss determined in the previous update cycle; if the reconstruction loss is greater than or equal to the loss threshold, the initial camera pose is discarded.
[0177] Step 412: If the update order of the previous update cycle is within the preset range, generate a reconstructed image of the 3DGS model in the target view based on the 3DGS model and the updated camera pose.
[0178] In this process, if the update order of the previous update cycle is within a preset range, the cloud generates a reconstructed image of the 3DGS model from the target viewpoint based on the 3DGS model and the updated camera pose.
[0179] Step 414: Based on the original image and the reconstructed image, determine the reconstruction loss of the target viewpoint in the current update cycle; based on the reconstruction loss, update the 3DGS model.
[0180] In this process, the cloud determines the reconstruction loss of the target viewpoint in the current update cycle based on the original image and the reconstructed image; and updates the 3DGS model based on the reconstruction loss.
[0181] Step 416: Encode the reconstructed 3DGS model to obtain encoded data.
[0182] After obtaining the 3DGS model, the cloud encodes the 3DGS model to obtain encoded data, and then sends the encoded data to the mobile device via end-to-cloud transmission.
[0183] Step 418: Decode the encoded data to obtain point cloud data.
[0184] The mobile device decodes the encoded data to obtain the point cloud data corresponding to the 3DGS model, including the position, covariance matrix, transparency, and color of each Gaussian point.
[0185] Step 420: Based on the point cloud data, perform mobile endpoint cloud reconstruction and rendering.
[0186] Among them, the mobile terminal performs 3DGS reconstruction and rendering based on point cloud data to obtain a 3DGS model.
[0187] Step 422: Perform 6DoF interaction on the 3DGS model to obtain updated video data.
[0188] After obtaining the 3DGS model, the mobile device can perform 6DoF interaction on the 3DGS model, and perform 2D image rendering on the 3DGS model from different perspectives to obtain video data.
[0189] In the embodiments described in this specification, the camera's rotation matrix and translation vector are introduced into the 3DGS model training. Feedback adjustment of the camera rotation matrix and translation vector are also introduced into the 3DGS model training. A threshold value for the maximum error of the effective camera viewpoint is also introduced into the 3DGS model training. By simultaneously introducing feedback optimization of the initial camera pose during the training of the Gaussian point cloud, information mutual exclusion caused by initial camera pose matching deviations is reduced, severely erroneous camera points are eliminated, and the reconstruction quality of single-object 3D reconstruction is improved.
[0190] Corresponding to the above method embodiments, this specification also provides embodiments of a three-dimensional Gaussian model update device. Figure 5A schematic diagram of a three-dimensional Gaussian model updating device according to one embodiment of this specification is shown. Figure 5 As shown, the device includes:
[0191] The acquisition module 510 is configured to acquire a 3D Gaussian model constructed based on a 3D scene, a reconstructed image of the 3D Gaussian model from the target view, and the original image. The 3D Gaussian model includes multiple 3D Gaussian points, and the reconstructed image is determined based on the multiple 3D Gaussian points and the updated camera extrinsic parameters from the target view. The updated camera extrinsic parameters are determined based on the reconstruction loss of the target view from the previous update cycle.
[0192] The determination module 520 is configured to determine the reconstruction loss of the target viewpoint in the current update cycle based on the original image and the reconstructed image.
[0193] The update module 530 is configured to update the 3D Gaussian model based on the reconstruction loss by updating the attribute information of multiple 3D Gaussian points.
[0194] In one optional embodiment of this specification, the acquisition module 510 is further configured to acquire the updated camera extrinsic parameters under the target viewpoint in the current update cycle; and, if the update order of the previous update cycle is within a preset range, generate a reconstructed image of the 3D Gaussian model under the target viewpoint based on the 3D Gaussian model and the updated camera extrinsic parameters.
[0195] In one optional embodiment of this specification, the acquisition module 510 is further configured to acquire the reconstruction loss of the target view in the previous update cycle; and based on the initial camera extrinsic parameters and reconstruction loss determined in the previous update cycle, determine the updated camera extrinsic parameters of the target view in the current update cycle.
[0196] In one optional embodiment of this specification, the acquisition module 510 is further configured to determine the updated camera extrinsic parameters under the target viewpoint in the current update cycle based on the initial camera extrinsic parameters and reconstruction loss determined in the previous update cycle when the reconstruction loss is less than the loss threshold; and to remove the initial camera extrinsic parameters when the reconstruction loss is greater than or equal to the loss threshold.
[0197] In one optional embodiment of this specification, the determining module 520 is further configured to replace the background of the original image and the background of the reconstructed image with the target background to obtain the target original image and the target reconstructed image; and to determine the reconstruction loss of the target viewpoint in the current update cycle based on the target original image and the target reconstructed image.
[0198] In one optional embodiment of this specification, the determining module 520 is further configured to acquire target background information of the target background; while keeping the foreground of the original image and the foreground of the reconstructed image unchanged, the background of the original image and the background of the reconstructed image are replaced based on the target background information to obtain the target original image and the target reconstructed image.
[0199] In one optional embodiment of this specification, the attribute information includes position; the update module 530 is further configured to determine the position gradient of any three-dimensional Gaussian point based on the reconstruction loss; perform density control on any three-dimensional Gaussian point if the position gradient is greater than the gradient threshold; and update the three-dimensional Gaussian model based on the density-controlled three-dimensional Gaussian point.
[0200] The above is a schematic scheme of a three-dimensional Gaussian model updating device according to this embodiment. It should be noted that the technical solution of this updating device and the technical solution of the above-described three-dimensional Gaussian model updating method belong to the same concept. For details not described in detail in the technical solution of the three-dimensional Gaussian model updating device, please refer to the description of the technical solution of the above-described three-dimensional Gaussian model updating method.
[0201] Figure 6 A structural block diagram of a computing device according to one embodiment of this specification is shown. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.
[0202] Computing device 600 also includes access device 640, which enables computing device 600 to communicate via one or more networks 660. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. Access device 640 may include one or more of any type of wired or wireless network interface (e.g., Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Networks (WLAN) interface, a Wi-MAX (World Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0203] In one embodiment of this specification, the above-described components of the computing device 600 and Figure 6 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 6 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0204] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server.
[0205] The processor 620 is used to execute computer programs / instructions, which, when executed by the processor, implement the steps of the above-described method for updating the three-dimensional Gaussian model.
[0206] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-described method for updating the three-dimensional Gaussian model belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-described method for updating the three-dimensional Gaussian model.
[0207] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described method for updating the three-dimensional Gaussian model.
[0208] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the above-described method for updating the three-dimensional Gaussian model. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the above-described method for updating the three-dimensional Gaussian model.
[0209] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method for updating the three-dimensional Gaussian model.
[0210] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-described method for updating a three-dimensional Gaussian model belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-described method for updating a three-dimensional Gaussian model.
[0211] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0212] Computer instructions include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Computer-readable media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in computer-readable media can be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0213] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0214] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0215] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A method for updating a three-dimensional Gaussian model, characterized in that, include: Acquire a 3D Gaussian model constructed based on a 3D scene, a reconstructed image of the 3D Gaussian model from a target viewpoint, and an original image. The 3D Gaussian model includes multiple 3D Gaussian points. The reconstructed image is determined based on the multiple 3D Gaussian points and the updated camera extrinsic parameters from the target viewpoint. The updated camera extrinsic parameters are determined based on the reconstruction loss of the target viewpoint in the previous update cycle. Based on the original image and the reconstructed image, determine the reconstruction loss of the target viewpoint in the current update cycle; Based on the reconstruction loss, the three-dimensional Gaussian model is updated by updating the attribute information of the plurality of three-dimensional Gaussian points.
2. The method according to claim 1, characterized in that, Obtaining the reconstructed image of the 3D Gaussian model from the target viewpoint includes: Obtain the updated camera extrinsic parameters from the target viewpoint during the current update cycle; If the update order of the previous update cycle is within a preset range, a reconstructed image of the 3D Gaussian model under the target viewpoint is generated based on the 3D Gaussian model and the updated camera extrinsic parameters.
3. The method according to claim 2, characterized in that, The process of obtaining the updated camera extrinsic parameters from the target viewpoint during the current update cycle includes: Obtain the reconstruction loss of the target view described in the previous update cycle; Based on the initial camera extrinsic parameters and reconstruction loss determined in the previous update cycle under the target viewpoint, the updated camera extrinsic parameters under the target viewpoint in the current update cycle are determined.
4. The method according to claim 3, characterized in that, The process of determining the updated camera extrinsics for the target viewpoint in the current update cycle based on the initial camera extrinsics and reconstruction loss determined in the previous update cycle includes: If the reconstruction loss is less than the loss threshold, the updated camera extrinsic parameters under the target viewpoint determined in the previous update cycle and the reconstruction loss are used to determine the updated camera extrinsic parameters under the target viewpoint in the current update cycle. If the reconstruction loss is greater than or equal to the loss threshold, the initial camera extrinsic parameters are discarded.
5. The method according to claim 1, characterized in that, The step of determining the reconstruction loss of the target viewpoint in the current update cycle based on the original image and the reconstructed image includes: The backgrounds of the original image and the reconstructed image are replaced with the target background to obtain the target original image and the target reconstructed image. Based on the original image of the target and the reconstructed image of the target, the reconstruction loss of the target viewpoint in the current update cycle is determined.
6. The method according to claim 5, characterized in that, The step of replacing the background of the original image and the background of the reconstructed image with the target background to obtain the target original image and the target reconstructed image includes: Obtain target background information; While keeping the foreground of the original image and the foreground of the reconstructed image unchanged, the background of the original image and the background of the reconstructed image are replaced based on the target background information to obtain the target original image and the target reconstructed image.
7. The method according to any one of claims 1 to 6, characterized in that, The attribute information includes location. The step of updating the 3D Gaussian model based on the reconstruction loss by updating the attribute information of the plurality of 3D Gaussian points includes: For any three-dimensional Gaussian point, the position gradient of the point is determined based on the reconstruction loss. When the position gradient is greater than the gradient threshold, density control is applied to any three-dimensional Gaussian point. The three-dimensional Gaussian model is updated based on the density-controlled three-dimensional Gaussian points.
8. A device for updating a three-dimensional Gaussian model, characterized in that, include: The acquisition module is configured to acquire a 3D Gaussian model constructed based on a 3D scene, a reconstructed image of the 3D Gaussian model from the target viewpoint, and the original image. The 3D Gaussian model includes multiple 3D Gaussian points, and the reconstructed image is determined based on the multiple 3D Gaussian points and the updated camera extrinsic parameters from the target viewpoint. The updated camera extrinsic parameters are determined based on the reconstruction loss of the target viewpoint in the previous update cycle. The determination module is configured to determine the reconstruction loss of the target viewpoint in the current update cycle based on the original image and the reconstructed image; The update module is configured to update the three-dimensional Gaussian model based on the reconstruction loss by updating the attribute information of the plurality of three-dimensional Gaussian points.
9. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 7.
11. A computer program product, characterized in that, Includes a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 7.