3D Gaussian model training method, map reconstruction method, equipment and medium

In 3D Gaussian model training, the normal information of the real image is combined with the depth information of the rendered image, and the normal information is supervised and trained, which solves the problem of poor geometric consistency in map reconstruction and improves the accuracy and consistency of map reconstruction.

CN120236167APending Publication Date: 2025-07-01NINGBO LOTUS ROBOTICS CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510368033.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the pure vision-based map reconstruction scheme, the geometric consistency of the reconstruction object is difficult to ensure, resulting in low map reconstruction accuracy, affecting subsequent downstream tasks.

Method used

Through the training method based on the 3D Gaussian model, the first normal information of the real image and the depth information of the rendered image are used to supervise the conversion of depth information into normal information, and the geometric consistency of the 3D Gaussian model is improved.

Benefits of technology

Improve the geometric consistency of the 3D Gaussian model reconstruction object, enhance the accuracy of map reconstruction, and improve support for subsequent downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236167A_ABST
    Figure CN120236167A_ABST
Patent Text Reader

Abstract

The invention provides a training method of a 3D Gaussian model, a map reconstruction method, equipment and a medium. The method comprises the following steps: acquiring first normal information of a real image based on the real image of a target scene under multiple visual angles, and carrying out sparse reconstruction on the target scene to obtain a sparse point cloud of the target scene; initializing a 3D Gaussian model based on the sparse point cloud, projecting the 3D Gaussian model to a pixel space to obtain rendered images of the target scene under multiple visual angles, and determining second normal information based on depth information of the rendered images; and taking the real image as an optimization target of the rendered image, taking the first normal information as an optimization target of the second normal information, and training the 3D Gaussian model to obtain a 3D Gaussian model for the target scene. According to the scheme, the depth information is converted into the normal information to serve as the supervision signal, the influence caused by the fact that the depth information of a real image does not comprise scale information is eliminated, and the geometric consistency of a 3D Gaussian model reconstruction object is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present application relate to the technical field of three-dimensional reconstruction, and in particular, to a training method for a 3D Gaussian model, a map reconstruction method, a device, and a medium. Background Art

[0002] With the rapid iteration and development of autonomous driving technology, the BEV Former (Bird's Eye View Transformer) technology for vehicle environment display and the Occupancy Network technology for obstacle detection have been successively deployed on mass-produced vehicles and have achieved rapid development. However, both the BEV Former technology and the Occupancy Network technology require a large amount of scene data for training.

[0003] Currently, most scenarios are constructed using lidar point clouds or pure vision solutions. However, with the increase in the deployment scale of mass-produced vehicles, the acquisition scale and efficiency of lidar-based data collection vehicles are difficult to meet the iterative development needs of in-vehicle models. Therefore, the pure vision-based map reconstruction solution has received increasing attention. However, in the pure vision-based map reconstruction solution, it is difficult to ensure the geometric consistency of the reconstruction objects, resulting in low map reconstruction accuracy and affecting subsequent downstream tasks. Summary of the Invention

[0004] The present application provides a training method for a 3D Gaussian model and a map reconstruction method to solve the deficiencies in the related art.

[0005] According to a first aspect of one or more embodiments of the present application, a training method for a 3D Gaussian model is provided. The method includes:

[0006] Based on real images of a target scene from multiple perspectives, obtain first normal information of the real images, and perform sparse reconstruction on the target scene to obtain a sparse point cloud of the target scene;

[0007] Initialize a 3D Gaussian model based on the sparse point cloud of the target scene, project multiple 3D Gaussian points in the 3D Gaussian model into pixel space to obtain rendered images of the target scene from the multiple perspectives, and determine second normal information of the rendered images based on depth information of the rendered images;

[0008] Use the real images as the optimization target of the rendered images, and use the first normal information as the optimization target of the second normal information, and train the 3D Gaussian model to obtain a 3D Gaussian model for the target scene.

[0009] According to a second aspect of one or more embodiments of the present application, a method for map reconstruction of a target scene is provided. The method includes:

[0010] Based on the internal and external parameters of cameras corresponding to multiple perspectives of the target scene, project multiple 3D Gaussian points in the 3D Gaussian model of the target scene into the pixel space to obtain rendered images of the target scene under the multiple perspectives. Among them, the 3D Gaussian model of the target scene is trained by the method described in the above first aspect;

[0011] Construct a scene map of the target scene according to the obtained rendered images.

[0012] According to a third aspect of one or more embodiments of the present application, an electronic device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method described in the embodiment of the above first aspect or implements the steps of the method described in the embodiment of the above second aspect.

[0013] According to a fourth aspect of one or more embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the method described in the embodiment of the above first aspect or implements the steps of the method described in the embodiment of the above second aspect.

[0014] According to a fifth aspect of one or more embodiments of the present application, a computer program product is provided, including a computer program. When the computer program is executed by a processor, it implements the steps of the method described in the embodiment of the above first aspect or implements the steps of the method described in the embodiment of the above second aspect.

[0015] As can be seen from the above technical solutions, in one or more embodiments of the present application, since the depth information of the real image does not include scale information, and the depth information of the rendered image includes scale information, therefore, the depth information of the real image is relative depth information, while the depth information of the rendered image is absolute depth information. If the depth information of the real image is used to supervise the depth information of the rendered image to train the 3D Gaussian model, the geometric consistency of the object reconstructed by the 3D Gaussian model will be poor. Considering the conversion relationship between depth information and normal information, therefore, the embodiments of the present application convert the depth information into normal information to supervise the training of the 3D Gaussian model, eliminating the influence brought by the fact that the depth information of the real image does not include scale information, and improving the geometric consistency of the object reconstructed by the 3D Gaussian model.

[0016] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0018] Figure 1 It is a schematic flowchart of a method for training a 3D Gaussian model provided by an exemplary embodiment.

[0019] Figure 2 It is a schematic flowchart of a method for training a 3D Gaussian model provided by an exemplary embodiment.

[0020] Figure 3 It is a schematic flowchart of a method for map reconstruction of a target scene provided by an exemplary embodiment.

[0021] Figure 4 It is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment.

[0022] Figure 5 It is a block diagram of the structure of a device for training a 3D Gaussian model provided by an exemplary embodiment.

[0023] Figure 6 It is a block diagram of the structure of a device for map reconstruction of a target scene provided by an exemplary embodiment. Detailed Description of the Embodiment

[0024] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0025] It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in the present application. In some other embodiments, the steps included in the method may be more or less than those described in the present application. In addition, a single step described in the present application may be decomposed into multiple steps for description in other embodiments; and multiple steps described in the present application may also be combined into a single step for description in other embodiments.

[0026] Next, one or more embodiments of the present application will be described in detail.

[0027] Figure 1 It is a flowchart of a method for training a 3D Gaussian model provided by an exemplary embodiment. As Figure 1 shown, the method may include the following steps:

[0028] S101. Based on the real images of the target scene from multiple perspectives, obtain the first normal information of the real images, and perform sparse reconstruction on the target scene to obtain the sparse point cloud of the target scene.

[0029] Among them, the target scene can be any scene, and the embodiments of the present application do not limit the target scene. In one embodiment, the target scene can be any scene passed through or experienced by the mobile platform during movement. For example, a parking lot scene, a vehicle collision scene, etc. Among them, the mobile platform is any movable device, such as a mass-produced vehicle, a data collection vehicle, a drone, etc., and the embodiments of the present application do not limit the mobile platform.

[0030] The real images of the target scene from multiple perspectives are images obtained by the camera shooting the target scene from multiple perspectives. In one embodiment, the real images of the target scene from multiple perspectives are multi-perspective images at the same moment, that is to say, these real images are obtained by shooting the target scene from different perspectives at the same time point. In another embodiment, the real images of the target scene from multiple perspectives are multi-perspective images within the same time period, that is to say, these real images are obtained by continuously or intermittently shooting from multiple perspectives within a time period. Optionally, a camera is configured on the mobile platform, and the real images of the target scene from multiple perspectives are pictures obtained by the camera during the movement of the mobile platform. Exemplarily, during the movement of the mobile platform, the configured camera shoots the current environment, and the images obtained during the target duration are stored as an image data packet. The real images of the target scene from multiple perspectives in step S101 above are the images in this image data packet. Among them, during the movement of the mobile platform, the camera can shoot the current environment at any image acquisition frequency. Among them, the image acquisition frequency can be a fixed value, for example, 10 Hz (hertz), 15 Hz (hertz), etc.; the image acquisition frequency can also change dynamically. For example, the image acquisition frequency of the camera changes with the movement speed of the mobile platform; the faster the movement speed of the mobile platform, the higher the image acquisition frequency; the slower the movement speed of the mobile platform, the lower the image acquisition frequency. The embodiments of the present application only exemplarily illustrate the image acquisition frequency and do not limit the image acquisition frequency.

[0031] It should be noted that the number of cameras configured on the mobile platform can be one or more, and the embodiments of the present application do not limit the number and position of the cameras on the mobile platform. Exemplarily, two cameras are configured in the front of the mobile platform, and one camera is configured in each of the front left, front right, rear left, rear right, and rear of the mobile platform to realize the observation of the surrounding environment of the mobile platform.

[0032] In the embodiments of the present application, only the first normal information and the second normal information are used to distinguish the normal information obtained by different methods, and the content of the normal information is not limited. Among them, the first normal information is the normal information of the real image, and the second normal information is the normal information determined based on the depth information of the rendered image. Optionally, the normal information of the image includes the normal corresponding to any pixel point in the image.

[0033] In one embodiment, the first normal information is indirectly obtained based on the depth information of the real image. Among them, obtaining the first normal information of the real image includes: performing monocular depth estimation on the real image to obtain the depth information of the real image; and determining the first normal information of the real image based on the depth information of the real image. Exemplarily, performing monocular depth estimation on the real image to obtain the depth map of the real image, calculating the gradient at each pixel point in the depth map to estimate the normal corresponding to the pixel point, and obtaining the first normal information of the real image.

[0034] In another embodiment, the first normal information is directly obtained based on the real image. Among them, obtaining the first normal information of the real image includes: inputting the real image into a pre-trained neural network model so that the neural network model outputs the first normal information of the real image. The neural network model will learn the mapping relationship from the input image to the output normal during the training process. Therefore, when the real image is input into the neural network model, the neural network model can output the first normal information of the real image. Exemplarily, inputting the real image into a pre-trained neural network model so that the neural network model outputs the normal map of the real image.

[0035] The sparse point cloud of the target scene is used to represent the sparse geometric structure of the target scene. In one embodiment, the sparse point cloud includes three-dimensional spatial points with obvious features in the target scene, such as the edge points and corners of a building. Exemplarily, if there is a table in the target scene, the sparse point cloud of the target scene may include points with obvious features such as the four corners of the table.

[0036] Any sparse reconstruction algorithm can be used for sparse reconstruction of the target scene. For example, algorithms such as Structure from Motion (SfM) algorithm and Multi-View Stereo (MVS) algorithm. The embodiments of the present application do not limit the sparse reconstruction algorithm.

[0037] S102. Initialize a 3D Gaussian model based on the sparse point cloud of the target scene, project multiple 3D Gaussian points in the 3D Gaussian model into the pixel space to obtain the rendered images of the target scene from multiple perspectives, and determine the second normal information of the rendered images based on the depth information of the rendered images.

[0038] The 3D Gaussian model includes multiple 3D Gaussian points, and each 3D Gaussian point corresponds to an ellipsoid (therefore, the 3D Gaussian point can also be called a 3D Gaussian sphere). Initializing the 3D Gaussian model based on the sparse point cloud of the target scene is to endow the attributes of the points in the sparse point cloud to the 3D Gaussian points in the 3D Gaussian model. That is, each point in the sparse point cloud is used to initialize a 3D Gaussian point, so as to obtain multiple 3D Gaussian points. These multiple 3D Gaussian points constitute the 3D Gaussian model. Therefore, initializing the 3D Gaussian model based on the sparse point cloud of the target scene can also be regarded as converting the sparse point cloud into a 3D Gaussian model.

[0039] In one embodiment, initializing the 3D Gaussian model based on the sparse point cloud of the target scene includes: initializing a 3D Gaussian point for each point in the sparse point cloud (that is, generating a 3D Gaussian point for each point in the sparse point cloud), initializing the Gaussian position of the 3D Gaussian point with the point coordinates of this point (that is, using the point coordinates of this point as the mean vector of the 3D Gaussian point), initializing the scale information of the 3D Gaussian point based on the minimum distance between this point and its adjacent points (exemplarily, the minimum distance between this point and its adjacent points is denoted as min_dis, and the scale information is denoted as [min_dis, min_dis, 0]. That is, the horizontal scale of the 3D Gaussian point is the minimum horizontal distance between this point and its adjacent points, the vertical scale of the 3D Gaussian point is the minimum vertical distance between this point and its adjacent points, and the vertical scale of the 3D Gaussian point is 0), initializing the spherical harmonic function of the 3D Gaussian point based on the color information of this point, initializing the pose information of the 3D Gaussian point as the identity matrix (that is, using the identity matrix as the pose information of the 3D Gaussian point), and initializing the opacity of the 3D Gaussian point to the default value (exemplarily, setting the opacity of the 3D Gaussian point to 0.1).

[0040] It should be noted that the embodiments of this application only exemplarily illustrate the process of "initializing the 3D Gaussian model based on the sparse point cloud of the target scene". Which attributes are inherited from the sparse point cloud and which attributes use the default values during the initialization process can be set according to actual needs, and the embodiments of this application do not limit this.

[0041] In one embodiment, the real images under multiple viewpoints are obtained by the camera shooting under different internal and external camera parameters. When projecting the multiple 3D Gaussian points in the 3D Gaussian model into the pixel space to obtain the rendered images of the target scene under multiple viewpoints, the projection can be performed based on the internal and external camera parameters corresponding to the multiple viewpoints, so that the projected rendered images and the corresponding real images belong to the same viewpoint.

[0042] Optionally, the internal and external camera parameters include the internal camera parameters and the external camera parameters; the internal camera parameters are used to describe the geometric and / or optical characteristics inside the camera. For example, the internal camera parameters include the focal length, the principal point, the distortion coefficient, etc. The external camera parameters are used to describe the pose of the camera. For example, the external camera parameters include the camera position and the camera attitude.

[0043] In one embodiment, when the camera captures real images of the target scene from multiple perspectives, the camera records the current internal and external camera parameters each time it captures an image, stores the internal and external camera parameters corresponding to the real images, and when projecting multiple 3D Gaussian points in the 3D Gaussian model into the pixel space, any stored internal and external camera parameters can be used to project the multiple 3D Gaussian points in the 3D Gaussian model into the pixel space to obtain a rendered image. The rendered image and the real image corresponding to the internal and external camera parameters are images from the same perspective, and the 3D Gaussian model can be trained based on the difference between the real image and the rendered image, that is, the attribute parameters of the 3D Gaussian points are optimized.

[0044] In another embodiment, the internal and external camera parameters corresponding to multiple perspectives can be obtained by performing pose estimation on the real images from multiple perspectives. Optionally, the internal and external camera parameters corresponding to the multiple perspectives are obtained during the sparse reconstruction of the target scene. That is to say, during the sparse reconstruction of the target scene, not only the sparse point cloud of the target scene is obtained, but also the internal and external camera parameters corresponding to multiple perspectives are obtained.

[0045] Exemplarily, using the structure from motion algorithm, the sparse point cloud of the target scene and the camera pose are estimated from the real images of the target scene from multiple perspectives, and the sparse point cloud of the target scene and the internal and external camera parameters corresponding to each real image are obtained. It should be noted that one real image corresponds to one perspective, and the internal and external camera parameters corresponding to multiple perspectives are also the internal and external camera parameters corresponding to multiple real images.

[0046] Optionally, the calculation of the internal and external camera parameters corresponding to the multiple perspectives is performed independently of the calculation of the sparse reconstruction. The embodiments of the present application do not limit the calculation timing and calculation method of the internal and external camera parameters corresponding to the multiple perspectives.

[0047] Next, an exemplary description is given of the process of "projecting multiple 3D Gaussian points in the 3D Gaussian model into the pixel space to obtain a rendered image of the target scene from any perspective".

[0048] In one embodiment, the expression of the 3D Gaussian point is:

[0049]

[0050] Σ = RSS T R T (2)

[0051] Among them, G(x) represents the Gaussian distribution of the 3D Gaussian point at the x position, x represents any position in the target scene, u represents the Gaussian position (i.e., the mean vector of this 3D Gaussian point); R represents the pose information of the 3D Gaussian point, S represents the scale information of the 3D Gaussian point, T represents the transpose of the matrix, and Σ represents the covariance matrix of the 3D Gaussian point.

[0052] Obtain the internal and external camera parameters corresponding to any perspective, filter out the 3D Gaussian points within this perspective according to the internal and external camera parameters, and perform differentiable rasterization on the filtered 3D Gaussian points. Specifically:

[0053] Based on the internal and external camera parameters corresponding to any perspective, draw a ray from the camera origin to a certain pixel point in the real image under this perspective, sort the 3D Gaussian points intersecting with this ray according to depth, and at the same time calculate the intersection position of each 3D Gaussian point and this ray. This intersection position is x in the above formula (1); project the 3D Gaussian points intersecting with this ray into the pixel space to obtain the two-dimensional Gaussian distribution G′(x′):

[0054]

[0055] ∑' = JW∑W T J T (4)

[0056] Among them, x′ represents the pixel point coordinates, u′ is the image pixel point coordinates generated by the 3D Gaussian position u through perspective transformation and projection transformation, W is the gradient information of the perspective transformation, J is the gradient information of the image projection transformation, and Σ′ represents the two-dimensional covariance matrix of the 3D Gaussian point.

[0057] Among them, the opacity of the projected 2D Gaussian point can be determined by the opacity of the corresponding 3D Gaussian point and the two-dimensional Gaussian distribution. Specifically: α′ = αG′(x′). Among them, α′ represents the opacity of the 2D Gaussian point, α represents the opacity of the 3D Gaussian point, and G′(x′) represents the two-dimensional Gaussian distribution of the 3D Gaussian point.

[0058] After that, grid-based rasterization sorts the 2D Gaussian points according to depth and uses the weighted average method to render the final color, depth, and normal to obtain the rendered image.

[0059]

[0060] Among them, N represents the number of 3D Gaussian points on the ray, c i represents the color of the i-th 3D Gaussian point, d i represents the depth of the i-th 3D Gaussian point, m iIt represents the normal of the i-th 3D Gaussian point. The color of the 3D Gaussian point can be obtained through the spherical harmonic function of the color of the 3D Gaussian point. The normal of the 3D Gaussian point can be obtained through the pose information of the 3D Gaussian point (this normal is the eigenvector corresponding to the minimum eigenvalue of the 3D Gaussian point, that is, the direction of the shortest axis of the ellipsoid corresponding to the 3D Gaussian point), and the depth can be obtained from the intersection position of the ray and the 3D Gaussian point.

[0061] Through the above method, by successively using the internal and external camera parameters corresponding to each perspective, projecting multiple 3D Gaussian points in the 3D Gaussian model into the pixel space, the rendered images of the target scene in each perspective can be obtained.

[0062] It should be noted that during the process of generating the rendered image, the normal of the rendered image is determined based on the normal of the 3D Gaussian point used for projection. However, the second normal information in step S102 above is not determined based on the normal of the 3D Gaussian point used for projection, but is determined based on the depth information of the rendered image. The depth information of the rendered image is determined based on the intersection position of the ray and the 3D Gaussian point.

[0063] In one embodiment, the depth information of the rendered image includes the depth of each pixel point in the rendered image, and the second normal information of the rendered image includes the normal of each pixel point in the rendered image. Determining the second normal information of the rendered image based on the depth information of the rendered image includes: calculating the gradient at each pixel point in the rendered image based on the depth of each pixel point in the rendered image, and estimating the normal corresponding to the pixel point based on the gradient at each pixel point to obtain the second normal information of the rendered image.

[0064] S103. Use the real image as the optimization target of the rendered image, and use the first normal information as the optimization target of the second normal information to train the 3D Gaussian model to obtain the 3D Gaussian model for the target scene.

[0065] In one embodiment, the 3D Gaussian model is trained based on only one perspective each time. First, using the internal and external camera parameters of one perspective, project multiple 3D Gaussian points in the 3D Gaussian model into the pixel space to obtain the rendered image under this perspective. Use the real image under this perspective as the optimization target for the rendered image under this perspective, and use the first normal information of this real image as the optimization target for the second normal information of this rendered image to train the 3D Gaussian model. Next, use the internal and external camera parameters of the next perspective to project multiple 3D Gaussian points in the 3D Gaussian model obtained from the previous training into the pixel space to obtain the rendered image under this perspective. Use the real image under this perspective as the optimization target for the rendered image under this perspective, and use the first normal information of this real image as the optimization target for the second normal information of this rendered image to train the 3D Gaussian model. Repeat the above steps until the trained 3D Gaussian model can project accurate rendered images under each perspective.

[0066] In another embodiment, the 3D Gaussian model is trained based on multiple perspectives together. First, using the internal and external camera parameters of multiple perspectives, project multiple 3D Gaussian points in the 3D Gaussian model into the pixel space to obtain the rendered images under these multiple perspectives. Use the real images under these multiple perspectives as the optimization targets for the rendered images under these multiple perspectives, and use the first normal information of the real images under these multiple perspectives as the optimization targets for the second normal information of the rendered images under these multiple perspectives to train the 3D Gaussian model. Next, again use the internal and external camera parameters of these multiple perspectives to project multiple 3D Gaussian points in the 3D Gaussian model obtained from the previous training into the pixel space to obtain the rendered images under these multiple perspectives. Use the real images under these multiple perspectives as the optimization targets for the rendered images under these multiple perspectives, and use the first normal information of the real images under these multiple perspectives as the optimization targets for the second normal information of the rendered images under these multiple perspectives to train the 3D Gaussian model. Repeat the above steps until the trained 3D Gaussian model can project accurate rendered images under each perspective.

[0067] Among them, when using the real images under these multiple perspectives as the optimization targets for the rendered images under these multiple perspectives, the real image under the same perspective is used as the optimization target for the rendered image under this perspective. That is to say, use the real image under perspective A as the optimization target for the rendered image under perspective A, use the real image under perspective B as the optimization target for the rendered image under perspective B, use the real image under perspective C as the optimization target for the rendered image under perspective C, and so on.

[0068] In some training methods of 3D Gaussian models, the depth information of the real image is used as the optimization target for the depth information of the rendered image to train the 3D Gaussian model. However, the depth information of the real image is obtained by methods such as monocular depth estimation of the real image. This depth information is relative depth information and does not include scale information. That is, the obtained depth information can only reflect the relative distance relationship between points in the target scene, and cannot provide the exact distance between points in the target scene. However, the depth information of the rendered image is determined based on the intersection position of the ray and the 3D Gaussian model. Therefore, the depth information of the rendered image is not relative depth information, but absolute depth information, which includes scale information. That is, the depth information of the rendered image can provide the exact distance between points in the target scene. It can be seen that if the depth information of the real image is used as the optimization target for the depth information of the rendered image, due to the fact that the depth information of the rendered image includes scale information while the depth information of the real image does not include scale information, there will be a problem of scale inconsistency between the two. This scale inconsistency will lead to deviations in the optimization process, making it difficult for the trained 3D Gaussian model to accurately reflect the geometric structure and size ratio in the real scene, resulting in a problem of poor geometric consistency of the reconstructed object.

[0069] In the embodiments of the present application, considering that there is a conversion relationship between depth information and normal information, and normal information is not affected by scale information, therefore, the embodiments of the present application convert depth information into normal information to supervise the training of the 3D Gaussian model, eliminating the influence caused by the fact that the depth information of the real image does not include scale information, and improving the geometric consistency of the object reconstructed by the 3D Gaussian model. Among them, converting depth information into normal information to supervise the training of the 3D Gaussian model means using the first normal information of the real image as the optimization target for the second normal information of the rendered image.

[0070] In one embodiment, using the real image as the optimization target for the rendered image and using the first normal information as the optimization target for the second normal information to train the 3D Gaussian model to obtain a 3D Gaussian model for the target scene includes: determining a first loss value based on the difference between the rendered image and the real image, and determining a second loss value based on the difference between the first normal vector and the second normal vector; based on the first loss value and the second loss value, determining the prediction loss value of the 3D Gaussian model, and updating the model parameters of the 3D Gaussian model according to the prediction loss value until the termination condition is met, to obtain a 3D Gaussian model for the target scene.

[0071] Among them, the termination condition can be any condition such as the prediction loss value converges, the number of training iterations reaches a specified number, the prediction loss value is less than a preset value, etc. The embodiments of the present application do not limit the termination condition. When determining the first loss value, any loss function can be used, and the embodiments of the present application do not limit this. When determining the second loss value, any loss function can also be used, and the embodiments of the present application do not limit this.

[0072] Optionally, based on the first loss value and the second loss value, determining the prediction loss value of the 3D Gaussian model includes: performing weighted averaging on the first loss value and the second loss value to obtain the prediction loss value of the 3D Gaussian model. Among them, the weights of the first loss value and the second loss value can be hyperparameters, empirical values, etc. The embodiments of the present application do not limit the weights of the first loss value and the second loss value.

[0073] In a possible implementation manner, on the basis of using the depth information of the rendered image for supervised training, the normal information of the rendered image can also be introduced as an additional supervision signal to further optimize and improve the training effect of the 3D Gaussian model. Since the first normal information of the real image has been used as the optimization target for the second normal information of the rendered image in the embodiments of the present application, the first normal information is no longer suitable as the optimization target for the third normal information of the rendered image. Considering the inherent geometric relationship between the depth information and the normal information, the embodiments of the present application perform self-supervision based on the depth information of the rendered image and the third normal information of the rendered image. Among them, using the real image as the optimization target for the rendered image and using the first normal information as the optimization target for the second normal information to train the 3D Gaussian model to obtain a 3D Gaussian model for the target scene includes: using the real image as the optimization target for the rendered image, using the first normal information as the optimization target for the second normal information, and performing self-supervision based on the depth information of the rendered image and the third normal information of the rendered image to train the 3D Gaussian model to obtain a 3D Gaussian model for the target scene; among them, the third normal information of the rendered image is determined based on the direction of the shortest axis of the ellipsoid corresponding to multiple 3D Gaussian points.

[0074] Among them, performing self-supervision based on the depth information of the rendered image and the third normal information of the rendered image can be to determine the second normal information of the rendered image based on the depth information of the rendered image, and perform self-supervision based on the second normal information of the rendered image and the third normal information of the rendered image.

[0075] In one embodiment, a real image is used as an optimization target for a rendered image, a first normal information is used as an optimization target for a second normal information, and self-supervision is performed based on the depth information of the rendered image and the third normal information of the rendered image to train a 3D Gaussian model to obtain a 3D Gaussian model for a target scene, including: determining a first loss value based on the difference between the rendered image and the real image, determining a second loss value based on the difference between the first normal vector and the second normal vector, and determining a third loss value based on the difference between the second normal vector and the third normal vector; determining a prediction loss value of the 3D Gaussian model based on the first loss value, the second loss value, and the third loss value, and updating the model parameters of the 3D Gaussian model according to the prediction loss value until a termination condition is met to obtain a 3D Gaussian model for the target scene.

[0076] Wherein, any loss function can be used to determine the third loss value, and the embodiments of the present application do not limit this.

[0077] Optionally, determining a prediction loss value of the 3D Gaussian model based on the first loss value, the second loss value, and the third loss value includes: performing weighted averaging on the first loss value, the second loss value, and the third loss value to obtain a prediction loss value of the 3D Gaussian model. Wherein, the first loss value, the second loss value, and the third loss value can be hyperparameters, empirical values, etc., and the embodiments of the present application do not limit this.

[0078] Exemplarily, an image loss value L1 and an image structural similarity loss value L ssim (L1 + λ ssim L ssim is the first loss value) are determined based on the difference between the rendered image and the real image, a cosine loss value L normal (which is also the second loss value) is determined based on the difference between the first normal vector and the second normal vector, and a self-supervised constraint loss value L self-super (which is also the third loss value) is determined based on the difference between the second normal vector and the third normal vector. The prediction loss value is: L = L1 + λ ssim L ssim + λ ssself-superm L self-super + λ normal L normal . Wherein, λ ssim is the weight of L ssim , λss self-superm is the weight of L self-super , λ normal is the weight of L normal , and λ ssim , λ ssself-superm , λ normal are hyperparameters.

[0079] As can be seen from the above technical solutions, in one or more embodiments of the present application, since the depth information of the real image does not include scale information, and the depth information of the rendered image includes scale information, therefore, the depth information of the real image is relative depth information, while the depth information of the rendered image is absolute depth information. If the depth information of the real image is used to supervise the depth information of the rendered image to train the 3D Gaussian model, the geometric consistency of the object reconstructed by the 3D Gaussian model will be poor. Considering the conversion relationship between depth information and normal information, therefore, the embodiments of the present application convert the depth information into normal information to supervise the training of the 3D Gaussian model, eliminating the influence caused by the fact that the depth information of the real image does not include scale information, and improving the geometric consistency of the object reconstructed by the 3D Gaussian model.

[0080] When reconstructing based on the 3D Gaussian model, the reconstructed dynamic objects often appear blurred or ghosted. To address this problem, the embodiments of the present application provide a solution, which is illustrated exemplarily through the Figure 2 embodiments shown. Figure 2 FIG. is a flowchart of a method for training a 3D Gaussian model provided for an exemplary embodiment. As Figure 2 shown, the method may include the following steps:

[0081] S201. Based on the real images of the target scene from multiple perspectives, obtain the two-dimensional projection positions of the dynamic objects in the target scene on the real images.

[0082] Among them, the dynamic object is an object that is moving in the target scene. For example, vehicles, flying birds, pedestrians, etc. The embodiments of the present application do not limit the type of dynamic objects. Based on the real images of the target scene from multiple perspectives, obtaining the two-dimensional projection positions of the dynamic objects in the target scene on the real images means: for each real image, determine the position of the dynamic object in the real image.

[0083] In one embodiment, the method of instance segmentation can be used to obtain the two-dimensional projection positions of the dynamic objects in the target scene on the real images. Among them, instance segmentation can not only distinguish different object categories, but also generate accurate bounding boxes for each object instance. Exemplarily, using mask2former to perform instance segmentation on the real images of the target driving scene from multiple perspectives can output the two-dimensional projection positions of multiple object instances such as roads, skies, vehicles, etc. on the real images.

[0084] In another embodiment, the method of object detection can be used to obtain the two-dimensional projection positions of the dynamic objects in the target scene on the real images. Among them, the object detection algorithm can identify and locate the objects of specific categories in the image and return the bounding boxes of these objects.

[0085] It should be noted that the embodiments of the present application only exemplarily illustrate the method for obtaining the two-dimensional projection position of a dynamic object in a real image in a target scene. Of course, other methods can also be adopted in the embodiments of the present application, such as the optical flow method, stereo vision, etc. The embodiments of the present application do not limit this.

[0086] S202. Based on the real images of the target scene from multiple perspectives, obtain the first normal information of the real images, and perform sparse reconstruction on the target scene to obtain the sparse point cloud of the target scene.

[0087] The above step S202 is the same as the above step S101, and reference can be made to the above step S101, which will not be elaborated here one by one.

[0088] S203. Initialize a 3D Gaussian model based on the sparse point cloud of the target scene, project multiple 3D Gaussian points in the 3D Gaussian model into the pixel space to obtain the rendered images of the target scene from multiple perspectives, and determine the second normal information of the rendered images based on the depth information of the rendered images.

[0089] The above step S203 is the same as the above step S102, and reference can be made to the above step S102, which will not be elaborated here one by one.

[0090] S204. Use the real image as the optimization target of the rendered image, use the first normal information as the optimization target of the second normal information, and use the two-dimensional projection position of the dynamic object as the optimization target of the projection position of the 3D Gaussian point corresponding to the dynamic object in the rendered image, and train the 3D Gaussian model to obtain a 3D Gaussian model for the target scene.

[0091] When reconstructing based on the 3D Gaussian model, the reconstructed dynamic object often appears blurred or ghosted. To solve this problem, in the embodiments of the present application, the two-dimensional projection position of the dynamic object is used as the optimization target of the projection position of the 3D Gaussian point corresponding to the dynamic object in the rendered image, so that the 3D Gaussian model can accurately project the dynamic object.

[0092] In one embodiment, the 3D Gaussian point corresponding to the dynamic object is the 3D Gaussian point located at the three-dimensional spatial position of the dynamic object. In one embodiment, the method further includes: obtaining the three-dimensional spatial position of the dynamic object in the target scene; determining the 3D Gaussian point located at the three-dimensional spatial position as the 3D Gaussian point corresponding to the dynamic object. Among them, the 3D Gaussian points located at the three-dimensional spatial position include all the 3D Gaussian points located at the three-dimensional spatial position and some of the 3D Gaussian points located at the three-dimensional spatial position.

[0093] Optionally, during the movement of the mobile platform, not only real images of the target scene are captured by the camera, but also the obstacle perception module performs obstacle recognition on the target scene to obtain the three-dimensional spatial positions of dynamic objects (the output result of the obstacle perception module can be the 3D bounding box of the dynamic object including temporal information). The acquisition of the three-dimensional spatial positions of dynamic objects in the target scene in step S204 above is to obtain the output result of the obstacle perception module. Among them, when the obstacle perception module performs obstacle recognition on the target scene, it can be based on the real images captured by the camera or not. When the obstacle perception module does not perform obstacle recognition based on the real images captured by the camera, it is necessary to ensure that the working frequency of the obstacle perception module is the same as that of the camera to achieve that the results of the two are for the same moment and the same perspective.

[0094] In another embodiment, a key point detection algorithm is used to identify key feature points on the dynamic object, and then these key points are used as the basis for determining the 3D Gaussian points corresponding to the dynamic object. Of course, other methods can also be used to determine the 3D Gaussian points corresponding to the dynamic object, and the embodiments of the present application do not limit the method for determining the 3D Gaussian points corresponding to the dynamic object.

[0095] In one embodiment, the real image is used as the optimization target of the rendered image, the first normal information is used as the optimization target of the second normal information, and the two-dimensional projection position of the dynamic object is used as the optimization target of the projection position of the 3D Gaussian point corresponding to the dynamic object in the rendered image to train the 3D Gaussian model to obtain the 3D Gaussian model for the target scene, including: determining a first loss value based on the difference between the rendered image and the real image, determining a second loss value based on the difference between the first normal vector and the second normal vector, and determining a fourth loss value based on the difference between the two-dimensional projection position of the dynamic object and the projection position of the 3D Gaussian point corresponding to the dynamic object in the rendered image; determining the prediction loss value of the 3D Gaussian model based on the first loss value, the second loss value, and the fourth loss value, and updating the model parameters of the 3D Gaussian model according to the prediction loss value until the termination condition is met to obtain the 3D Gaussian model for the target scene.

[0096] Among them, any loss function can be used to determine the fourth loss value, and the embodiments of the present application do not limit this.

[0097] It should be noted that Figure 1 The embodiments shown Figure 2 can be arbitrarily combined, and the embodiments of the present application do not limit this.

[0098] It should be noted another point that in Figure 1 and Figure 2In the illustrated embodiment, when training a 3D Gaussian model, not only can the attribute parameters of 3D Gaussian points be adjusted, but 3D Gaussian points can also be copied, split, or deleted. Next, an exemplary description of the process of copying, splitting, or deleting 3D Gaussian points will be given:

[0099] During the process of training a 3D Gaussian model, calculate the average gradient of 3D Gaussian points, and statistically calculate the average gradient of the position gradients that are relatively large during N training processes. If this average gradient is greater than a first threshold (such as 0.001) and the maximum scale of the 3D Gaussian point corresponding to this average gradient is less than a second threshold (such as 0.1), then copy the 3D Gaussian point corresponding to this average gradient. If this average gradient is greater than a first threshold (such as 0.001) and the maximum scale of the 3D Gaussian point corresponding to this average gradient is greater than a second threshold (such as 0.1), then split the 3D Gaussian point corresponding to this average gradient. The splitting method is to sample along the maximum scale direction to generate two Gaussian points. To prevent floating Gaussians in the area near the camera, it can be determined whether the opacity of the 3D Gaussian point is less than a third threshold (such as 0.01), and delete the 3D Gaussian points with opacity less than the third threshold. Additionally, after a certain number of training iterations, if there are some 3D Gaussian points with particularly large scales, then these 3D Gaussian points will cause the problem of blurred rendered images. Therefore, it can be determined whether the maximum scale of the 3D Gaussian point is greater than a fourth threshold (such as 1.0), and delete the 3D Gaussian points with the maximum scale greater than the fourth threshold.

[0100] In the above technical solution, for the problem of blurring or ghosting in the reconstructed dynamic object, when training a 3D Gaussian model, the two-dimensional projection position of the dynamic object is used as the optimization target for the projection position of the 3D Gaussian point corresponding to the dynamic object in the rendered image, so that the 3D Gaussian model can accurately project the dynamic object, effectively alleviating the situation of blurring or ghosting in the reconstructed dynamic object.

[0101] Figure 3 It is a flowchart of a method for map reconstruction of a target scene provided for an exemplary embodiment. As Figure 3 shown, the method may include the following steps:

[0102] S301. Based on the internal and external parameters of the cameras corresponding to multiple perspectives of the target scene, project multiple 3D Gaussian points in the 3D Gaussian model for the target scene into the pixel space to obtain the rendered images of the target scene from multiple perspectives. Among them, the 3D Gaussian model for the target scene is trained by the method described in the embodiment shown in Figure 1 or Figure 2 the embodiment shown above.

[0103] It should be noted that when projecting multiple 3D Gaussian points in the 3D Gaussian model for the target scene into the pixel space, not only can a rendered image be generated, but also a depth map and a normal map can be generated. Among them, whether to generate the depth map and the normal map can be determined according to actual needs, and the embodiments of the present application do not limit this.

[0104] S302. Construct a scene map of the target scene according to the obtained rendered image.

[0105] By generating meshes for the rendered images from multiple viewpoints or for the rendered images and depth maps from multiple viewpoints, the scene map of the target scene can be obtained. Among them, the scene map of the target scene is a 3D scene map of the target scene.

[0106] In one embodiment, the scene map of the target scene can be uniformly sampled to generate point cloud information including color and coordinates for subsequent manual annotation / auto-annotation.

[0107] In the above technical solutions, by improving the accuracy of the 3D Gaussian model of the target scene, a scene map with higher reconstruction accuracy can be further constructed based on the 3D Gaussian model of the target scene, reducing the impact on subsequent downstream tasks.

[0108] It should be noted that the user permission information, user information (including but not limited to user login accounts, user login passwords, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this specification are all information and data authorized by the user or fully authorized by all parties. And the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0109] Corresponding to the above method embodiments, the present application also provides an embodiment of an xx device.

[0110] Figure 4 It is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of the present application. Refer to Figure 4 , at the hardware level, the electronic device includes a processor 402, an internal bus 404, a network interface 406, a memory 408, and a non-volatile memory 410. Of course, other hardware required for other services may also be included. The processor 402 reads the corresponding computer program from the non-volatile memory 410 into the memory 408 and then runs it. Of course, in addition to the software implementation, the present application does not exclude other implementation methods, such as logic devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or logic devices.

[0111] Figure 5 It is a block diagram of a training device for a 3D Gaussian model shown according to an exemplary embodiment of the present application. Referring to Figure 5 , the device includes:

[0112] An acquisition unit 501, configured to obtain first normal information of the real image based on real images of a target scene from multiple perspectives;

[0113] A reconstruction unit 502, configured to perform sparse reconstruction on the target scene to obtain a sparse point cloud of the target scene;

[0114] An initialization unit 503, configured to initialize a 3D Gaussian model based on the sparse point cloud of the target scene;

[0115] A projection unit 504, configured to project multiple 3D Gaussian points in the 3D Gaussian model into the pixel space to obtain rendered images of the target scene from the multiple perspectives;

[0116] The acquisition unit 501 is further configured to determine second normal information of the rendered image based on depth information of the rendered image;

[0117] A training unit 505, configured to use the real image as an optimization target for the rendered image, and use the first normal information as an optimization target for the second normal information, and train the 3D Gaussian model to obtain a 3D Gaussian model for the target scene.

[0118] Optionally, the training unit 505 is configured to use the real image as an optimization target for the rendered image, use the first normal information as an optimization target for the second normal information, and perform self-supervision based on depth information of the rendered image and third normal information of the rendered image, and train the 3D Gaussian model to obtain a 3D Gaussian model for the target scene;

[0119] Wherein, the third normal information of the rendered image is determined based on the direction of the shortest axis of the ellipsoid corresponding to the multiple 3D Gaussian points.

[0120] Optionally, the training unit 505 is configured to determine a first loss value based on a difference between the rendered image and the real image, determine a second loss value based on a difference between the first normal vector and the second normal vector, and determine a third loss value based on a difference between the second normal vector and the third normal vector; determine a prediction loss value of the 3D Gaussian model based on the first loss value, the second loss value, and the third loss value, and update model parameters of the 3D Gaussian model according to the prediction loss value until a termination condition is met, to obtain a 3D Gaussian model for the target scene.

[0121] Optionally, an acquisition unit 501 is configured to perform monocular depth estimation on the real image to obtain depth information of the real image; and determine first normal information of the real image based on the depth information of the real image.

[0122] Optionally, the acquisition unit 501 is further configured to obtain two-dimensional projection positions of dynamic objects in the target scene on the real image based on real images of the target scene from multiple perspectives;

[0123] A training unit 505 is configured to use the real image as an optimization target for the rendered image, use the first normal information as an optimization target for the second normal information, and use the two-dimensional projection positions of the dynamic objects as optimization targets for the projection positions of the corresponding 3D Gaussian points of the dynamic objects in the rendered image, and train the 3D Gaussian model to obtain a 3D Gaussian model for the target scene.

[0124] Optionally, the acquisition unit 501 is further configured to obtain three-dimensional spatial positions of dynamic objects in the target scene; and determine 3D Gaussian points located in the three-dimensional spatial positions as the 3D Gaussian points corresponding to the dynamic objects, where the 3D Gaussian points located in the three-dimensional spatial positions include all 3D Gaussian points located in the three-dimensional spatial positions and partial 3D Gaussian points located in the three-dimensional spatial positions.

[0125] Figure 6 It is a block diagram of a map reconstruction device for a target scene shown according to an exemplary embodiment of the present application. Refer to Figure 6 and the device includes:

[0126] A projection unit 601 is configured to project multiple 3D Gaussian points in a 3D Gaussian model for the target scene into pixel space based on internal and external camera parameters corresponding to multiple perspectives of the target scene to obtain rendered images of the target scene from the multiple perspectives, where the 3D Gaussian model for the target scene is trained by the method described in the embodiment shown in Figure 1 or Figure 2 the method described in the embodiment shown;

[0127] A construction unit 602 is configured to construct a scene map of the target scene according to the obtained rendered images.

[0128] The implementation processes of the functions and roles of each module in the above device are specifically detailed in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.

[0129] The devices or modules illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, laptop computer, cellular phone, camera phone, smart phone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or a combination of any several of these devices.

[0130] In a typical configuration, a computer includes one or more processors, including a central processing unit (CPU) and a graphics processing unit (GPU), an input / output interface, a network interface, and memory. Among them, the central processing unit is used for computing simulation, and the graphics processing unit is used for outputting high-quality three-dimensional images.

[0131] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of a computer-readable medium.

[0132] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, disk storage, quantum memory, graphene-based storage media, or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0133] Corresponding to the embodiments of the foregoing xx method, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of any one of the foregoing embodiments of the 3D Gaussian model training method or the target scene map reconstruction method.

[0134] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of protection of the present application.

Claims

1. A 3D Gaussian model training method, characterized in that: The method comprises: Based on real images of a target scene at multiple viewing angles, first normal information of the real image is acquired, and sparse reconstruction is performed on the target scene to obtain a sparse point cloud of the target scene; Initializing a 3D Gaussian model based on a sparse point cloud of the target scene, projecting multiple 3D Gaussian points in the 3D Gaussian model into a pixel space to obtain a rendered image of the target scene at the multiple viewing angles, and determining second normal information of the rendered image based on depth information of the rendered image; The real image is used as an optimization target of the rendered image, and the first normal information is used as an optimization target of the second normal information, and the 3D Gaussian model is trained to obtain a 3D Gaussian model for the target scene.

2. The method according to claim 1, characterized in that The method of taking the real image as an optimization target of the rendered image and taking the first normal information as an optimization target of the second normal information, and training the 3D Gaussian model to obtain a 3D Gaussian model for the target scene includes: Taking the real image as an optimization target of the rendered image, taking the first normal information as an optimization target of the second normal information, and performing self-supervision based on the depth information of the rendered image and the third normal information of the rendered image to train the 3D Gaussian model, so as to obtain a 3D Gaussian model for the target scene; The third normal information of the rendered image is determined based on the direction of the shortest axis of the ellipsoid corresponding to the multiple 3D Gaussian points.

3. The method according to claim 2, characterized in that The method of using the real image as an optimization target of the rendered image, using the first normal information as an optimization target of the second normal information, and performing self-supervision based on the depth information of the rendered image and the third normal information of the rendered image to train the 3D Gaussian model to obtain a 3D Gaussian model for the target scene includes: Determine a first loss value based on a difference between the rendered image and the real image, determine a second loss value based on a difference between the first normal vector and the second normal vector, and determine a third loss value based on a difference between the second normal vector and the third normal vector; Based on the first loss value, the second loss value and the third loss value, a predicted loss value of the 3D Gaussian model is determined, and model parameters of the 3D Gaussian model are updated according to the predicted loss value until a termination condition is met, thereby obtaining a 3D Gaussian model for the target scene.

4. The method according to claim 1, characterized in that: The obtaining of first normal information of the real image includes: Performing monocular depth estimation on the real image to obtain depth information of the real image; Based on the depth information of the real image, first normal information of the real image is determined.

5. The method according to claim 1, characterized in that The method further comprises: Based on the real images of the target scene at multiple viewing angles, obtaining the two-dimensional projection positions of dynamic objects in the target scene on the real images; The method of taking the real image as an optimization target of the rendered image and taking the first normal information as an optimization target of the second normal information, and training the 3D Gaussian model to obtain a 3D Gaussian model for the target scene includes: The real image is used as the optimization target of the rendered image, the first normal information is used as the optimization target of the second normal information, and the two-dimensional projection position of the dynamic object is used as the optimization target of the projection position of the 3D Gaussian point corresponding to the dynamic object in the rendered image, and the 3D Gaussian model is trained to obtain a 3D Gaussian model for the target scene.

6. The method according to claim 5, characterized in that The method further comprises: Acquire the three-dimensional spatial position of the dynamic object in the target scene; The 3D Gaussian points located in the three-dimensional space position are determined as the 3D Gaussian points corresponding to the dynamic object, wherein the 3D Gaussian points located in the three-dimensional space position include all 3D Gaussian points located in the three-dimensional space position and some 3D Gaussian points located in the three-dimensional space position.

7. A method for reconstructing a map of a target scene, characterized in that: The method comprises: Based on the internal and external parameters of the camera corresponding to the multiple perspectives of the target scene, multiple 3D Gaussian points in the 3D Gaussian model for the target scene are projected into the pixel space to obtain the rendered images of the target scene at the multiple perspectives, wherein the 3D Gaussian model for the target scene is trained by the method described in any one of claims 1 to 6 above; A scene map of the target scene is constructed according to the obtained rendered image.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Indoor scene surface reconstruction method and device, equipment and storage medium

    CN120599179A