Three-dimensional reconstruction method and system

By introducing depth information into the 3DGS algorithm to adjust the Gaussian ellipsoid density, the problem of blurred scenes in the existing technology is solved, and a higher quality three-dimensional reconstruction effect is achieved.

CN120279175APending Publication Date: 2025-07-08ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510344787.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing 3DGS algorithms often have blurred areas when reconstructing certain scenarios, resulting in poor reconstruction results.

Method used

By using the depth information of the target scene in the real image to adjust the density of the Gaussian ellipsoid in the three-dimensional model, the scene representation ability of the three-dimensional model is improved.

Benefits of technology

The three-dimensional model's representation ability of the target scene is improved, and the fuzzy areas in the reconstruction scene are avoided or reduced, achieving higher quality three-dimensional reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279175A_ABST
    Figure CN120279175A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a three-dimensional reconstruction method and system. In the three-dimensional reconstruction method, the three-dimensional reconstruction system can obtain an image set obtained by shooting a target scene from a plurality of visual angles, obtain a to-be-trained three-dimensional model comprising a plurality of Gaussian ellipsoids, and then train the three-dimensional model by using the image set. And adjusting the density of a Gaussian ellipsoid in the three-dimensional model based on the depth information of the target scene in each image of the image set in the training process so as to improve the characterization capability of the three-dimensional model for the target scene, and taking the trained three-dimensional model as a reconstruction model corresponding to the target scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the fields of 3D graphics and computer vision, and particularly to a 3D reconstruction method and system. Background Art

[0002] With the rapid development of technology, people have put forward higher requirements for high-quality 3D reconstruction and real-time rendering experience of scenes. 3D reconstruction technology can recover the 3D geometric shape and appearance information of an object or scene from 2D data.

[0003] For example, the 3Dimensional Gaussian Splatting (3DGS) algorithm can achieve low-cost and high-quality scene reconstruction by collecting a scene video or image and combining the implicit discrete Gaussian point cloud representation of the scene. A large number of studies have applied this algorithm to various 3D reconstruction scenarios, such as high-precision face reconstruction, city-level 3D reconstruction, etc. However, the existing 3DGS algorithm often has blurred areas when reconstructing certain scenes, resulting in poor reconstruction effects.

[0004] Therefore, a 3D reconstruction method capable of achieving better reconstruction effects is needed.

[0005] The content in the background art section is only the information known to the inventor personally, and does not mean that the above information has entered the public domain before the filing date of this disclosure, nor does it mean that it can become the prior art of this disclosure. Summary of the Invention

[0006] This specification provides a 3D reconstruction method and system. In this method, by using the depth information of the target scene in real images to adjust the density of Gaussian ellipsoids in the 3D model, the scene representation ability of the 3D model is improved.

[0007] In a first aspect, this specification provides a 3D reconstruction method. The 3D reconstruction method includes: obtaining an image set captured from multiple perspectives of a target scene; obtaining a 3D model to be trained, where the 3D model includes a plurality of Gaussian ellipsoids; training the 3D model using the image set, and during the training process, adjusting the density of the Gaussian ellipsoids in the 3D model based on the depth information of the target scene in each image of the image set to improve the representation ability of the 3D model for the target scene; and using the trained 3D model as the reconstruction model corresponding to the target scene.

[0008] In some embodiments, for any target image in the image set, adjusting the density of Gaussian ellipsoids in the three-dimensional model based on the depth information of the target scene in the target image includes: determining a plurality of feature points of the target scene; and for any target feature point among the plurality of feature points, determining the actual depth of the target feature point in the target image and determining the predicted depth of the target feature point in a predicted image, where the predicted image is an image formed by projecting the three-dimensional model onto the camera plane corresponding to the target image, and adjusting the density of Gaussian ellipsoids in the three-dimensional model based on the predicted depth and the actual depth.

[0009] In some embodiments, determining the predicted depth of the target feature point in the predicted image includes: determining a first ellipsoid set in the three-dimensional model, where each Gaussian ellipsoid in the first ellipsoid set is located on a target ray, and the target ray points from the position of the camera center corresponding to the target image to the position of the target feature point; and in the case where the first ellipsoid set is not empty, determining the predicted depth of the target feature point in the predicted image based on the depths of the Gaussian ellipsoids in the first ellipsoid set in the predicted image.

[0010] In some embodiments, determining the predicted depth of the target feature point in the predicted image based on the depths of the Gaussian ellipsoids in the first ellipsoid set in the predicted image includes: determining weight coefficients corresponding to the Gaussian ellipsoids in the first ellipsoid set based on the transparencies corresponding to the Gaussian ellipsoids in the first ellipsoid set; and performing weighted averaging on the depths corresponding to the Gaussian ellipsoids in the first ellipsoid set according to the corresponding weight coefficients, and taking the result of the weighted averaging as the predicted depth.

[0011] In some embodiments, the method further includes: in the case where the first ellipsoid set is empty, determining the predicted depth of the target feature point in the predicted image to be infinite.

[0012] In some embodiments, determining the actual depth of the target feature point in the target image includes: obtaining sparse reconstruction data corresponding to the target scene, where the sparse reconstruction data includes three-dimensional coordinate information of a plurality of point clouds; obtaining three-dimensional coordinate information of a target point cloud corresponding to the target feature point from the sparse reconstruction data; and determining the actual depth of the target feature point in the target image based on the camera parameters corresponding to the target image and the three-dimensional coordinate information of the target point cloud.

[0013] In some embodiments, obtaining the sparse reconstruction data corresponding to the target scene includes: performing point cloud reconstruction based on the image set to obtain the sparse reconstruction data.

[0014] In some embodiments, each pixel in the target image corresponds to depth information. Determining the actual depth of the target feature point in the target image includes: determining a target pixel corresponding to the target feature point from the target image; and using the depth information corresponding to the target pixel as the actual depth of the target feature point in the target image.

[0015] In some embodiments, adjusting the density of the Gaussian ellipsoids in the three-dimensional model based on the actual depth and the predicted depth includes: performing at least one of a first adjustment operation or a second adjustment operation based on the actual depth and the predicted depth to adjust the density of the Gaussian ellipsoids in the three-dimensional model, where the first adjustment operation includes: generating a Gaussian ellipsoid at a position corresponding to the actual depth in the three-dimensional model, and the second adjustment operation includes: determining a second ellipsoid set in the three-dimensional model and performing a decomposition process on each Gaussian ellipsoid in the second ellipsoid set, where each Gaussian ellipsoid in the second ellipsoid set is located on a target ray and occludes the target feature point, and the target ray points from the position of the camera optical center corresponding to the target image to the position of the target feature point.

[0016] In some embodiments, for any target Gaussian ellipsoid in the second ellipsoid set, the decomposition process includes: decomposing the target Gaussian ellipsoid into a plurality of Gaussian ellipsoids, and the volume of each decomposed Gaussian ellipsoid is smaller than the volume of the target Gaussian ellipsoid.

[0017] In some embodiments, performing at least one of the first adjustment operation or the second adjustment operation based on the actual depth and the predicted depth includes: performing the first adjustment operation and the second adjustment operation when the predicted depth is less than the actual depth; or performing the first adjustment operation when the predicted depth is greater than the actual depth.

[0018] In some embodiments, determining multiple feature points of the target scene includes: performing feature point detection on the target image to obtain multiple feature points of the target scene.

[0019] In some embodiments, obtaining the three-dimensional model to be trained includes: performing point cloud reconstruction based on the image set to obtain sparse reconstruction data, where the sparse reconstruction data includes three-dimensional position information of multiple point clouds; and performing parameterization processing on each point cloud in the sparse reconstruction data based on a three-dimensional Gaussian distribution to obtain the multiple Gaussian ellipsoids.

[0020] In some embodiments, the method further includes: during the training process, adjusting the density of the Gaussian ellipsoid in the three-dimensional model based on the color information of each image in the image set.

[0021] In a second aspect, this specification also provides a three-dimensional reconstruction system. The system includes: at least one storage medium storing at least one instruction set for performing three-dimensional reconstruction; and at least one processor communicatively connected to the at least one storage medium. When the three-dimensional reconstruction system runs, the at least one processor reads the at least one instruction set and executes the three-dimensional reconstruction method according to any one of the first aspect as instructed.

[0022] Other functions of the three-dimensional reconstruction method and system provided in this specification will be partially listed in the following description. The creative aspects of the three-dimensional reconstruction method and system provided in this specification can be fully explained by practicing or using the methods, devices, and combinations described in the detailed examples below. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the technical solutions in the embodiments of this specification, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0024] Figure 1 FIG. shows a schematic diagram of an application scenario of three-dimensional reconstruction provided according to an embodiment of this specification;

[0025] Figure 2 FIG. shows a basic flowchart of three-dimensional reconstruction using the 3DGS algorithm provided according to an embodiment of this specification;

[0026] Figure 3 FIG. shows a hardware structure diagram of a computing system provided according to an embodiment of this specification;

[0027] Figure 4 FIG. shows a flowchart of a three-dimensional reconstruction method provided according to an embodiment of this specification; and

[0028] Figure 5 FIG. shows a partial process schematic diagram of determining the predicted depth provided according to an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The following description provides specific application scenarios and requirements of this specification, aiming to enable those skilled in the art to manufacture and use the content in this specification. For those skilled in the art, various partial modifications to the disclosed embodiments are obvious, and the general principles defined here can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the disclosed embodiments, but has the broadest scope consistent with the claims.

[0030] The terms used herein are for the purpose of describing specific example embodiments only and are not restrictive. For example, unless the context clearly indicates otherwise, as used herein, the singular forms "a", "an", and "the" may also include the plural forms. When used in this specification, the terms "comprising", "including", and / or "containing" mean that the associated integers, steps, operations, elements, and / or components exist, but do not exclude the existence of one or more other features, integers, steps, operations, elements, components, and / or groups, or the addition of other features, integers, steps, operations, elements, components, and / or groups in the system / method.

[0031] Considering the following description, these features of this specification and other features, as well as the operations and functions of the related elements of the structure, and the economy of the combination and manufacture of the components can be significantly improved. Referring to the accompanying drawings, all of which form a part of this specification. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.

[0032] The flowcharts used in this specification illustrate the operations implemented by the system according to some embodiments in this specification. It should be clearly understood that the operations in the flowchart may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.

[0033] The application scenarios of this specification are introduced below.

[0034] The embodiments of this specification can be applied to scenarios that require three-dimensional reconstruction of the target scenario. For example, the embodiments of this specification can be used in multiple fields such as computer vision, graphics, robotics, medicine, archaeology, architecture, etc. In the medical field, three-dimensional reconstruction can convert the image data of computed tomography (CT) into a three-dimensional model to help doctors plan surgeries more accurately. In the field of architecture, three-dimensional reconstruction can convert the data of ancient buildings into a three-dimensional model to provide accurate geometric information for building renovation.

[0035] It should be noted that the above application scenarios for three-dimensional reconstruction of the target scenario are only some examples among multiple usage scenarios. The three-dimensional reconstruction method provided in this specification can be applied not only to the scenarios listed above, but also to other scenarios that require three-dimensional reconstruction. Those skilled in the art should understand that when the three-dimensional reconstruction method provided in this specification is applied to other usage scenarios, its implementation manner and technical effects are similar.

[0036] Figure 1 FIG. shows a schematic diagram of an application scenario 100 for three-dimensional reconstruction provided according to an embodiment of this specification. As Figure 1 shown, the application scenario 100 may include a three-dimensional reconstruction system 110.

[0037] The three-dimensional reconstruction system 110 can realize the three-dimensional reconstruction of the target scenario. For example, the three-dimensional reconstruction system 110 can obtain an image set captured from multiple perspectives of the target scenario, train a three-dimensional model to be trained based on the images in the image set, and use the trained three-dimensional model as the reconstruction model corresponding to the target scenario.

[0038] As an example, Figure 1 the target scenario in is a chair, and multiple images are captured from four perspectives of the target scenario by an image acquisition device. The three-dimensional reconstruction system 110 can perform three-dimensional reconstruction on the target scenario based on the images captured from the four perspectives to obtain a reconstruction model.

[0039] When performing three-dimensional reconstruction on the target scenario, the three-dimensional reconstruction system 110 can use a voxel rendering algorithm based on Neural Radiance Field (NeRF), or can also use the 3DGS algorithm. Hereinafter, the 3DGS algorithm will be mainly introduced as an example.

[0040] The 3DGS algorithm mainly realizes the efficient representation of the target scenario by collecting multi-perspective images of the target scenario, combining the training of the cutting-edge differentiable neural radiance field, and using multiple Gaussian ellipsoids. Among them, each Gaussian ellipsoid contains 3D attributes observed from different perspectives in the scenario. The 3DGS algorithm then trains and optimizes the 3D Gaussian ellipsoids through a network based on the differentiable rendering process supervised by multiple perspective images, so as to achieve high-quality reconstruction and representation of the target scenario.

[0041] Figure 2 FIG. shows a basic flowchart of three-dimensional reconstruction using the 3DGS algorithm provided according to an embodiment of this specification. As Figure 2As shown, during the training process, the 3D reconstruction system 110 can determine a target image from a set of images, and then determine the camera parameters corresponding to the target image. The 3D reconstruction system 110 can project the current 3D model based on the camera parameters to obtain a predicted image. The above projection process is actually a process of projecting each Gaussian ellipsoid in the target scene onto a two-dimensional plane along a specified viewing angle. Then, the 3D reconstruction system 110 renders the Gaussian ellipsoids into a predicted image through differentiable Gaussian rasterization technology, and then calculates the error between the predicted image and the target image. The above process can be called the forward propagation process (shown by the solid line). The 3D reconstruction system 110 can update the parameters of the Gaussian ellipsoids in a gradient backpropagation manner according to the error, and perform density control on the Gaussian ellipsoids. Among them, performing adaptive density control on the Gaussian ellipsoids can include the generation of new Gaussian ellipsoids. The above process can be called the backward propagation process (shown by the dashed line).

[0042] In some embodiments, when performing density control on Gaussian ellipsoids, only color information is used as the basis for determining whether to generate new Gaussian ellipsoids. For example, color information such as position gradient information, color gradient information, or spectrogram distribution on the target image is used to determine whether new Gaussian ellipsoids need to be generated. However, in the actual application of 3DGS, when only using color information to perform density control on Gaussian ellipsoids, the model often has a poor reconstruction effect. For example, the reconstructed scene often has blurred areas. After a large amount of research and in-depth thinking, the inventor found that the reason for the blurred areas is that the front-back position relationship of some Gaussian ellipsoids is incorrect. This is because when performing density control on Gaussian ellipsoids, only the color information of the image is considered, and the color information often cannot reflect the depth information of the scene.

[0043] This specification also provides a 3D reconstruction method and system. In the 3D reconstruction method, the density of Gaussian ellipsoids in the 3D model is adjusted by the depth information in the target image of the real shot of the target scene, so as to improve the scene representation ability of the 3D model. Specifically, by introducing depth information, the relationship between the true depth of the feature points in the target scene in the target image and the predicted depth in the predicted image obtained by projecting the Gaussian ellipsoids in the 3D model of the target feature points is examined, so as to perform operations of decomposing and / or generating the Gaussian ellipsoids in the 3D model, so as to improve the scene performance ability of the 3D model, so that the scene reconstructed by the 3D model avoids having blurred areas, or reduces the possibility of having blurred areas in the scene, and the scene reconstructed by the 3D model tends to the target scene.

[0044] Figure 3 Shows a hardware structure diagram of a computing system 300 provided according to an embodiment of this specification.

[0045] The computing system 300 may serve as the Figure 1 3D reconstruction system 110 in

[0046] and execute the 3D reconstruction method described in this specification. Figure 3 As shown, the computing system 300 may include at least one storage medium 330 and at least one processor 320. In some embodiments, the computing system 300 may further include a communication port 350 and an internal communication bus 310. The computing system 300 may also include I / O components 360.

[0047] The internal communication bus 310 may connect different system components. For example, the internal communication bus 310 may connect the storage medium 330, the processor 320, the communication port 350, and the I / O components 360, etc.

[0048] The I / O components 360 support input / output between the computing system 300 and other components.

[0049] The communication port 350 is used for data communication between the computing system 300 and the outside world. For example, the communication port 350 may be used for data communication between the computing system 300 and the network 140. The communication port 350 may be a wired communication port or a wireless communication port.

[0050] The storage medium 330 may include a data storage device. The data storage device may be a non-transitory storage medium or a transitory storage medium. For example, the data storage device may include one or more of a magnetic disk 332, a read-only storage medium (ROM) 334, or a random access storage medium (RAM) 335. The storage medium 330 further includes at least one instruction set stored in the data storage device. The instruction set may include computer program code, and the computer program code may include programs, routines, objects, components, data structures, processes, modules, and so on.

[0051] At least one processor 320 can be communicatively connected to at least one storage medium 330. When the computing system 300 is running, at least one processor 320 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the three-dimensional reconstruction method provided in this specification. The processor 320 can execute the steps included in the three-dimensional reconstruction method. The processor 320 can be in the form of one or more processors. In some embodiments, the processor 320 can include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field-programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of executing one or more functions, etc., or any combination thereof.

[0052] For illustrative purposes only, only one processor 320 is shown in the computing system 300 in the drawings. However, it should be noted that the computing system 300 in this specification can also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification can be executed by one processor or jointly by multiple processors. For example, if it is described in this specification that the processor 320 of the computing system 300 executes step A and step B, it should be understood that step A and step B can also be executed jointly or separately by two different processors 320 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).

[0053] Figure 4 A flowchart of a three-dimensional reconstruction method P400 provided according to an embodiment of this specification is shown. The three-dimensional reconstruction system 110 can execute the three-dimensional reconstruction method P400. Specifically, the processor 220 in the three-dimensional reconstruction method P400 can read the instruction set stored in its local storage medium and then, according to the provisions of the instruction set, execute the three-dimensional reconstruction method P400 of this specification. As Figure 4 shown, the three-dimensional reconstruction method P400 can include steps S410 - S440.

[0054] S410: Obtain a set of images obtained by photographing a target scene from multiple perspectives.

[0055] The target scene can be any object or scene that a user expects to obtain or reconstruct, such as a building, a car, a person, a sea view, etc. The image set is a group of images including multiple perspective images obtained by photographing the target scene from different angles. The images in the image set contain the real content of the target scene from different perspectives and can provide multi-faceted information for the reconstruction of the target scene.

[0056] Different target scenes can use different image acquisition devices and select different shooting perspectives, which are not limited in this specification.

[0057] S420: Obtain a three-dimensional model to be trained, where the three-dimensional model includes a plurality of Gaussian ellipsoids.

[0058] A Gaussian ellipsoid is a geometric shape based on a three-dimensional Gaussian distribution and is used to represent data points or objects in three-dimensional space. Each Gaussian ellipsoid can be defined by parameters such as position (mean), covariance matrix, volume density (transparency), and spherical harmonic coefficients. A plurality of Gaussian ellipsoids combined together can construct a three-dimensional model.

[0059] In some embodiments, obtaining the three-dimensional model to be trained may include: performing point cloud reconstruction based on the image set to obtain sparse reconstruction data, where the sparse reconstruction data includes three-dimensional coordinate information of a plurality of point clouds; and performing parameterization processing on each point cloud in the sparse reconstruction data based on a three-dimensional Gaussian distribution to obtain the plurality of Gaussian ellipsoids.

[0060] Point cloud reconstruction is a technique for extracting three-dimensional information from two-dimensional images. Point cloud reconstruction techniques may include, but are not limited to, Structure from Motion (SfM), Multi-View Stereo (MVS), and so on. As described above, the image set includes information of the target scene from different perspectives. Performing point cloud reconstruction based on the image set can obtain or recover the three-dimensional information of the target scene from the two-dimensional images, thereby obtaining sparse reconstruction data. The sparse reconstruction data may include three-dimensional coordinate information of a plurality of point clouds. The three-dimensional coordinate information of the point clouds characterizes the geometric shape and spatial distribution of the target scene, etc. The plurality of point cloud data included in the sparse reconstruction data only cover some key scenes in the target scene and are not a complete scene.

[0061] The three-dimensional reconstruction system 110 can initialize a plurality of point clouds, that is, perform parameterization processing on these point cloud data based on a three-dimensional Gaussian distribution, so as to convert the originally scattered point cloud information into a continuous and high-dimensional mathematical representation, enabling the three-dimensional model to be trained in a more standardized form.

[0062] As an example, the three-dimensional model to be trained may be Figure 2Multiple Gaussian ellipsoids shown.

[0063] S430: Train the three-dimensional model using the image set, and during the training process, adjust the density of the Gaussian ellipsoids in the three-dimensional model based on the depth information of the target scene in each image of the image set to enhance the representation ability of the three-dimensional model for the target scene.

[0064] The depth information of an image refers to the distance information of each pixel point in the image relative to the image acquisition device (such as a camera). The depth information describes the position of the target scene in the depth direction (such as the optical axis direction of the camera). The density of the Gaussian ellipsoids can reflect their distribution and coverage in three-dimensional space, thereby affecting the capture and representation of scene details by the three-dimensional model. When Gaussian ellipsoids appear at incorrect positions in the three-dimensional model or are generated, they may cover or interfere with the details that actually exist in the scene, and may also produce blurred or inaccurate structures, especially in edge or complex regions where blurred or inaccurate structures are likely to occur.

[0065] The image set contains the real content of the target scene from different perspectives. Therefore, during the training process, by using the image set to introduce real depth information to adjust the density of the Gaussian ellipsoids in the three-dimensional model, the target scene can be assisted in reconstruction, so as to achieve a more refined capture and more efficient representation of the scene, and further enhance the representation ability of the three-dimensional model for the target scene.

[0066] The three-dimensional reconstruction system can train the three-dimensional model using each image in the image set, or can also use partial images corresponding to representative perspectives in the image set to train the three-dimensional model. For example, the image set includes 3 images. The three-dimensional reconstruction system can adjust the density of the Gaussian ellipsoids in the three-dimensional model based on the depth information of the target scene in the first image, then adjust the density of the Gaussian ellipsoids in the three-dimensional model based on the depth information in the second image, and then adjust the density of the Gaussian ellipsoids in the three-dimensional model based on the depth information of the target scene in the third image. For another example, the viewing angle difference between the first image and the second image is very small, and the three-dimensional reconstruction system can adjust the density of the Gaussian ellipsoids in the three-dimensional model only using the depth information in the first and third images.

[0067] The following introduces the method of density adjustment based on the depth information of the target image in the image set. The target image can be any image in the image set, or can also be an image with relatively high information content in the image set, such as an image corresponding to a representative perspective.

[0068] In some embodiments, for any target image in the image set, adjusting the density of the Gaussian ellipsoids in the three-dimensional model based on the depth information of the target scene in the target image may include the following steps.

[0069] (1) Determine multiple feature points of the target scene.

[0070] The feature points of the target scene can be points that can represent the saliency and uniqueness of the scene in the target scene, such as corner points, edge points, points with strong texture feature changes, etc. Feature points can usually be stably detected from different perspectives and can also remain stable for changes such as rotation, scaling, and illumination.

[0071] In some embodiments, determining multiple feature points of the target scene may include: performing feature point detection on the target image to obtain multiple feature points of the target scene.

[0072] Among them, the feature points of the target scene can be two-dimensional feature points (feature data). For example, the feature points of the target scene are the feature points of the target image. The 3D reconstruction system can determine the feature points of the target image and use them as the feature points of the target scene. The feature points of the target scene can also be three-dimensional feature points. For example, the feature points of the target scene are the geometric representations of the two-dimensional feature points of the target image in three-dimensional space, and the feature points of the target image are the two-dimensional projections of the three-dimensional feature points of the target scene.

[0073] As an example, the feature points in each target image in the image set can be determined through the SfM algorithm, and then the feature points in each target image in the image set are matched through the feature matching algorithm and the positions of the feature points in three-dimensional space are calculated through triangulation, so as to obtain points in three-dimensional space, that is, the feature points of the target scene. In some embodiments, the multiple point clouds included in the sparse reconstruction data are the multiple feature points of the target scene, and the three-dimensional coordinate information of the multiple point clouds is the three-dimensional coordinate information of the feature points of the target scene. That is, the multiple feature points of the target scene constitute the sparse reconstruction data.

[0074] Among them, feature point detection on the target image can be performed through algorithms such as Scale Invariant Feature Transform (SIFT) algorithm, (Oriented FAST and Rotated BRIEF, ORB) algorithm, etc.

[0075] In some embodiments, multiple three-dimensional feature points of the target scene can also be obtained by laser scanning the target scene.

[0076] (2) For any feature point among the multiple feature points, perform the following steps.

[0077] (1) Determine the actual depth of the target feature point in the target image, and determine the predicted depth of the target feature point in the predicted image, where the predicted image is an image formed by projecting the three-dimensional model onto the camera plane corresponding to the target image.

[0078] The actual depth of the target feature point in the target image can reflect the actual position information after the target scene is projected onto the target image. The predicted depth of the target feature point in the target image can reflect the position information after the scene constructed by the three-dimensional model is projected onto the camera plane.

[0079] The following separately introduces how to determine the actual depth and the predicted depth of the target feature point in the target image.

[0080] Since the image set contains the real content of the target scene from different perspectives, the target images in the image set may include the actual depth information of the target feature point. Therefore, in some embodiments, each pixel in the target image corresponds to depth information. Determining the actual depth of the target feature point in the target image may include: determining a target pixel corresponding to the target feature point from the target image; and using the depth information corresponding to the target pixel as the actual depth of the target feature point in the target image.

[0081] A pixel is the smallest unit of an image, representing a point in the image. The target image may include multiple pixels (pixel points). Each pixel point may include position information (such as its actual position in the target image), color information, depth information, etc. The two-dimensional feature point can be considered as a point with significant features extracted and selected from the pixels. The feature point can be considered as a subset of the pixel points. For example, in a 1024×1024 image, there may be more than one million pixels, but after feature extraction, there may be only a few hundred feature points. The three-dimensional reconstruction system can determine the target pixel corresponding to the target feature point and use the depth information corresponding to the target pixel as the actual depth of the target feature point.

[0082] In some embodiments, the RGBD camera can obtain the depth information of the scene while capturing the target image by combining a color image sensor and a depth sensor. Therefore, using an RGBD camera can obtain the depth information of each pixel point in the image.

[0083] Since the image set contains the real content of the target scene from different perspectives, and the sparse reconstruction data is obtained by performing point cloud reconstruction based on the image set, the sparse reconstruction data also includes the information of the actual depth. Therefore, in some embodiments, determining the actual depth of the target feature point in the target image may include the following steps:

[0084] (a) Obtain the sparse reconstruction data corresponding to the target scene, where the sparse reconstruction data includes the three-dimensional coordinate information of multiple point clouds.

[0085] The method for obtaining the sparse reconstruction data corresponding to the target scene can refer to the above text and will not be elaborated here.

[0086] (b) Obtain the three-dimensional coordinate information of the target point cloud corresponding to the target feature point from the sparse reconstruction data.

[0087] The target feature point can find its corresponding target point cloud among the multiple point clouds included in the sparse reconstruction data. In some embodiments, when the target feature point is two-dimensional feature data from the target image, by generating a matrix for converting two-dimensional coordinates to three-dimensional coordinates during the process of generating the sparse reconstruction data, the point cloud corresponding to the target feature point in the three-dimensional space can be found. In some embodiments, when the target feature point is three-dimensional feature data, as described above, multiple feature points of the target scene can form the sparse reconstruction data. That is, each point cloud in the sparse reconstruction data is a feature point, and the three-dimensional coordinate information of the feature point is also the three-dimensional coordinate information of its corresponding point cloud. Through the three-dimensional coordinate information of the target feature point and the three-dimensional coordinate information of the point cloud, the target point cloud corresponding to the target feature point and the three-dimensional coordinate information of the target point cloud can be found.

[0088] (c) Based on the camera parameters corresponding to the target image and the three-dimensional coordinate information of the target point cloud, determine the actual depth of the target feature point in the target image.

[0089] The camera parameters corresponding to the target image may include the camera internal parameters and camera external parameters for taking the target image. The camera external parameters describe the position and orientation of the camera in the world coordinate system. The camera external parameters may include a rotation matrix, a translation vector, etc. The camera external parameters can be used to transform points in the world coordinate system to the camera coordinate system. The camera internal parameters describe the internal optical characteristics of the camera. The camera internal parameters may include the focal length of the camera, the position of the principal point, distortion coefficients, etc. The camera internal parameters can be used to transform points in the camera coordinate system to the image coordinate system.

[0090] The camera internal parameters and external parameters are usually combined into a calibration matrix for transforming points in the world coordinate system to the image coordinate system. That is, through the camera parameters, the target point cloud in the world coordinate system can be converted to the image coordinate system, and the three-dimensional coordinate information of the target point cloud can be converted to the coordinate information in the image coordinate system. The three-dimensional reconstruction system can determine the actual depth of the target feature point at the corresponding pixel point in the target image through the transformed coordinate information. Further, the three-dimensional reconstruction system can use the following formula (1) to determine the coordinates of the target point cloud in the image coordinate system:

[0091] p image = K[R|t]Pworld

[0092] In formula (1), P world represents the coordinates of the target point cloud and can be a 4×1 vector; K represents the intrinsic matrix of the camera and can be a 3×4 matrix, [R|t] represents the extrinsic matrix of the camera and can be a 4×4 matrix; represents the rotation matrix, and t represents the translation vector. The operation result p image represents the projection coordinates of the target point cloud in the image coordinate system and can be a 3×1 vector. For example, p image can be expressed as p image =[u, v, w] T , where u and v are the coordinates of the pixels corresponding to the target point cloud, and w is the depth information of the pixel, that is, the actual depth of the target point cloud. In some embodiments, [u, v, w] T represents the position of the projected point on the image plane. In some embodiments, [u / w, v / w] T represents the position of the actual pixel of the projected point on the image plane.

[0093] Therefore, the extrinsic matrix [R|t] of the camera is used to transform the target point cloud from the world coordinate system to the camera coordinate system, and then the intrinsic matrix K of the camera is used to project the points in the camera coordinate system onto the image coordinate system to determine the pixels corresponding to the target point cloud on the target image and the depth corresponding to the pixels, and the depth corresponding to the pixel point is determined as the actual depth corresponding to the target point cloud.

[0094] In some embodiments, the depth information of the target feature points can also be obtained by a lidar. The lidar can emit laser pulses to scan the target scene, record the distance of each measurement point to form a 3D point cloud, and then project all the 3D point clouds onto the scanning plane of the lidar to generate a depth image. Each pixel value in the depth image represents the depth information of the pixel point.

[0095] The three-dimensional reconstruction system can also determine the predicted depth of the target feature points in the predicted image. The predicted image is an image formed by projecting a three-dimensional model onto the camera plane corresponding to the target image.

[0096] The camera plane may refer to the plane on which an image is formed. For example, the camera plane may refer to the plane where the camera sensor is located. The 3D reconstruction system may determine the camera plane according to the parameters of the camera, combine multiple Gaussian ellipsoids in the 3D model with the camera parameters, and project the Gaussian ellipsoids in the 3D space onto the 2D plane through perspective projection. The projection result of each Gaussian ellipsoid will form a point or a region on the image plane, presenting the object form in the target image. The projection result maps the geometric information of the 3D Gaussian ellipsoid onto the corresponding camera plane of the target image, and finally forms a predicted image. The predicted image presents a simulation of the target scene, reflecting the result of mapping multiple Gaussian ellipsoids onto the camera plane.

[0097] Figure 5 Fig. shows a partial process schematic diagram for determining the predicted depth according to an embodiment of the present specification. The following will describe in conjunction with Figure 5 how to determine the predicted depth of the target feature point in the predicted image.

[0098] (a) Determine a first set of ellipsoids in the 3D model, where each Gaussian ellipsoid in the first set of ellipsoids is located on the target ray, and the target ray points from the position of the camera optical center corresponding to the target image to the position of the target feature point.

[0099] The position of the camera optical center corresponding to the target image may be the position where the camera optical center for shooting the target image is located. The 3D reconstruction system may determine the target ray with the position of the camera optical center as the starting point and the direction from the optical center position to the target feature point in the 3D space. In the imaging process of the camera, this ray is a set representation of the points in the 3D space projected onto the 2D plane. That is, all the points in the 3D space on the target ray will be projected onto the same point on the imaging plane (camera plane).

[0100] Each Gaussian ellipsoid in the first set of ellipsoids is located on the target ray. That is, each Gaussian ellipsoid in the first set of ellipsoids will be projected onto the same point on the imaging plane, and this point is the pixel point corresponding to the target feature point in the predicted image. All the Gaussian ellipsoids in the first set of ellipsoids act together to render the pixel point corresponding to the target feature point in the predicted image. Therefore, each Gaussian ellipsoid in the first set of ellipsoids affects the depth information of this pixel point and the predicted depth information of the target feature point.

[0101] For example, as Figure 5 shown, the target ray points from the camera optical center to the position of the target feature point, and the Gaussian ellipsoid A, the Gaussian ellipsoid B, and the Gaussian ellipsoid C are located on the target ray. It should be noted that Figure 5 is only for the convenience of explanation, and the actual 3D model may be more complex than the Figure 5 3D model shown in.

[0102] (b) When the first ellipsoid set is not empty, based on the depths of the Gaussian ellipsoids in the first ellipsoid set in the predicted image, determine the predicted depth of the target feature point in the predicted image.

[0103] Since each Gaussian ellipsoid in the first ellipsoid set affects the depth shown by the target feature point at the corresponding pixel in the predicted image, the predicted depth of the target feature point in the predicted image can be determined based on the depths of all the Gaussian ellipsoids in the first ellipsoid set in the predicted image.

[0104] In some embodiments, the determining the predicted depth of the target feature point in the predicted image based on the depths of the Gaussian ellipsoids in the first ellipsoid set in the predicted image may include: determining the weight coefficients corresponding to the Gaussian ellipsoids in the first ellipsoid set based on the transparencies corresponding to the Gaussian ellipsoids in the first ellipsoid set; and performing weighted averaging on the depths corresponding to the Gaussian ellipsoids in the first ellipsoid set according to the corresponding weight coefficients, and taking the result of the weighted averaging as the predicted depth.

[0105] The transparency of a Gaussian ellipsoid can be used to describe the degree of light occlusion of the Gaussian ellipsoid during the rendering process, and the transparency can also reflect the contribution degree of the Gaussian ellipsoid to the depth information. The value range of the transparency is usually [0, 1]. When the transparency is 0, it means the Gaussian ellipsoid is completely transparent and light can completely penetrate the Gaussian ellipsoid, without contributing to the depth information in the predicted image. When the transparency is 1, it means the Gaussian ellipsoid is completely opaque, light is completely blocked by it, and the Gaussian ellipsoids behind it are completely invisible, with the greatest contribution to the depth information in the predicted image. That is, the greater the transparency, the greater the contribution degree of the Gaussian ellipsoid to the depth information in the predicted image, and the lower the transparency, the smaller the contribution degree of the Gaussian ellipsoid to the depth information in the predicted image.

[0106] As mentioned above, there is a corresponding relationship between the target feature point and a certain pixel in the image. When multiple Gaussian ellipsoids in the first ellipsoid set are projected onto the same two-dimensional pixel point, the depth information of this pixel point can be considered as the weighted synthesis of the transparencies of all the Gaussian ellipsoids in the first ellipsoid set. Further, the 3D reconstruction system can use the following formula (2) to determine the final depth value Z of a pixel point after multiple Gaussian ellipsoids are projected onto it:

[0107]

[0108] In formula (2), α i represents the transparency of the i-th Gaussian ellipsoid, z iCharacterize the depth (depth value) of the i-th Gaussian ellipsoid. The depth of the Gaussian ellipsoid can be the coordinate value of the center point of the Gaussian ellipsoid in the optical axis direction in the camera coordinate system. After transforming the center position of the Gaussian ellipsoid to the camera coordinate system through the external parameters of the camera, the coordinate value of the center point of the Gaussian ellipsoid in the optical axis direction in the camera coordinate system can be determined, and thus the depth of the Gaussian ellipsoid is determined. Further, the 3D reconstruction system can determine the predicted depth of the target feature point in the predicted image, that is, the predicted depth of the corresponding pixel of the target feature point in the predicted image.

[0109] For example, after the 3D reconstruction system performs weighted averaging on the Gaussian ellipsoid A, Gaussian ellipsoid B, and Gaussian ellipsoid C in Figure 5 using the above formula (2), the predicted depth is obtained. Figure 5 The position L2 corresponding to the predicted depth is shown in

[0110] (2) Based on the predicted depth and the actual depth, adjust the density of the Gaussian ellipsoids in the 3D model.

[0111] The difference between the predicted depth and the actual depth can reflect the fitting degree of the current 3D model to the target scene, which is of guiding significance for the optimization of the 3D model. Adjust the density of the Gaussian ellipsoids in the 3D model based on the difference between the two, so that the 3D model can more accurately represent the target scene, improve the ability of the 3D model to capture details of the target scene, avoid the reconstructed scene of the 3D model having fuzzy regions or reduce the possibility of the existence of fuzzy regions, and make the scene represented by the 3D model visually closer to the real target scene.

[0112] In some embodiments, the adjusting the density of the Gaussian ellipsoids in the 3D model based on the actual depth and the predicted depth may include: performing at least one of a first adjustment operation or a second adjustment operation based on the actual depth and the predicted depth to adjust the density of the Gaussian ellipsoids in the 3D model. Specifically:

[0113] (a) When the predicted depth is less than the actual depth, perform the first adjustment operation and the second adjustment operation.

[0114] Among them, the first adjustment operation may include: generating a Gaussian ellipsoid at the position corresponding to the actual depth in the 3D model.

[0115] The second adjustment operation may include: determining a second ellipsoid set in the 3D model and performing decomposition processing on each Gaussian ellipsoid in the second ellipsoid set, where each Gaussian ellipsoid in the second ellipsoid set is located on the target ray and forms an occlusion to the target feature point, and the target ray points from the position of the camera optical center corresponding to the target image to the position of the target feature point.

[0116] When the predicted depth is less than the actual depth, it means that the depth information in the predicted image is inaccurate. From the above analysis, it can be seen that if some Gaussian ellipsoids have high transparency and shallow depth (closer to the camera), they will occupy a dominant position in the depth synthesis, masking the influence of the Gaussian ellipsoids with a farther depth, which may make the final predicted depth value smaller. Therefore, when the predicted depth is less than the actual depth, there is a high probability that a Gaussian ellipsoid that should not exist is generated or exists before the position corresponding to the actual depth. In other words, the target ray pointing from the camera optical center position to the location of the target feature point should point to the actual depth, but before pointing to the actual depth, it may pass through other Gaussian ellipsoids. These other Gaussian ellipsoids can be considered as Gaussian ellipsoids that are incorrectly generated or should not be generated at this position during the training of the three-dimensional model. In order to correct the improper representation of the target scene by the three-dimensional model and retain the detailed information of the target scene, the three-dimensional reconstruction system can decompose these Gaussian ellipsoids.

[0117] In some embodiments, for any target Gaussian ellipsoid in the second ellipsoid set, the decomposition process may include: decomposing the target Gaussian ellipsoid into multiple Gaussian ellipsoids, and the volume of each decomposed Gaussian ellipsoid is smaller than the volume of the target Gaussian ellipsoid. For example, the target Gaussian ellipsoid is decomposed into two Gaussian ellipsoids, four or more smaller Gaussian ellipsoids.

[0118] For example, Figure 5 The position L1 corresponding to the actual depth is shown, and there are Gaussian ellipsoid A and Gaussian ellipsoid B before the position D, and there is Gaussian ellipsoid C after the position L1. Gaussian ellipsoid A and Gaussian ellipsoid B can be considered as incorrectly generated ellipsoids, and the second ellipsoid set includes Gaussian ellipsoid A and Gaussian ellipsoid B. The three-dimensional reconstruction system can decompose Gaussian ellipsoid A into two Gaussian ellipsoids, A1 and A2; and decompose Gaussian ellipsoid B into two Gaussian ellipsoids, B1 and B2.

[0119] When the predicted depth is less than the actual depth, the depth information corresponding to the target feature point is missing or offset because the wrong Gaussian ellipsoid is decomposed. Therefore, in addition to performing the second operation, the 3D reconstruction system can also perform the first adjustment operation to generate a Gaussian ellipsoid at the position corresponding to the actual depth in the 3D model. Since the position of the newly generated Gaussian ellipsoid is the same as the position corresponding to the actual depth, the Gaussian ellipsoid can accurately show the actual depth or position information of the target scene, so that the scene reconstructed by the Gaussian ellipsoid does not appear or reduces the possibility of blurred areas in the scene.

[0120] For example, Figure 5 The position L1 corresponding to the actual depth is shown, and a new Gaussian ellipsoid D is generated at the position L1.

[0121] (b) In the case where the predicted depth is greater than the actual depth, perform the first adjustment operation.

[0122] When the predicted depth is greater than the actual depth, it indicates that the Gaussian ellipsoid corresponding to the actual depth is not correctly generated. Therefore, the 3D reconstruction system can perform the first adjustment operation to generate a Gaussian ellipsoid at the position corresponding to the actual depth.

[0123] It should be noted that the above density adjustment of the Gaussian ellipsoid can be performed during the iterative training process of the 3D model. Among them, during each iterative training, after the 3D reconstruction system decomposes the Gaussian ellipsoid, the small Gaussian ellipsoids obtained by decomposition may still be located on the target ray and block the target feature points. In this case, the small Gaussian ellipsoids still need to be decomposed. During the next iterative training process, the system can repeat the above 3D reconstruction method to decompose the small Gaussian ellipsoids again. The specific processing process will not be elaborated here.

[0124] In some embodiments, when the difference between the predicted depth and the actual depth is less than a preset threshold, the 3D reconstruction system can consider that the depth exhibited by the 3D model is the same as or approximately the same as the actual depth. During the subsequent iterative training process, even if there are still Gaussian ellipsoids on the target ray that block the target feature points, the 3D reconstruction system may not perform the above decomposition process.

[0125] In some embodiments, the 3D reconstruction system can simultaneously use depth information and color information to adjust the density of the Gaussian ellipsoids in the 3D model. Therefore, the 3D reconstruction method may further include: adjusting the density of the Gaussian ellipsoids in the 3D model based on the color information of each image in the image set during the training process. As mentioned above, the color information of the image may include position gradient information, color gradient information, or spectrogram distribution, etc. Adjusting the density based on the color information and depth information of each image in the image set can be, for example: The image set includes 3 images. The 3D reconstruction system can adjust the density of the Gaussian ellipsoids in the 3D model based on the depth information and color information of the target scene in the first image, then adjust the density of the Gaussian ellipsoids in the 3D model based on the depth information and color information of the second image, and then adjust the density of the Gaussian ellipsoids in the 3D model based on the depth information and color information of the target scene in the third image.

[0126] Continue to refer to Figure 3 , the 3D reconstruction method P400 further includes:

[0127] S440: Use the trained 3D model as the reconstruction model corresponding to the target scene.

[0128] By adjusting and optimizing the Gaussian ellipsoid density of the three-dimensional model using depth information, a more refined three-dimensional reconstruction model for representing the target scene can ultimately be obtained. This three-dimensional model can avoid the occurrence of blurred regions or reduce the possibility of blurred regions in the three-dimensional model, and thus can be used as the reconstruction model corresponding to the target scene.

[0129] In summary, the three-dimensional reconstruction method provided in this specification adjusts the Gaussian ellipsoid density in the three-dimensional model by introducing real depth information using an image set to assist in reconstructing the target scene, thereby achieving a more refined capture and more efficient representation of the scene by the three-dimensional model, and further enhancing the representation ability of the three-dimensional model for the target scene. Specifically, by introducing depth information, the relationship between the true depth of the feature points in the target scene in the target image and the predicted depth of the target feature points in the predicted image obtained by projecting the Gaussian ellipsoid in the three-dimensional model is compared, so as to perform operations of decomposing and / or generating the Gaussian ellipsoid in the three-dimensional model to enhance the scene performance ability of the three-dimensional model, thereby avoiding or reducing the possibility of blurred regions in the reconstructed scene, and making the scene reconstructed by the three-dimensional model tend to the target scene.

[0130] On the other hand, this specification provides a non-transitory storage medium storing at least one set of executable instructions for performing 3D reconstruction. When the executable instructions are executed by a processor, the executable instructions direct the processor to perform the steps of the 3D reconstruction method P400 described in this specification. In some possible implementation manners, various aspects of this specification may also be implemented in the form of a program product, which includes program code. When the program product runs on a 3D reconstruction system, the program code is used to cause the 3D reconstruction system to perform the steps of the 3D reconstruction method P400 described in this specification. The program product for implementing the above method may adopt a portable compact disc read-only memory (CD-ROM) including program code and may run on a 3D reconstruction system. However, the program product of this specification is not limited thereto. In this specification, the readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system. The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium may include a data signal propagated in a baseband or as a part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium other than the readable storage medium, and the readable medium may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above. The program code for performing the operations of this specification may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the 3D reconstruction system, partially on the 3D reconstruction system, executed as an independent software package, partially on the 3D reconstruction system and partially on a remote computing device, or entirely on the remote computing device.

[0131] The above description has been made of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require a particular order or a sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0132] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and is not necessarily limiting. Although not explicitly stated herein, those skilled in the art will understand that this specification is intended to encompass various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.

[0133] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with that embodiment may be included in at least one embodiment of this specification. Thus, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. Additionally, the specific features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.

[0134] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature, and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, drawing, or its description. However, this does not mean that the combination of these features is necessary, and those skilled in the art are fully likely to mark out some of the devices as separate embodiments for understanding when reading this specification. That is to say, the embodiments in this specification can also be understood as the integration of multiple sub - embodiments. And the content of each sub - embodiment is also valid when it has fewer features than all the features of a single foregoing disclosed embodiment.

[0135] Each patent, patent application, publication of patent application, and other materials cited herein, such as articles, books, specifications, publications, documents, items, etc., except those that are inconsistent with or conflict with this document, or those that have a limiting effect on the broadest scope of the claims, may be incorporated herein by reference and used for all purposes now or hereafter associated with this document. In addition, in the event of any inconsistency or conflict between the description, definition, and / or use of relevant terms in any material and the description, definition, and / or use of relevant terms in this document, the terms in this document shall prevail.

[0136] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.

Claims

1. A three-dimensional reconstruction method, comprising: Obtaining a set of images captured from multiple perspectives of a target scene; Obtaining a three-dimensional model to be trained, where the three-dimensional model includes a plurality of Gaussian ellipsoids; Training the three-dimensional model using the set of images, and during the training process, adjusting the density of the Gaussian ellipsoids in the three-dimensional model based on the depth information of the target scene in each image of the set of images to enhance the representation ability of the three-dimensional model for the target scene; and Taking the trained three-dimensional model as the reconstruction model corresponding to the target scene.

2. The method according to claim 1, wherein, For any target image in the set of images, adjusting the density of the Gaussian ellipsoids in the three-dimensional model based on the depth information of the target scene in the target image, including: Determining a plurality of feature points of the target scene; and For any target feature point among the plurality of feature points, Determining the actual depth of the target feature point in the target image, and determining the predicted depth of the target feature point in a predicted image, where the predicted image is an image formed by projecting the three-dimensional model onto the camera plane corresponding to the target image, Adjusting the density of the Gaussian ellipsoids in the three-dimensional model based on the predicted depth and the actual depth.

3. The method according to claim 2, wherein The determining the predicted depth of the target feature point in the predicted image includes: Determining a first ellipsoid set in the three-dimensional model, where each Gaussian ellipsoid in the first ellipsoid set is located on a target ray, and the target ray points from the position of the camera center corresponding to the target image to the position of the target feature point; and When the first ellipsoid set is not empty, determining the predicted depth of the target feature point in the predicted image based on the depths of the Gaussian ellipsoids in the first ellipsoid set in the predicted image.

4. The method according to claim 3, wherein The determining the predicted depth of the target feature point in the predicted image based on the depths of the Gaussian ellipsoids in the first ellipsoid set in the predicted image includes: Determining the weight coefficients corresponding to the Gaussian ellipsoids in the first ellipsoid set based on the transparencies corresponding to the Gaussian ellipsoids in the first ellipsoid set; and Performing weighted averaging on the depths corresponding to the Gaussian ellipsoids in the first ellipsoid set according to the corresponding weight coefficients, and taking the result of the weighted averaging as the predicted depth.

5. The method according to claim 3, wherein The method further includes: When the first ellipsoid set is empty, determining the predicted depth of the target feature point in the predicted image to be infinite.

6. The method according to claim 2, wherein The determining the actual depth of the target feature point in the target image includes: Obtaining sparse reconstruction data corresponding to the target scene, where the sparse reconstruction data includes three-dimensional coordinate information of a plurality of point clouds; Obtaining the three-dimensional coordinate information of a target point cloud corresponding to the target feature point from the sparse reconstruction data; and Determining the actual depth of the target feature point in the target image based on the camera parameters corresponding to the target image and the three-dimensional coordinate information of the target point cloud.

7. The method according to claim 6, wherein The obtaining the sparse reconstruction data corresponding to the target scene includes: Performing point cloud reconstruction based on the set of images to obtain the sparse reconstruction data.

8. The method according to claim 2, wherein Each pixel in the target image corresponds to depth information. Determining the actual depth of the target feature point in the target image includes: Determining a target pixel corresponding to the target feature point from the target image; and Taking the depth information corresponding to the target pixel as the actual depth of the target feature point in the target image.

9. The method according to claim 2, wherein the adjusting the density of the Gaussian ellipsoids in the three-dimensional model based on the actual depth and the predicted depth includes: Performing at least one of a first adjustment operation or a second adjustment operation based on the actual depth and the predicted depth to adjust the density of the Gaussian ellipsoids in the three-dimensional model, where the first adjustment operation includes generating a Gaussian ellipsoid at a position corresponding to the actual depth in the three-dimensional model, the second adjustment operation includes determining a second ellipsoid set in the three-dimensional model and performing decomposition processing on each Gaussian ellipsoid in the second ellipsoid set, where each Gaussian ellipsoid in the second ellipsoid set is located on a target ray and forms an occlusion for the target feature point, and the target ray points from the position of the camera optical center corresponding to the target image to the position of the target feature point.

10. The method according to claim 9, wherein, For any target Gaussian ellipsoid in the second ellipsoid set, the decomposition processing includes decomposing the target Gaussian ellipsoid into a plurality of Gaussian ellipsoids, and the volume of each decomposed Gaussian ellipsoid is smaller than the volume of the target Gaussian ellipsoid.

11. The method according to claim 9, wherein The performing at least one of the first adjustment operation or the second adjustment operation based on the actual depth and the predicted depth includes: Performing the first adjustment operation and the second adjustment operation when the predicted depth is less than the actual depth; or Performing the first adjustment operation when the predicted depth is greater than the actual depth.

12. The method according to claim 2, wherein, The determining a plurality of feature points of the target scene includes: Performing feature point detection on the target image to obtain a plurality of feature points of the target scene.

13. The method according to claim 1, wherein Obtaining a three-dimensional model to be trained includes: Performing point cloud reconstruction based on the image set to obtain sparse reconstruction data, where the sparse reconstruction data includes three-dimensional coordinate information of a plurality of point clouds; and Performing parameterization processing on each point cloud in the sparse reconstruction data based on a three-dimensional Gaussian distribution to obtain the plurality of Gaussian ellipsoids.

14. The method according to claim 1, wherein, The method further includes: During the training process, adjusting the density of the Gaussian ellipsoids in the three-dimensional model based on the color information of each image in the image set.

15. A three-dimensional reconstruction system, including: At least one storage medium storing at least one instruction set for performing three-dimensional reconstruction; And At least one processor communicatively connected to the at least one storage medium, where when the three-dimensional reconstruction system runs, the at least one processor reads the at least one instruction set and executes the three-dimensional reconstruction method according to any one of claims 1-14 based on the indication of the at least one instruction set.

Citation Information

Cited By

  • Three-dimensional model real-time generation system based on multi-source data synchronous acquisition equipment

    CN121582497A