Multi-view image reconstruction method and device based on three-dimensional Gaussian sputtering representation
By redesigning the projection scheme of three-dimensional Gaussian elements and adding differentiable rendering of normals and depths, the problems of insufficient multi-view consistency and surface fitting accuracy in the prior art are solved, and higher quality element distribution and depth information are achieved, and grid models that are consistent with the reconstruction target topology can be output.
Patent Information
- Application Number
- CN202510114489.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing three-dimensional Gaussian sputtering method has shortcomings in multi-view consistency, Gaussian element distribution neatness, and surface fitting accuracy. It is impossible to output a grid model consistent with the topology of the reconstruction target, and lacks gradient transfer paths of depth and normals, making it difficult to provide effective supervision information for the reconstruction optimization process.
By redesigning the projection scheme of three-dimensional Gaussian elements, based on the basic principles of projective geometry and quadratic surfaces, differentiable rendering of normals and depth is realized, and the self-supervised Laplace loss function of cumulative normals and three-dimensional rotation loss based on screen space normals are added in the model training step, and the morphological representation and shape constraints of the element are optimized.
The model quality of the reconstruction results is improved, the primitive distribution becomes denser and more reasonable, providing better depth information, and the projection results are consistent and accurate from multiple perspectives, and the grid model that is consistent with the topology of the reconstruction target can be output.
Smart Images

Figure CN120070799A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to three-dimensional reconstruction technology, and in particular to a multi-view image reconstruction method, apparatus, electronic device, computer-readable storage medium, and computer program product based on three-dimensional Gaussian splatting characterization. Background Art
[0002] Three-dimensional Gaussian splatting (3DGS for short) is a three-dimensional reconstruction method that uses a 3D Gaussian probability distribution as the basic primitive. Its main parts include: a differentiable renderer, a visualization tool, and a reconstruction module. Among them, the differentiable rendering is used to rasterize and render 3D Gaussian data, provide the mapping index between pixels and Gaussian primitives, and implement the gradient backpropagation of pixels. The visualization module provides a functional interaction interface for interactive switching of viewpoints and different rendering modes. The reconstruction module is responsible for evaluating the difference between the 3D Gaussian data and the reference image, and gradually optimizing the 3D Gaussian model in combination with the differentiable renderer.
[0003] The Gaussian primitive of the three-dimensional Gaussian reconstruction method is defined by three-dimensional pose (position, rotation, scale) and color attributes. Each primitive has a total of 58 optimizable attributes (one transparency attribute, three attributes for position, rotation, and scale each, and a total of (3 + 1)×2×3 = 48 RGB color attributes defined by third-order spherical harmonics). The differentiable renderer calculates the screen coordinates corresponding to the centers of each Gaussian primitive, the projection depth of the line-of-sight direction, the sampled color in the implementation direction, and the linear approximation result of the non-linear perspective transformation at the primitive center under the given viewport information (including camera pose, viewport size, and clipping distance), and applies this linear approximation transformation to the Gaussian primitive to obtain the projected shape of the three-dimensional Gaussian primitive on the imaging plane (the projection result is a two-dimensional Gaussian distribution). Subsequently, the corresponding index between the screen pixel and the Gaussian primitive is recorded according to the two-dimensional projected shape. All primitives on each screen pixel are sorted from near to far according to the projection depth. Finally, each pixel mixes the RGB colors of the primitives with the two-dimensional Gaussian probability corresponding to each primitive as the transparency to obtain the final rendering result. Among them, the RGB color is used as the direct optimization target, and the parameter gradient is backpropagated along two paths: RGB → spherical harmonic function color, RGB → transparency → two-dimensional Gaussian distribution covariance → projection approximation transformation → three-dimensional pose to realize the supervised training of the Gaussian model. In addition, the reconstruction module also includes an adaptive splitting / cloning function. For primitives with a large transparency gradient, primitives with a scaling factor greater than the threshold are sampled into multiple small primitives according to their three-dimensional probability distribution, or multiple small primitives are cloned from a primitive with a scaling factor less than the threshold.
[0004] The 3D Gaussian reconstruction method has achieved good visual effects in both indoor and outdoor scene reconstruction tasks. Image quality metrics such as Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) have exceeded previous multi-view reconstruction methods. However, the results of 3D Gaussian reconstruction still have deficiencies in multi-view consistency, the neatness of Gaussian primitive distribution, and surface fitting accuracy. It cannot output a mesh model topologically consistent with the reconstruction target and cannot meet the needs of actual scene reconstruction work. At the same time, its differentiable rendering pipeline lacks the gradient transfer paths for depth and normal, making it difficult to provide effective supervision information for the reconstruction optimization process. In addition, existing 3D Gaussian visualization tools only include the results of 3D Gaussian sputtering and the visualization results of 3D Gaussian primitives, which are not convenient for intuitively evaluating the surface reconstruction quality of the reconstruction model, and there is no direct interaction with Gaussian primitives in the visualization tool, so direct editing and adjustment of the reconstruction results cannot be achieved.
[0005] Specifically, the reconstruction module of the 3DGS method only uses RGB images as supervision information. At the same time, the optimization freedom of 3D Gaussian primitives is relatively large, and the model is in an underconstrained state. The optimization process is inevitably prone to overfitting. Its specific manifestations are as follows: the distribution of 3D Gaussian primitives in the model is disorderly, there are a large number of disorderly interpenetrations between primitives, the morphological distribution cannot reflect the surface characteristics of the model, and the visual effect of the local surface may be pieced together by multiple parts at a relatively far distance, and there is only a good visual effect from a specific perspective, and the overall shape does not match the actual situation of the model. Summary of the Invention
[0006] When analyzing the projection scheme of 3D Gaussian primitives, the inventors found that the optimization freedom of 3D Gaussian primitives is too large, the overall training process is in an underconstrained state, and it is prone to situations that do not meet the actual model requirements or even overfitting. At the same time, they realized that the surface normal continuity constraint can effectively guide the Gaussian primitives to fit the actual surface of the model. In addition, the original projection scheme of 3D Gaussian primitives is not based on strict projective geometry derivation, and there is a problem of inconsistent projection results under multiple views, which will interfere with the disparity information required for model training. Therefore, the inventors started from the basic principles of projective geometry and quadric surfaces, redesigned the projection scheme of 3D Gaussian primitives, and realized the differentiable rendering of normal and depth on this basis.
[0007] Aiming at the deficiencies of the prior art, as Figure 3 shown, the present invention proposes a multi-view image reconstruction method based on 3D Gaussian sputtering characterization, which includes:
[0008] A data acquisition step of acquiring multi-view images of the target object to be reconstructed. The multi-view images are the captured images of the target object from multiple views, and each image has corresponding pose information;
[0009] A model training step, according to the multi-view image and the pose information, through single-view rendering, loss calculation, gradient backpropagation and model densification, a three-dimensional Gaussian sputtering model is trained to obtain a final Gaussian primitive;
[0010] In the image reconstruction step, each part of the final Gaussian primitive is labeled, and a rendering method corresponding to the label is used to reconstruct the final Gaussian primitive into a mesh model, and a rendered depth image and an RGB image are generated.
[0011] The multi-view image reconstruction method based on three-dimensional Gaussian sputtering characterization, wherein the model training step includes:
[0012] In the preprocessing step, the Gaussian primitives behind the camera are removed in the camera coordinate system; the position of the camera in the local coordinates of the Gaussian primitive is calculated according to the camera position, the position, rotation and scaling of the Gaussian primitive, that is, S -1 R -1 (CP), where C is the camera position, P, R, and S are the position vector, rotation matrix, and scaling matrix of the Gaussian primitive respectively; remove Gaussian primitives whose camera position vector length in the local coordinate system is less than the preset value; calculate the tangent value t of the cone angle and the direction vector d of the vector according to the length r of the camera position vector in the local coordinate system of the Gaussian primitive; calculate the cone covariance cov of the Gaussian primitive in the local coordinate system l =I-(t 2 +1)*d*d T , where I is the unit matrix; calculate the cone covariance in the camera coordinate system, that is, cov = (S -1 R -1 V -1 )^T*cov l *(S -1 R -1 V -1 ), where V is the camera rotation matrix; the intersection equation cov on the viewing plane is calculated according to the cone equation -1 , where zfar is the viewing plane distance, W and H are the width and height of the screen;
[0013]
[0014] Cone equation: X T cov -1 X=0
[0015] Combined with z=zfar, we can get the intercept equation:
[0016] ax^2+by^2+cz^2+2dxy+2exz+2fyz=0
[0017] a(x - l)^2 + b(y - h)^2 + 2d(x - l)(y - h) = C
[0018] Corresponding:
[0019]
[0020] Corresponding to the two - dimensional Gaussian distribution in screen space:
[0021] Covariance:
[0022] Mean:
[0023] Filter out the Gaussian primitives outside the screen and the Gaussian primitives projected as hyperboloids according to the screen center position and covariance; Calculate the depth reference plane equation in the camera coordinate system according to r and d. In the local coordinate system where x and y are the coordinates of pixels in screen space, the depth reference plane equation is:
[0024]
[0025] In the camera coordinate system, the depth reference plane equation is:
[0026]
[0027] In the camera coordinate system, the ray parameter equation of the line of sight corresponding to each pixel is:
[0028]
[0029] Solving the equation gives the ray parameter u and the depth D = u * zfar;
[0030] Calculate the primitive normal according to R and V, and adjust the normal so that the angle with the camera Z - direction is less than 90 degrees; Calculate the primitive coverage range;
[0031] Calculate the primitive color: Calculate the RGB components separately from the spherical harmonics; When the type attribute is greater than 0, adjust the RGB components according to the configuration;
[0032] Calculate the depth of the primitive at the center of each block according to the reference plane; Sort according to the position ID and the center depth; Each block collects the RGB, two - dimensional covariance, depth plane, and normal information of the relevant primitives; Calculate the probability density function pixel - by - pixel, and screen out the primitives with opacity lower than the preset value; Calculate the depth of the primitive at each pixel according to the reference plane; Mix and blend the RGB, depth, and normal according to the transparency, and normalize the normal blending result;
[0033] Post - processing step: Calculate the screen space normal according to the depth image and normalize the depth image.
[0034] The multi-view image reconstruction method based on three-dimensional Gaussian sputtering characterization, wherein the model training steps include:
[0035] Loss calculation step, calculating the RGB loss L according to the rendering result and the reference image RGB , adding a self-supervised Laplacian loss function R for the cumulative normal N and a three-dimensional rotation loss R based on the screen space normal SN , and a shape constraint regularization term R EC ;
[0036] Calculating the difference of the local normal of the three-dimensional Gaussian sputtering model through the Laplace operator, and minimizing the difference during the optimization process of the three-dimensional Gaussian sputtering model; approximately detecting the edge of the three-dimensional Gaussian sputtering model through the depth image gradient information as the shielding area;
[0037]
[0038] Based on the viewing angle size and depth information, obtaining the camera space coordinates corresponding to the pixels, taking the cross product of the vectors between the three-dimensional points corresponding to the adjacent pixels in the x and y directions of the pixel screen to obtain the normal direction at the pixel, and converting it to the world space coordinate system according to the camera transformation matrix, so as to calculate the supervision difference with the normal image based on the screen space. The calculation of this supervision difference uses a three-dimensional rotation loss function to adapt to the problem that the rasterization result of the three-dimensional Gaussian primitive is non-orientable;
[0039]
[0040] To constrain the primitive shape, the reconstruction module adds a shape constraint regularization term R EC , which is used to guide the primitive to change to a flat and uniform shape:
[0041]
[0042] The optimization target loss function of the reconstruction module is: Loss = L RGB +Mask*(θ D *R D +θ SN *R SN )+R ec , where Mask is the depth image local gradient mask.
[0043] The multi-view image reconstruction method based on three-dimensional Gaussian sputtering characterization, wherein the gradient backpropagation includes: calculating the transparency gradient, depth reference plane gradient, normal gradient, and RGB gradient of each primitive according to the depth, normal, RGB gradient of each pixel, and the transparency of each primitive; calculating the corresponding two-dimensional covariance gradient according to the transparency gradient of each primitive; calculating the rotation, displacement, and scaling gradients of the primitive according to the two-dimensional covariance gradient, depth reference plane gradient, and normal gradient of each primitive; calculating the spherical harmonic function gradient according to the RGB gradient of each primitive;
[0044] The model densification includes: classifying primitives into large primitives and small primitives according to the y-axis length of the primitive and the scene size, cloning the gradient of the small primitive and splitting the large primitive; removing primitives whose size exceeds a preset value; removing primitives whose z-axis length is greater than the x-axis length by a preset multiple and removing primitives whose eccentricity is greater than a preset value.
[0045] The multi-view image reconstruction method based on three-dimensional Gaussian sputtering characterization, wherein the three-dimensional Gaussian sputtering model is labeled and adjusted through a visualization window:
[0046] According to the normal, screen space normal, and depth visualization options, the effect of the three-dimensional Gaussian sputtering model is viewed in real time during training, and by labeling the type of the three-dimensional Gaussian sputtering model, different types of primitives in the three-dimensional Gaussian sputtering model are rendered in different ways.
[0047] As Figure 4 shown, the present invention also proposes a multi-view image reconstruction device based on three-dimensional Gaussian sputtering characterization, which includes:
[0048] A data acquisition module that acquires multi-view images of the target object to be reconstructed. The multi-view images are the captured images of the target object from multiple perspectives, and each image has corresponding pose information;
[0049] A model training module that trains a three-dimensional Gaussian sputtering model based on the multi-view images and the pose information through single-view rendering, Loss calculation, gradient backpropagation, and model densification to obtain the final Gaussian primitives;
[0050] An image reconstruction module that labels each part of the final Gaussian primitives and reconstructs the final Gaussian primitives into a mesh model by using a rendering method corresponding to the label, and generates a rendered depth image and an RGB image.
[0051] The multi-view image reconstruction device based on three-dimensional Gaussian sputtering characterization, wherein the model training module includes:
[0052] Preprocessing module, which eliminates Gaussian primitives behind the camera in the camera coordinate system; calculates the position of the camera in the local coordinates of the Gaussian primitive according to the camera position, the position, rotation, and scaling of the Gaussian primitive, i.e., S -1 R -1 (C - P), where C is the camera position, and P, R, S are the position vector, rotation matrix, and scaling matrix of the Gaussian primitive respectively; eliminates Gaussian primitives with the length of the camera position vector in the local coordinate system of the Gaussian primitive less than a preset value; calculates the tangent value t of the cone surface inclination angle and the direction vector d of this vector according to the length r of the camera position vector in the local coordinate system of the Gaussian primitive; calculates the cone covariance cov in the local coordinate system of the Gaussian primitive l =I - (t 2 + 1)*d*d T , where I is the identity matrix; calculates the cone covariance in the camera coordinate system, i.e., cov = (S -1 R -1 V -1 )^T * cov l *(S -1 R -1 V -1 ), where V is the camera rotation matrix; calculates the intercept equation on the view plane according to the cone equation cov -1 , where zfar is the view plane distance, and W, H are the width and height of the screen;
[0053]
[0054] Cone equation: X T cov -1 X = 0
[0055] Combining z = zfar, the intercept equation can be obtained:
[0056] ax^2 + by^2 + cz^2 + 2dxy + 2exz + 2fyz = 0
[0057] a(x - l)^2 + b(y - h)^2 + 2d(x - l)(y - h) = C
[0058] Corresponding to:
[0059]
[0060] Corresponding to the two-dimensional Gaussian distribution in screen space:
[0061] Covariance:
[0062] Mean:
[0063] Filter out Gaussian primitives outside the screen and Gaussian primitives projected as hyperboloids according to the screen center position and covariance; calculate the depth reference plane equation in the camera coordinate system according to r and d. The depth reference plane equation in the local coordinate system with x and y being the coordinates of the pixels in the screen space is as follows:
[0064]
[0065] In the camera coordinate system, the depth reference plane equation is:
[0066]
[0067] In the camera coordinate system, the ray parameter equation of the line of sight corresponding to each pixel is:
[0068]
[0069] Solving the equation gives the ray parameter u and the depth D = u * zfar;
[0070] Calculate the primitive normal according to R and V, and adjust the normal so that the angle with the camera Z direction is less than 90 degrees; calculate the primitive coverage range;
[0071] Calculate the primitive color: calculate the RGB components separately from the spherical harmonics; when the type attribute is greater than 0, adjust the RGB components according to the configuration;
[0072] Calculate the depth of the primitive at the center of each block according to the reference plane; sort according to the position ID and the center depth; each block collects the RGB, two-dimensional covariance, depth plane, and normal information of the relevant primitives; calculate the probability density function for each pixel, and screen out the primitives with opacity lower than the preset value; calculate the depth of the primitive at each pixel according to the reference plane; mix and blend the RGB, depth, and normal according to the transparency, and normalize the normal blending result;
[0073] The post-processing module calculates the screen space normal according to the depth image and normalizes the depth image;
[0074] The model training module includes:
[0075] The loss calculation module calculates the RGB loss L according to the rendering result and the reference image RGB , add the self-supervised Laplacian loss function R for the cumulative normal N and the three-dimensional rotation loss R based on the screen space normal SN , as well as the shape constraint regularization term R EC ;
[0076] Calculate the difference in local normal vectors of the three-dimensional Gaussian sputtering model through the Laplace operator, and minimize this difference during the optimization process of the three-dimensional Gaussian sputtering model; approximately detect the edges of the three-dimensional Gaussian sputtering model through the depth image gradient information as the shielding area;
[0077]
[0078] Based on the viewing angle size and depth information, obtain the camera space coordinates corresponding to the pixels. Calculate the cross product of the vectors between the three-dimensional points corresponding to the adjacent pixels in the x and y directions of the pixel screen to obtain the normal direction at this pixel, and convert it to the world space coordinate system according to the camera transformation matrix, so as to calculate the supervision difference with the normal image based on the screen space. The calculation of this supervision difference uses a three-dimensional rotation loss function to adapt to the problem that the rasterization result of the three-dimensional Gaussian primitive is non-orientable;
[0079]
[0080] To constrain the shape of the primitive, the reconstruction module adds a shape constraint regularization term R EC , which is used to guide the primitive to change towards a flat and uniform shape:
[0081]
[0082] The optimization objective loss function of the reconstruction module is: Loss = L RGB +Mask*(θ D *R D +θ SN *R SN )+R ec , where Mask is the local gradient mask of the depth image;
[0083] This gradient backpropagation includes: calculating the transparency gradient, depth reference plane gradient, normal gradient, and RGB gradient of each primitive according to the depth, normal, RGB gradient of each pixel and the transparency of each primitive; calculating the corresponding two-dimensional covariance gradient according to the transparency gradient of each primitive; calculating the rotation, displacement, and scaling gradients of the primitive according to the two-dimensional covariance gradient, depth reference plane gradient, and normal gradient of each primitive; calculating the spherical harmonic function gradient according to the RGB gradient of each primitive;
[0084] The densification of this model includes: classifying the primitives into large primitives and small primitives according to the y-axis length of the primitive and the scene size, cloning the gradient of the small primitive and splitting the large primitive; removing the primitives whose size exceeds the preset value; removing the primitives whose z-axis length is greater than the x-axis length by a preset multiple and removing the primitives whose eccentricity is greater than the preset value.
[0085] Annotate and adjust the three-dimensional Gaussian sputtering model through the visualization window:
[0086] According to the normal vector, screen space normal vector, and depth visualization options, the effect of the three-dimensional Gaussian splatter model can be viewed in real time during training. By annotating the types for the three-dimensional Gaussian splatter model, the primitives of different types in the three-dimensional Gaussian splatter model are rendered in different ways.
[0087] The present invention also provides an electronic device, including the multi-view image reconstruction device based on three-dimensional Gaussian splatter characterization described above. The electronic device is either connected to an information display device, which is used to display the mesh model, the rendered depth image, and the RGB image according to the display parameters, attributes set by the user, or through an artificial intelligence model.
[0088] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the multi-view image reconstruction method based on three-dimensional Gaussian splatter characterization are implemented.
[0089] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the multi-view image reconstruction method based on three-dimensional Gaussian splatter characterization are implemented.
[0090] As can be seen from the above solutions, the advantages of the present invention are as follows:
[0091] The quality of the reconstructed result model is high, and the distribution of primitives is dense and reasonable. The present invention can be applied to the reconstruction work of plant leaves, can better perceive the change of leaf surface curvature, and output a three-dimensional Gaussian model with reasonable primitive distribution, flat shape, and fitting the page. At the same time, it is also applied to the reconstruction of outdoor open scenes and indoor scenes. The reconstructed result can provide better depth information compared with the original three-dimensional Gaussian reconstruction method. The reconstruction results of the present invention in various scenes are significantly improved compared with the original version. By collecting multi-view depth rendering results and performing TSDF processing, a better mesh model can be obtained. Subsequently, the model can be further adjusted through three-dimensional modeling software, and it can be more conveniently integrated into the industrial three-dimensional reconstruction process. Its convenience and high efficiency can greatly reduce the reconstruction operation cycle, reduce the reconstruction cost, and improve the reconstruction quality. Description of the Drawings
[0092] Figure 1 It is a comparison schematic diagram of the rasterization scheme of the present invention and the original rasterization scheme;
[0093] Figure 2 It is a schematic diagram of the depth plane D, projected depth, and normal vector calculation scheme;
[0094] Figure 3 It is a flowchart of the method of the present invention;
[0095] Figure 4 It is a module diagram of the device of the present invention;
[0096] Figure 5 Structural schematic diagram of the first electronic device of the present invention;
[0097] Figure 6 Structural schematic diagram of the application environment of the first electronic device of the present invention;
[0098] Figure 7 Structural schematic diagram of the second electronic device of the present invention.
[0099] Reference numerals:
[0100] A - First electronic device;
[0101] B - Multi - perspective image reconstruction device based on three - dimensional Gaussian sputtering characterization;
[0102] C - Data acquisition device;
[0103] D - Information display device;
[0104] 1000 - Second electronic device;
[0105] Ⅰ - Computing unit;
[0106] Ⅱ - ROM;
[0107] Ⅲ - RAM;
[0108] Ⅳ - Bus;
[0109] Ⅴ - Interface;
[0110] Ⅵ - Input unit;
[0111] Ⅶ - Output unit;
[0112] Ⅷ - Storage medium;
[0113] Ⅸ - Communication unit. Detailed implementation manners
[0114] It should be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non - exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0115] Without further limitation, an element qualified by the statement "comprising an..." does not exclude the presence of additional identical elements in a process, method, article, or apparatus that comprises the element.
[0116] The processor described in the present invention is the control center of an electronic device, which can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), or it can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0117] Optionally, the processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.
[0118] In a specific implementation, as an embodiment, the processor can include one or more CPUs. Each of these processors can be a single-CPU or a multi-CPU. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The electronic device can include: servers, desktop computers, laptop computers, smartphones, tablets, embedded computers, etc., where the embedded computer includes vehicles and robots, etc.
[0119] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.
[0120] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto. The actual knowledge structure recognition device may include more or fewer components than shown in the drawings, or combine certain components, or have different component arrangements.
[0121] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0122] It should also be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context.
[0123] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0124] It should also be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0125] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0126] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0127] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0128] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store program codes.
[0129] Aiming at the problems of the 3DGS method, such as lack of effective supervision information, poor surface fitting quality, inconsistent multi-view rendering, and imperfect visualization tools, the present invention proposes a dense three-dimensional Gaussian reconstruction method based on envelope surfaces and surface constraints, optimizes the three-dimensional Gaussian primitive rasterization method, provides a consistent and accurate two-dimensional projection range under multiple views, supplements the normal and depth gradient transfer paths, optimizes the primitive shape representation scheme, constrains the primitive shape, improves the reconstruction training efficiency, adds a self-supervised Laplacian loss function for cumulative normals and a three-dimensional rotation loss based on screen-space normals, optimizes the arrangement characteristics and spatial poses of the three-dimensional Gaussian primitives, obtains a result that better conforms to the surface characteristics of the reconstruction object, and supplements the normal and depth visualization functions to intuitively, conveniently, and at any time view the model details and evaluate the model quality. At the same time, according to actual usage needs, the present invention designs a three-dimensional Gaussian annotation tool to achieve rapid annotation and classification display of three-dimensional Gaussian primitives.
[0130] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are hereinafter given and described in detail in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are merely for illustrative purposes. The scope of protection of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.
[0131] The specific use of the present invention is mainly divided into three stages, including data acquisition, three-dimensional reconstruction, annotation, and meshing.
[0132] In the first stage of data acquisition, multi-view image materials and pose information required for reconstruction need to be obtained. The multi-view images can be obtained by taking photos or shooting videos. When shooting, the camera focal length, exposure, and color temperature need to be adjusted according to the actual shooting environment, and the camera parameters and ambient light need to be unified or fixed during the shooting process. The shooting subject should be kept as still as possible during shooting, and the selected shooting angles should cover the shooting subject as completely as possible. When shooting, according to the reconstruction object or actual needs, various schemes such as hand-held shooting, shooting with a rack or rail, and shooting with a drone can be selected. If shooting a video, video frames with clear and non-shaky images need to be extracted from the video as reconstruction materials. For each reconstruction object, it is advisable to use more than 40 images / frames of reconstruction materials. The camera pose information corresponding to each image / frame of reconstruction materials can be obtained by calculating with the Colmap software. If the materials are shot with a drone or robotic arm or rail with high-precision positioning and gyroscopes, the positioning data of the shooting equipment itself can also be used. The size of the reconstructed images can be adjusted according to the reconstruction accuracy requirements and equipment performance. When the length and width are about 1000px, the reconstruction time and quality are relatively good, and the best image ratio is that the length and width are the same, and the difference between the length and width should not be too large.
[0133] Finally, the acquired images and pose data files are classified and stored in the same folder, and the specific format is the same as the output format of Colmap.
[0134] In the second stage, the training code is called to generate a Gaussian model (3D Gaussian sputtering model). The command line input is:
[0135] python train.py -s data folder path [-m output model path] [--scale <ab|ef|raw>] [--local_normal_loss] [--local_normal_loss_after normal-depth constraint start round] [--vf_loss] [--vf_loss_after normal self-supervision start round]
[0136] Configure the startup parameters according to actual needs. In addition, add "--GUI" to the startup parameters to start the visualization tool while starting the training, and adjust the display mode through the UI slider of the visualization tool.
[0137] In the training code, in the order of data transfer, there are model loading here and scene initialization, single-view rendering, Loss calculation, gradient backpropagation, and model densification. The main code flow of the present invention is as follows:
[0138] Model loading (implemented in Python)
[0139] 1. Load the SfM point cloud model and randomly assign size attributes to each point cloud, including the X-axis length, eccentricity, and flatness ratio
[0140] 2. The size attribute is a three-dimensional vector, and the x, y, and z components correspond to the X-axis length, flatness ratio, and eccentricity respectively
[0141] 3. The X-axis length is e x
[0142] 4. The flatness ratio is the ratio of the Y-axis length to the X-axis length. If the flatness ratio is greater than 1, it is expressed as e y +1, and the Y-axis length of this point is e x *(e y +1)
[0143] 5. The eccentricity is the ratio of the Z-axis length to the Y-axis length, restricted to between 1 and 5, represented by This point's Z-axis length is
[0144] 6. Assign a type attribute to each point cloud and initialize it to 0
[0145] Single-view rendering is divided into four parts: preprocessing, depth sorting, rasterization, and postprocessing
[0146] Preprocessing (Cuda)
[0147] 13. Eliminate Gaussian primitives behind the camera in the camera coordinate system
[0148] 14. Calculate the position of the camera in the local coordinates of the Gaussian primitive, i.e., S, based on the camera position, the position, rotation, and scaling of the Gaussian primitive -1 R -1 (C - P), where C is the camera position, and P, R, S are the position vector, rotation matrix, and scaling matrix of the Gaussian primitive respectively
[0149] 15. Eliminate primitives with the length of the camera position vector less than 1 in the local coordinate system (projected as a virtual spherical surface)
[0150] 16. Calculate the tangent value t of the conical surface inclination angle and the direction vector d of this vector based on the length r of the camera position vector in the local coordinate system
[0151] 17. Calculate the conical covariance cov in the local coordinate system l = I - (t 2 + 1) * d * d T , where I is the identity matrix
[0152] 18. Calculate the conical covariance in the camera coordinate system, i.e., cov = (S -1 R -1 V -1 )^T * cov l * (S -1 R -1 V -1 ), where V is the camera rotation matrix
[0153] 19. Calculate the intercept equation on the view plane according to the conical equation, as follows. Where zfar is the view plane distance, and W, H are the width and height of the picture
[0154]
[0155] Conical equation: X T cov -1 X = 0
[0156] Combined with z = zfar, the intercept equation can be obtained as follows
[0157] ax^2 + by^2 + cz^2 + 2dxy + 2exz + 2fyz = 0
[0158] a(x - l)^2 + b(y - h)^2 + 2d(x - l)(y - h) = C
[0159] Corresponding to
[0160]
[0161] Two-dimensional Gaussian distribution in the corresponding screen space:
[0162] Covariance:
[0163] Mean:
[0164] 20. Filter out Gaussian primitives outside the screen and Gaussian primitives projected as hyperboloids according to the screen center position and covariance
[0165] 21. Calculate the depth reference plane equation in the camera coordinate system according to r and d, where x and y are the coordinates of the pixels in the screen space
[0166] The depth reference plane equation in the local coordinate system is
[0167]
[0168] In the camera coordinate system, the depth reference plane equation is
[0169]
[0170] In the camera coordinate system, the ray parametric equation of the line of sight corresponding to each pixel is
[0171]
[0172] Solving the equation gives the ray parameter u and the depth D = u * zfar
[0173] 22. Calculate the primitive normal according to R and V, and adjust the normal so that the angle with the camera Z direction is less than 90 degrees
[0174] 23. Calculate the primitive coverage
[0175] 24. Calculate the primitive color
[0176] a) Calculate the RGB components separately from the spherical harmonics
[0177] b) When the type attribute is greater than 0, adjust the RGB components according to the configuration, and at this time the color does not participate in the gradient transfer
[0178] Depth sorting (Cuda)
[0179] 1. Calculate the depth of the primitive at the center of each block according to the reference plane, where the block Patch is a square pixel block area, which is automatically allocated by Cuda during rasterization. It is a representation scheme used in the original Gaussian sputtering scheme to accelerate rendering, and generally the block size is 4x4.
[0180] 2. Sort according to the position ID and the center depth
[0181] Rasterization (Cuda)
[0182] 1. Each block collects the RGB, two-dimensional covariance, depth plane, and normal information of relevant primitives
[0183] 2. Calculate the probability density function pixel by pixel and filter out primitives with too low opacity
[0184] 3. Calculate the depth of primitives at each pixel according to the reference plane
[0185] 4. Blend RGB, depth, and normal according to transparency and normalize the normal blending result
[0186] 5. Post-processing (implemented in Pytorch)
[0187] 6. Calculate the screen space normal according to the depth image
[0188] 7. Normalize the depth image
[0189] Loss calculation (implemented in Pytorch)
[0190] 1. Calculate the L1 Loss and SSIM Loss according to the rendering result and the reference image
[0191] 2. Calculate the local gradient mask Mask of the depth image through the Sobel operator. Pixels with gradients less than \(\delta\) are valid regions
[0192] 3. Calculate the local depth difference regularization term through the Laplace operator
[0193] 4. Calculate the screen space normal regularization term through the three-dimensional rotation loss
[0194] 5. Calculate the shape constraint regularization term
[0195] 6. Combine the Loss
[0196] Loss = L RGB + Mask * (\(\theta\) D * R D + \(\theta\) SN * R SN + R ec
[0197] Here, R represents the loss, \(\theta\) is the corresponding weight factor, the subscript D represents the depth loss, SN represents the screen space normal loss, ec represents the eccentricity loss, LossRGB is the image L1Loss and SSIM Loss, and mask is the previous edge MASK
[0198] Gradient backpropagation (implemented in Cuda)
[0199] 5. Calculate the transparency gradient, depth reference plane gradient, normal gradient, and RGB gradient of each primitive according to the pixel depth, normal, RGB gradient, and the transparency of each primitive.
[0200] 6. Calculate the corresponding two-dimensional covariance gradient according to the transparency gradient of each primitive.
[0201] 7. Calculate the rotation, displacement, and scaling gradients of the primitive according to the two-dimensional covariance gradient, depth reference plane gradient, and normal gradient of each primitive.
[0202] 8. Calculate the spherical harmonic function gradient according to the RGB gradient of each primitive.
[0203] Model densification (implemented in Pytorch)
[0204] 1. Distinguish small primitives and large primitives according to the y-axis length of the primitive and the scene size, clone the small primitives with larger gradients, and split the large primitives with larger gradients.
[0205] 2. Remove overly large primitives.
[0206] 3. Remove primitives with a z-axis length much greater than the x-axis length and primitives with an eccentricity greater than 4.8.
[0207] In the third stage, before data annotation, a configuration file named "modifier" can be added to the model output folder. In the file, the visualization colors of different types of components are configured in order. Each line of the file is six floating-point numbers, and each line corresponds to the color offsets of nine types of components numbered 1-9 in sequence. The six floating-point numbers are the relative and absolute offset values of the color values of the three RGB channels respectively. The final color of each type of primitive is the original RGB * relative offset value + absolute offset value. Open the model through Sibr_GaussianViewer_app.exe. In the option box, visualization schemes such as normal, screen space normal, depth, and RGB can be selected, and the original three-dimensional Gaussian rasterization module and the new rasterization module used in the present invention can also be switched. Switch different annotation types through the numeric keys 1-9, clear the annotation type with the numeric key 0, click on the target area with the left mouse button to label the primitives in that area as the set type, and hold down the alt key and the left mouse button and move the mouse to control the camera view. After the annotation is completed, select "Save Gaussian Model" in the menu bar to save the annotation result as a PLY file. When a mesh model needs to be generated, run the script render.py, specify the model folder, generate the rendered depth image and RGB image, and run the TSDF reconstruction script to generate the mesh model. The mesh model is a three-dimensional mesh model, and rendered images from any perspective can be obtained with the help of a renderer.
[0208] Furthermore, in order to achieve the above technical effects, the present invention proposes the following key technical points:
[0209] Key point 1: Implementing 3D Gaussian primitive projection based on envelope surface projection.
[0210] The present invention re-derives the projection equation of 3D primitives from the underlying principles of projective geometry, which is used to calculate the pixel range covered by the primitive and the corresponding projection depth of each pixel. The specific derivation idea is as follows:
[0211] The isosurface of the 3D Gaussian distribution probability density function is a quadratic surface, and the covariance matrix of the 3D Gaussian distribution is a positive semi-definite matrix. Combining the quadratic surface discrimination conditions, it can be known that this quadratic surface is an ellipsoid, and its shape (the lengths of the three axes and the orientation) directly corresponds to the covariance matrix of the Gaussian distribution. We assume that the position E, scaling S, and rotation R of the 3D Gaussian primitive correspond to the expectation and covariance of the 3D Gaussian distribution. Then, a certain isosurface can be selected to make its shape the ellipsoid defined by the pose attributes of the Gaussian primitive. Selecting this ellipsoid as the projection object and examining it in the camera coordinate system, the envelope line of the projection of this ellipsoid on the imaging plane at a distance Z in front of the camera is the curve where the intersection points of all the tangents of the ellipsoid passing through the camera position are located on the imaging plane. That is, the envelope line of the ellipsoid projection is the intersection line I between the envelope surface E formed by the tangents of the ellipsoid passing through the camera position (0,0,0) and the imaging plane. The ellipsoid and the sphere can be transformed into each other through a linear transformation. Assuming that the linear transformation T makes the ellipsoid in the camera space transform into a unit sphere,
[0212] T = S -1 R -1 (-E)
[0213] In the transformed coordinate system (referred to as the local coordinate system), the surface formed by all the tangents of the unit sphere passing through the camera position T(0,0,0) is a conical surface C (or a virtual spherical surface, but this case needs to be excluded)
[0214] where θ is half of the cone angle of the cone
[0215] where
[0216] C s = T(0,0,0)
[0217] d = -Normalize(C s )
[0218] r = x - C s
[0219] Then, in the camera coordinate system, the envelope surface is
[0220]
[0221] Similarly, for the conical surface, the equation of the intersection line between the envelope surface and the imaging plane can be calculated. For the envelope line I, we choose to eliminate the parabola and hyperbola, which are two non-closed figures, according to the conic discriminant, and retain the elliptical curve. It is assumed that the ellipse corresponds to an isosurface of a two-dimensional Gaussian distribution, and the probability density function of the two-dimensional Gaussian distribution determines the attenuation of the projection transparency of the three-dimensional Gaussian primitive at each position on the imaging plane.
[0222] Key point 2: Determine the depth reference plane based on the planar Jacobi transformation and calculate the primitive depth.
[0223] The depth plane D, projection depth, and normal calculation scheme are as Figure 2 shown. At the same time, in the local coordinate system, the intersection line between the conical surface and the unit sphere lies on the same plane. We select this plane as the reference plane D for the projection depth of the Gaussian primitive.
[0224] d T X = (1 - sin 2 θ)|Cs|
[0225] The depth of observing the Gaussian primitive from different viewing directions is the distance from the viewing line to the reference plane. The normal direction of this plane is the d direction. In the camera coordinate system, the reference plane D' is obtained by transforming D through J(T), and the corresponding equation is
[0226] N c = (d T (VRS) -1 ) T
[0227]
[0228] The pixel point P i = (l, h, z far ) corresponds to the depth:
[0229]
[0230] In addition, we select the positive direction of the shortest axis of the ellipse as the normal direction of the three-dimensional Gaussian primitive. Since each plane has positive and negative normals, a unified direction needs to be selected when accumulating the normals. The dot product of the normal direction of each primitive and the positive viewing direction also needs to be calculated. The normal with a dot product less than 0 is reversed. The color sampling of the Gaussian primitive is the same as the original version and is sampled according to the camera direction.
[0231] The probability that a one-dimensional Gaussian distribution is between [-3σ, 3σ] is greater than 99%. Similarly, we calculate the maximum eigenvalue based on the covariance matrix and obtain the maximum variance σ of the two-dimensional Gaussian distribution corresponding to I. We select a pixel block with I as the center and 3σ as the half-side length to determine the rasterization range of the three-dimensional Gaussian primitive. The distance from the ray from the camera position to the center of each pixel block to the primitive depth reference plane D is used as the projection depth of the primitive in the pixel block.
[0232] Key point 3: Calculate screen space normals based on rendered depth image
[0233] Then, each pixel superimposes the color of the two-dimensional Gaussian primitive that may be covered on it according to the projection depth of the pixel block from near to far according to the attenuated transparency. At the same time, the depth is superimposed directly according to the transparency of the three-dimensional Gaussian primitive itself. After the depth image is output, the screen space normal and world space normal can be calculated according to the depth gradient and camera coordinate transformation. In addition, the superposition method of normals is exactly the same as RGB, and the cumulative normal image is obtained. The calculation of the gradient transfer from depth and normal pixels to Gaussian primitives is opposite to the superposition process, and the weights of each primitive are iteratively calculated from far to near.
[0234] At this point, the rasterization rendering of the three-dimensional Gaussian primitive is completed.
[0235] Key point 4: Modify the scaling representation scheme of the 3D Gaussian primitive and select the shortest axis of the primitive as the normal direction.
[0236] The primitives with relatively flat shapes have larger parallax under different viewing angles, which makes it easier to determine the normal direction and calculate the normal gradient. However, the scaling and rotation of the three-dimensional Gaussian can affect the shape of the Gaussian primitive, and the normal gradient is difficult to reasonably guide the changes in the rotation and scaling factors, resulting in the difficulty in converging the training effect. The present invention modifies the representation of the three-dimensional primitive, and the y- and z-axis scaling factors are based on the x-axis scaling factor, and are greater than 1, so that in the local coordinates of the primitive, the x-axis direction is always the shortest axis, constraining the solution space of the three-dimensional Gaussian model and making the normal guidance effect better. At the same time, the flatness and eccentricity are defined by the ratio of the x-, y-, and z-axis scaling factors, and during the training process, a portion of the primitives that are too flat, sharp, and may cause gradient explosions are eliminated according to the flatness and eccentricity.
[0237] Key point 5: Add Laplace self-supervised loss, 3D rotation loss, and shape constraint regularization terms to the cumulative normal, screen space normal, and primitive shape respectively.
[0238] In the reconstruction module, in addition to the RGB supervision information, a self-supervised Laplace loss function R of the cumulative normal is added N and the 3D rotation loss R based on the screen space normal SN , and the shape constraint regularization term R EC .
[0239] Based on the assumption of smooth model surface, it can be approximately considered that the normal direction of the model surface changes uniformly. The difference of the local normal of the model is calculated by the Laplace operator, and this difference should be minimized during the model optimization process. At the same time, because the normal smoothness assumption is limited to the same surface and the normals between different surfaces do not affect each other, the model edges can be approximately detected through the depth image gradient information as the shielding area.
[0240]
[0241] i and j are used to represent the xy-direction index coordinates of the image. Here, all pixels of the image are traversed, where Δ represents the Laplace operator;
[0242] Based on the viewing angle size and depth information, the camera space coordinates corresponding to the pixels can be estimated. The cross product of the vectors between the three-dimensional points corresponding to the adjacent pixels in the x and y directions of the pixel screen can approximately represent the normal direction at this pixel, and it can be transformed to the world space coordinate system according to the camera transformation matrix, so as to calculate the supervised difference with the normal image based on the screen space. The calculation of this difference uses a three-dimensional rotation loss function to adapt to the problem that the rasterization result of the three-dimensional Gaussian primitive is non-orientable. Similarly, the self-supervised loss based on the screen space normal also needs to shield the edge area.
[0243]
[0244] In addition, to constrain the primitive shape, the reconstruction module can also selectively add a shape constraint regularization term R EC , which is used to guide the primitive to change into a flat and uniform shape.
[0245]
[0246] r represents the radius, rx, ry, and rz respectively represent the radii of the Gaussian primitive in the x, y, and z directions, θ is the weight parameter, the subscript f represents the flatness coefficient, and the subscript e represents the ellipticity coefficient;
[0247] Finally, the optimization objective of the reconstruction module is:
[0248] L = Loss RGB + θ N * R N + θ SN * R SN + R EC
[0249] Key Point 6: Optimize the three-dimensional Gaussian visualization tool to provide rendering mode switching, model grouping annotation, and rendering functions
[0250] During training, the model effect can be viewed in real time through sibr_remoteGaussian_app_d.exe. To better supervise the training process, we added visualization options for normal vectors, screen space normal vectors, and depth, and controlled different display modes through sliders.
[0251] In our visualization tool, different annotation types can be switched by pressing keys 0-9, the visualization perspective can be switched in FPS mode, Gaussian primitives within a specified area can be marked interactively with the mouse button, and different types of primitives are rendered in different colors. The annotated model can be saved as a new model with classification information. In addition to switching multiple rendering results, the visualization tool can also seamlessly switch between the original rasterization rendering and the new rasterization method of this method in real time. The comparison of the rasterization scheme effect of the present invention with the original rasterization scheme is Figure 1 as shown.
[0252] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiment.
[0253] As Figure 4 shown, the present invention also proposes a multi-view image reconstruction device based on three-dimensional Gaussian sputtering characterization, which includes:
[0254] A data acquisition module that acquires multi-view images of the target object to be reconstructed. The multi-view images are captured images of the target object from multiple perspectives, and each image has corresponding pose information;
[0255] A model training module that trains a three-dimensional Gaussian sputtering model based on the multi-view images and the pose information through single-view rendering, Loss calculation, gradient backpropagation, and model densification to obtain the final Gaussian primitives;
[0256] An image reconstruction module that annotates each part of the final Gaussian primitives and reconstructs the final Gaussian primitives into a mesh model using a rendering method corresponding to the annotation, and generates a rendered depth image and an RGB image.
[0257] The multi-view image reconstruction device based on three-dimensional Gaussian sputtering characterization, wherein the model training module includes:
[0258] A preprocessing module that eliminates Gaussian primitives behind the camera in the camera coordinate system; calculates the position of the camera in the local coordinates of the Gaussian primitive according to the position, rotation, and scaling of the camera and the Gaussian primitive, that is, S -1 R -1(C - P), where C is the camera position, and P, R, S are the position vector, rotation matrix, and scaling matrix of the Gaussian primitive respectively; eliminate the Gaussian primitives whose length of the camera position vector in the local coordinate system of the Gaussian primitive is less than a preset value; calculate the tangent value t of the conical surface inclination angle and the direction vector d of this vector according to the length r of the camera position vector in the local coordinate system of the Gaussian primitive; calculate the conical covariance cov in the local coordinate system of the Gaussian primitive l = I - (t 2 + 1)*d*dT, where I is the identity matrix; calculate the conical covariance in the camera coordinate system, that is, cov = (S -1 R -1 V -1 )^T*cov l *(S -1 R -1 V -1 ), where V is the camera rotation matrix; calculate the intercept equation on the view plane according to the conical equation cov -1 , where zfar is the view plane distance, and W, H are the width and height of the picture;
[0259]
[0260] Conical equation: X T cov -1 X = 0
[0261] Combined with z = zfar, the intercept equation can be obtained:
[0262] ax^2 + by^2 + cz^2 + 2dxy + 2exz + 2fyz = 0
[0263] a(x - l)^2 + b(y - h)^2 + 2d(x - l)(y - h) = C
[0264] Corresponding to:
[0265]
[0266] Corresponding to the two-dimensional Gaussian distribution in the screen space:
[0267] Covariance:
[0268] Mean value:
[0269] Filter out the Gaussian primitives outside the picture and the Gaussian primitives projected as hyperboloids according to the screen center position and covariance; calculate the depth reference plane equation in the camera coordinate system according to r and d, where x and y are the coordinates of the pixels in the local coordinate system of the screen space. The depth reference plane equation is:
[0270]
[0271] In the camera coordinate system, the equation of the depth reference plane is:
[0272]
[0273] In the camera coordinate system, the ray parametric equation of the line of sight corresponding to each pixel is:
[0274]
[0275] Solving the equation gives the ray parameter u and the depth D = u * zfar;
[0276] Calculate the primitive normal according to R and V, and adjust the normal so that the angle with the camera Z direction is less than 90 degrees; calculate the coverage range of the primitive;
[0277] Calculate the primitive color: calculate the RGB components separately from the spherical harmonics; when the type attribute is greater than 0, adjust the RGB components according to the configuration;
[0278] Calculate the depth of the primitive at the center of each block according to the reference plane; sort according to the position ID and the center depth; each block collects the RGB, two-dimensional covariance, depth plane, and normal information of the relevant primitives; calculate the probability density function pixel by pixel, and filter out the primitives with opacity lower than the preset value; calculate the depth of the primitive at each pixel according to the reference plane; mix and blend the RGB, depth, and normal according to the transparency, and normalize the normal blending result;
[0279] The post-processing module calculates the screen space normal according to the depth image and normalizes the depth image;
[0280] The model training module includes:
[0281] The loss calculation module calculates the RGB loss L according to the rendering result and the reference image RGB , add the self-supervised Laplacian loss function R for the cumulative normal N and the 3D rotation loss R based on the screen space normal SN , and the shape constraint regularization term R EC ;
[0282] Calculate the difference of the local normal of the 3D Gaussian sputtering model through the Laplace operator, and the difference should be minimized during the optimization process of the 3D Gaussian sputtering model; approximately detect the edge of the 3D Gaussian sputtering model through the depth image gradient information as the shielding area;
[0283]
[0284] Based on the viewing angle size and depth information, the camera space coordinates corresponding to the pixels are obtained. The cross product of the position vectors between the three-dimensional points corresponding to the adjacent pixels in the x and y directions of the pixel screen is calculated to obtain the normal direction at that pixel, and it is transformed to the world space coordinate system according to the camera transformation matrix, so as to calculate the supervision difference with the normal image based on the screen space. The calculation of this supervision difference uses a three-dimensional rotation loss function to adapt to the problem that the rasterization result of the three-dimensional Gaussian primitive is non-orientable;
[0285]
[0286] To constrain the primitive shape, the reconstruction module adds a shape constraint regularization term R EC , which is used to guide the primitive to change to a flat and uniform shape:
[0287]
[0288] The optimization objective loss function of the reconstruction module is: Loss = L RGB +Mask*(θ D *R D +θ SN *R SN )+R ec , where Mask is the local gradient mask of the depth image;
[0289] This gradient backpropagation includes: calculating the transparency gradient, depth reference plane gradient, normal gradient, and RGB gradient of each primitive according to the depth, normal, RGB gradient of each pixel and the transparency of each primitive; calculating the corresponding two-dimensional covariance gradient according to the transparency gradient of each primitive; calculating the rotation, displacement, and scaling gradients of the primitive according to the two-dimensional covariance gradient, depth reference plane gradient, and normal gradient of each primitive; calculating the spherical harmonic function gradient according to the RGB gradient of each primitive;
[0290] The densification of this model includes: classifying the primitives into large primitives and small primitives according to the y-axis length of the primitive and the scene size, cloning the gradient of the small primitive and splitting the large primitive; removing the primitives whose size exceeds the preset value; removing the primitives whose z-axis length is greater than the x-axis length by a preset multiple and removing the primitives whose eccentricity is greater than the preset value.
[0291] The three-dimensional Gaussian sputtering model is labeled and adjusted through a visualization window:
[0292] According to the normal, screen space normal, and depth visualization options, the effect of the three-dimensional Gaussian sputtering model can be viewed in real time during training, and by labeling the type for this three-dimensional Gaussian sputtering model, different types of primitives in this three-dimensional Gaussian sputtering model are rendered in different ways.
[0293] Such as Figure 5As shown in the figure, in another embodiment, the present invention further provides a first electronic device A, including the multi-view image reconstruction device based on three-dimensional Gaussian sputtering characterization described above.
[0294] As Figure 6 shown, the first electronic device A can also be connected to a data acquisition device C and an information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to acquire multi-view images, such as the multi-view videos or pictures described in the embodiments of the present invention. The information display device D is used to display the mesh body model, the rendered depth image, and the RGB image analyzed by the present invention.
[0295] Among them, the information display device D can process and organize the data output by the first electronic device A based on an information display mechanism to improve the readability of the data output by the first electronic device A. This information display mechanism can be preset manually. For example, the data output by the first electronic device A is visually displayed, and it can display according to the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll and play, etc. The specified key information, such as the rendered depth image and RGB image of a specific perspective, is presented to the user, enabling the user to understand this information more promptly without having to access a secondary page or scroll the page, saving the user's operations. Or this information display mechanism can be an artificial intelligence AI display model, which can learn the user's key attention information based on the user's previous usage habits, such as viewing duration, click times, editing times, etc., and then automatically present rich and necessary key information to the user.
[0296] The present invention also provides a computer program product. The computer program product includes a computer program, which can be stored on a readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-view image reconstruction method based on three-dimensional Gaussian sputtering characterization provided by the above-mentioned various methods.
[0297] In another embodiment, the present invention further provides a storage medium VIII for storing a computer program for executing the multi-view image reconstruction method based on three-dimensional Gaussian sputtering characterization. It should be understood that the storage medium in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0298] Figure 7 FIG. shows a schematic block diagram of a second electronic device 1000 that can be used to implement the embodiments of the present invention. The second electronic device 1000 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present invention described and / or claimed herein. The second electronic device 1000 may be the same as or different from the first electronic device A.
[0299] The second electronic device 1000 includes a computing unit Ⅰ, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory Ⅱ (ROM) or a computer program loaded from a storage medium Ⅷ into a random access memory (RAM) Ⅲ. In the RAM Ⅲ, various programs and data required for the operation of the device 1000 can also be stored. The computing unit Ⅰ, the ROM Ⅱ, and the RAM Ⅲ are connected to each other through a bus Ⅳ. An input / output (I / O) interface Ⅴ is also connected to the bus Ⅳ.
[0300] Multiple components in the second electronic device 1000 are connected to the I / O interface Ⅴ, including: an input unit Ⅵ, such as a keyboard, a mouse, etc.; an output unit Ⅶ, such as various types of displays, speakers, etc.; a storage medium Ⅷ, such as a magnetic disk, an optical disc, etc.; and a communication unit Ⅸ, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit Ⅸ allows the second electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0301] The computing unit Ⅰ can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit Ⅰ include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit Ⅰ executes the various methods and processes described above, such as method steps S1 - S3. For example, in some embodiments, the method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage medium Ⅷ. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM Ⅱ and / or the communication unit Ⅸ. When the computer program is loaded into the RAM Ⅲ and executed by the computing unit Ⅰ, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit Ⅰ can be configured to execute the method in any other appropriate way (for example, by means of firmware).
[0302] Although the embodiments of the present invention have been disclosed as above, they are not limited to only the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the illustrations shown and described herein.
Claims
1. A multi-view image reconstruction method based on three-dimensional Gaussian sputtering characterization, characterized in that: include: A data acquisition step, acquiring multi-view images of the target object to be reconstructed, wherein the multi-view images are images of the target object taken at multiple viewpoints, and each image has corresponding posture information; A model training step, according to the multi-view image and the pose information, through single-view rendering, loss calculation, gradient backpropagation and model densification, a three-dimensional Gaussian sputtering model is trained to obtain a final Gaussian primitive; In the image reconstruction step, each part of the final Gaussian primitive is labeled, and a rendering method corresponding to the label is used to reconstruct the final Gaussian primitive into a mesh model, and a rendered depth image and an RGB image are generated.
2. The multi-view image reconstruction method based on three-dimensional Gaussian sputtering characterization according to claim 1, characterized in that: The model training steps include: In the preprocessing step, the Gaussian primitives behind the camera are removed in the camera coordinate system; the position of the camera in the local coordinates of the Gaussian primitive is calculated according to the camera position, the position, rotation and scaling of the Gaussian primitive, that is, S -1 R -1 (CP), where C is the camera position, P, R, and S are the position vector, rotation matrix, and scaling matrix of the Gaussian primitive respectively; remove Gaussian primitives whose camera position vector length in the local coordinate system is less than the preset value; calculate the tangent value t of the cone angle and the direction vector d of the vector according to the length r of the camera position vector in the local coordinate system of the Gaussian primitive; calculate the cone covariance cov of the Gaussian primitive in the local coordinate system l =I-(t 2 +1)*d*d T , where I is the unit matrix; calculate the cone covariance in the camera coordinate system, that is, cov = (S -1 R - 1 V -1 )^T*cov l *(S -1 R -1 V -1 ), where V is the camera rotation matrix; the intersection equation cov on the viewing plane is calculated according to the cone equation -1 , where zfar is the viewing plane distance, W and H are the width and height of the screen; Cone equation: X T cov -1 X=0 Combined with z=zfar, we can get the intercept equation: ax^2+by^2+cz^2+2dxy+2exz+2fyz=0 a(xl)^2+l(yh)^2+2d(xl)(yh)=C correspond: Corresponding to the two-dimensional Gaussian distribution in screen space: Covariance: Mean: According to the center position and covariance of the screen, the Gaussian primitives outside the screen and the Gaussian primitives projected as hyperboloids are filtered out; the depth reference plane equation in the camera coordinate system is calculated according to r and d, where x and y are the coordinates of the pixel in the screen space. The depth reference plane equation in the local coordinate system is: In the camera coordinate system, the depth reference plane equation is: In the camera coordinate system, the ray parameter equation of the line of sight corresponding to each pixel is: Solving the equations yields the ray parameter u and depth D = u*zfar; Calculate the primitive normal based on R and V, and adjust the normal to be less than 90 degrees with the camera Z direction; calculate the primitive coverage; Calculate the color of the primitive: calculate the RGB components from the spherical harmonics respectively; when the type attribute is greater than 0, adjust the RGB components according to the configuration; Calculate the depth of the primitive at the center of each block based on the reference plane; Sort by location ID and center depth; Collect RGB, 2D covariance, depth plane, and normal information of related primitives for each block; Calculate probability density function pixel by pixel to filter out primitives with opacity lower than a preset value; Calculate the depth of the primitive at each pixel based on the reference plane; Mix RGB, depth, and normal based on transparency, and normalize the normal mixing result; In the post-processing step, the screen space normals are calculated based on the depth image and the depth image is normalized.
3. The multi-view image reconstruction method based on three-dimensional Gaussian sputtering characterization according to claim 1 or 2, characterized in that: The model training steps include: Loss calculation step: calculate the RGB loss L based on the rendering result and the reference image RGB , add the self-supervised Laplace loss function R of the cumulative normal N and the 3D rotation loss R based on the screen space normal SN , and the shape constraint regularization term R EC ; The difference of local normals of the three-dimensional Gaussian sputtering model is calculated by Laplace operator, and the difference should be minimized during the optimization process of the three-dimensional Gaussian sputtering model; the edge of the three-dimensional Gaussian sputtering model is approximately detected by depth image gradient information as a shielding area; Based on the viewing angle and depth information, the camera space coordinates corresponding to the pixel are obtained, and the radial vector cross product between the three-dimensional points corresponding to the adjacent pixels in the x and y directions of the pixel screen is performed to obtain the normal direction of the pixel, and then converted to the world space coordinate system according to the camera transformation matrix, so as to calculate the supervised difference with the normal image based on the screen space. The calculation of the supervised difference adopts the three-dimensional rotation loss function to adapt to the problem that the rasterization result of the three-dimensional Gaussian primitive is not orientable; To constrain the shape of the primitive, the reconstruction module adds a shape constraint regularization term R EC , which is used to guide the primitives to change to a flat and uniform shape: The optimization objective loss function of the reconstruction module is: Loss = L RGB +Mask*(θ D *R D +θ SN *R SN )+R ec , where Mask is the local gradient mask of the depth image.
4. The multi-view image reconstruction method based on three-dimensional Gaussian sputtering characterization according to claim 1 or 2, characterized in that: The gradient return includes: calculating the transparency gradient, depth reference plane gradient, normal gradient, and RGB gradient of each primitive according to the depth, normal, RGB gradient of each pixel and the transparency of each primitive; calculating the corresponding two-dimensional covariance gradient according to the transparency gradient of each primitive; calculating the rotation, displacement, and scaling gradient of the primitive according to the two-dimensional covariance gradient, depth reference plane gradient, and normal gradient of each primitive; and calculating the spherical harmonic function gradient according to the RGB gradient of each primitive; The model densification includes: classifying primitives into large primitives and small primitives according to the y-axis length of the primitive and the scene size, cloning the gradient of the small primitive and splitting the large primitive; eliminating primitives whose primitive size exceeds a preset value; eliminating primitives whose z-axis length is greater than the x-axis length by a preset multiple and eliminating primitives whose eccentricity is greater than a preset value.
5. The multi-view image reconstruction method based on three-dimensional Gaussian sputtering characterization according to claim 1 or 2, characterized in that: Annotate and adjust the 3D Gaussian sputtering model through the visualization window: According to the normal, screen space normal, and depth visualization options, the effect of the 3D Gaussian sputtering model can be viewed in real time during training, and different types of primitives in the 3D Gaussian sputtering model can be rendered differently by labeling the type of the 3D Gaussian sputtering model.
6. A multi-view image reconstruction device based on three-dimensional Gaussian sputtering characterization, characterized in that: include: A data acquisition module collects and obtains multi-view images of the target object to be reconstructed, wherein the multi-view images are images of the target object taken at multiple viewpoints, and each image has corresponding posture information; A model training module, according to the multi-view image and the pose information, trains a three-dimensional Gaussian sputtering model through single-view rendering, loss calculation, gradient return and model densification to obtain a final Gaussian primitive; The image reconstruction module annotates each part of the final Gaussian primitive, and uses a rendering method corresponding to the annotation to reconstruct the final Gaussian primitive into a mesh model, and generates a rendered depth image and an RGB image.
7. The multi-view image reconstruction device based on three-dimensional Gaussian sputtering characterization according to claim 6, characterized in that: The model training module includes: The preprocessing module removes the Gaussian primitives behind the camera in the camera coordinate system; the camera position in the local coordinates of the Gaussian primitive is calculated according to the camera position, the position, rotation and scaling of the Gaussian primitive, that is, S -1 R -1 (CP), where C is the camera position, P, R, and S are the position vector, rotation matrix, and scaling matrix of the Gaussian primitive respectively; remove Gaussian primitives whose camera position vector length in the local coordinate system is less than the preset value; calculate the tangent value t of the cone angle and the direction vector d of the vector according to the length r of the camera position vector in the local coordinate system of the Gaussian primitive; calculate the cone covariance cov of the Gaussian primitive in the local coordinate system l = l-(t 2 +1)*d*d T , where I is the unit matrix; calculate the cone covariance in the camera coordinate system, that is, cov = (S -1 R - 1 V -1 )^T*cov l *(S -1 R -1 V -1 ), where V is the camera rotation matrix; the intersection equation cov on the viewing plane is calculated according to the cone equation -1 , where zfar is the viewing plane distance, W and H are the width and height of the screen; Cone equation: X T cov -1 X=0 Combined with z=zfar, we can get the intercept equation: ax^2+by^2+cz^2+2dxy+2exz+2fyz=0 a(xl)^2+b(yh)^2+2d(xl)(yh)=C correspond: Corresponding to the two-dimensional Gaussian distribution in screen space: Covariance: Mean: According to the center position and covariance of the screen, the Gaussian primitives outside the screen and the Gaussian primitives projected as hyperboloids are filtered out; the depth reference plane equation in the camera coordinate system is calculated according to r and d, where x and y are the coordinates of the pixel in the screen space. The depth reference plane equation in the local coordinate system is: In the camera coordinate system, the depth reference plane equation is: In the camera coordinate system, the ray parameter equation of the line of sight corresponding to each pixel is: Solving the equations yields the ray parameter u and depth D = u*zfar; Calculate the primitive normal based on R and V, and adjust the normal to be less than 90 degrees with the camera Z direction; calculate the primitive coverage; Calculate the color of the primitive: calculate the RGB components from the spherical harmonics respectively; when the type attribute is greater than 0, adjust the RGB components according to the configuration; Calculate the depth of the primitive at the center of each block based on the reference plane; Sort by location ID and center depth; Collect RGB, 2D covariance, depth plane, and normal information of related primitives for each block; Calculate probability density function pixel by pixel to filter out primitives with opacity lower than a preset value; Calculate the depth of the primitive at each pixel based on the reference plane; Mix RGB, depth, and normal based on transparency, and normalize the normal mixing result; The post-processing module calculates the screen space normal based on the depth image and normalizes the depth image; The model training module includes: Loss calculation module, calculates RGB loss L based on the rendering result and reference image RGB , add the self-supervised Laplace loss function R of the cumulative normal N and the 3D rotation loss R based on the screen space normal SN , and the shape constraint regularization term R EC ; The difference of local normals of the three-dimensional Gaussian sputtering model is calculated by Laplace operator, and the difference should be minimized during the optimization process of the three-dimensional Gaussian sputtering model; the edge of the three-dimensional Gaussian sputtering model is approximately detected by depth image gradient information as a shielding area; Based on the viewing angle and depth information, the camera space coordinates corresponding to the pixel are obtained, and the radial vector cross product between the three-dimensional points corresponding to the adjacent pixels in the x and y directions of the pixel screen is performed to obtain the normal direction of the pixel, and then converted to the world space coordinate system according to the camera transformation matrix, so as to calculate the supervised difference with the normal image based on the screen space. The calculation of the supervised difference adopts the three-dimensional rotation loss function to adapt to the problem that the rasterization result of the three-dimensional Gaussian primitive is not orientable; To constrain the shape of the primitive, the reconstruction module adds a shape constraint regularization term R EC , which is used to guide the primitives to change to a flat and uniform shape: The optimization objective loss function of the reconstruction module is: Loss = L RGB +Mask*(θ D *R D +θ SN *R SN )+R ec , where Mask is the local gradient mask of the depth image; The gradient return includes: calculating the transparency gradient, depth reference plane gradient, normal gradient, and RGB gradient of each primitive according to the depth, normal, RGB gradient of each pixel and the transparency of each primitive; calculating the corresponding two-dimensional covariance gradient according to the transparency gradient of each primitive; calculating the rotation, displacement, and scaling gradient of the primitive according to the two-dimensional covariance gradient, depth reference plane gradient, and normal gradient of each primitive; and calculating the spherical harmonic function gradient according to the RGB gradient of each primitive; The model densification includes: classifying primitives into large primitives and small primitives according to the y-axis length of the primitive and the scene size, cloning the gradient of the small primitive and splitting the large primitive; eliminating primitives whose primitive size exceeds a preset value; eliminating primitives whose z-axis length is greater than the x-axis length by a preset multiple and eliminating primitives whose eccentricity is greater than a preset value. Annotate and adjust the 3D Gaussian sputtering model through the visualization window: According to the normal, screen space normal, and depth visualization options, the effect of the 3D Gaussian sputtering model can be viewed in real time during training, and different types of primitives in the 3D Gaussian sputtering model can be rendered differently by labeling the type of the 3D Gaussian sputtering model.
8. An electronic device, characterized in that: It includes the multi-view image reconstruction device based on three-dimensional Gaussian sputtering characterization as described in claim 6 or 7, and the electronic device is connected to an information display device, and the information display device is used to display the grid model, the rendered depth image and the RGB image according to the display parameters and attributes set by the user or through an artificial intelligence model.
9. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the multi-view image reconstruction method based on three-dimensional Gaussian sputtering characterization described in any one of claims 1 to 5 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the multi-view image reconstruction method based on three-dimensional Gaussian sputtering characterization described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Scene three-dimensional reconstruction method based on prior depth and Gaussian sputtering model fusion
CN118351252A
Virtual human arbitrary view angle rendering method and system based on three-dimensional Gaussian spattering
CN118736092A
Reconstruction method and device of three-dimensional dynamic scene and storage medium
CN119295651A
Systems and methods for efficient floorplan generation from 3D scans of indoor scenes
US20210279950A1
Cited By
Method and system for enhancing visual positioning based on cross-domain three-dimensional Gaussian sputtering
CN120388074A
Non-rigid three-dimensional editing method and system based on cross-modal attention guidance
CN120655802A
Near-infrared assisted low-light scene three-dimensional reconstruction method based on 3D Gaussian splashing
CN121280638A
Urban building agent reconstruction method based on aerial images
CN122156490A