Gaussian sputtering three-dimensional reconstruction method based on joint attention mechanism

By introducing a joint attention mechanism in the Gaussian sputtering three-dimensional reconstruction method and optimizing position weights using channel and spatial attention mechanisms, the problem of insufficient attention to inter-Gaussian position correlation in the existing methods is solved, and the quality and local performance of the three-dimensional reconstruction images are significantly improved.

CN120014168AActive Publication Date: 2025-05-16CHONGQING UNIV

Patent Information

Application Number
CN202510103687.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction method of Gaussian sputtering has shortcomings in image reconstruction quality and local performance, mainly due to the lack of attention to inter-Gaussian position correlation.

Method used

The Gaussian sputtering three-dimensional reconstruction method based on the joint attention mechanism is adopted. By constructing a joint attention neural network model, using the channel attention mechanism and spatial attention mechanism, the position correlation between different Gaussians is discovered and utilized, and the position weight coefficient is optimized to improve image reconstruction.

Benefits of technology

The overall quality and local performance of the three-dimensional reconstruction images were improved, and the SSIM index with an average of 0.852 and the PSNR index with an average of 24.02 were obtained on the Tanks&Temples dataset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014168A_ABST
    Figure CN120014168A_ABST
Patent Text Reader

Abstract

The invention relates to the field of three-dimensional reconstruction, in particular to a Gaussian sputtering three-dimensional reconstruction method based on a joint attention mechanism, which combines a channel attention mechanism and a space attention mechanism, so that a model can discover and utilize possible position correlation between different gauss, and can realize the three-dimensional reconstruction of the gauss. Different parameters of the three-dimensional Gaussian are optimized through the position weight, and the optimized parameters are used for reconstruction and rendering, so that the overall quality and local performance of the three-dimensional reconstructed image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-dimensional reconstruction, and in particular to a Gaussian sputtering three-dimensional reconstruction method based on a joint attention mechanism. Background Art

[0002] Three-Dimensional Reconstruction (3D Reconstruction) refers to the establishment of a mathematical model suitable for computer representation and processing of three-dimensional objects. This process is the basis for processing, operating and analyzing the properties of three-dimensional objects in a computer environment. It is also a key technology for establishing virtual reality in computers to express the objective world. 3D reconstruction is widely used in many fields such as computer science and technology, computer applications, multimedia, computer graphics, and computer vision.

[0003] Gaussian sputtering technology simulates and reconstructs three-dimensional scenes through small Gaussian ellipsoid particles, which are defined by four important properties: position, covariance matrix, opacity, and spherical harmonics (SH) coefficients. Among them, the position determines the position of the center of the ellipsoid in space; the covariance matrix determines the degree of scaling and rotation of the ellipsoid; the opacity determines the degree to which the ellipsoid blocks light; and the spherical harmonics (SH) coefficients determine the color distribution on the surface of the ellipsoid. Compared with other methods that require a lot of memory and computing resources (such as NeRF), Gaussian sputtering technology still has certain advantages in memory and computing requirements. However, in reality, the position correlation between a group of Gaussians is often ignored in related research, which to some extent affects the overall quality and local performance of image reconstruction.

[0004] Due to the lack of attention to positional correlation, the current Gaussian sputtering 3D reconstruction methods often have certain deficiencies in image reconstruction quality and local image performance. Therefore, it is of great significance to find a method to adaptively pay attention to the possible positional correlation between different Gaussians. Summary of the invention

[0005] The present invention aims to provide a Gaussian sputtering 3D reconstruction method based on a joint attention mechanism to solve the technical problem that the current Gaussian sputtering 3D reconstruction method often has certain deficiencies in image reconstruction quality and local image performance due to the lack of attention to position correlation.

[0006] The Gaussian sputtering 3D reconstruction method based on the joint attention mechanism in the present invention comprises the following steps:

[0007] Step 1: Construct the joint attention neural network model for obtaining position weight coefficients according to the following strategy:

[0008] The input of the joint attention neural network model includes two parts, one of which is a pre-set three-plane data structure, which includes a trainable original three-plane data. The three-plane data structure is mainly used to store feature vectors representing Gaussian position coordinates. The feature vector is a trainable parameter, and its structure is three matrices of the same size (H, W, C), where H represents height, W represents width, and C represents depth or the number of channels;

[0009] The original three-plane data are respectively input into the channel attention network module for internal self-attention calculation and the spatial attention network module for spatial attention calculation;

[0010] The three-plane data after the channel attention network module and the spatial attention network module are orthogonalized to obtain the channel three-plane and the spatial three-plane;

[0011] Another part of the model input is the position coordinates of each Gaussian point cloud data in a set of 3D Gaussian point clouds constructed by the same 3D reconstruction scene; by projecting the position coordinates onto the channel three planes and the spatial three planes respectively, the first attention feature vector corresponding to the channel three planes and the second attention feature vector corresponding to the spatial three planes are obtained;

[0012] The first attention feature vector and the second attention feature vector are concatenated, and the concatenated feature vector is sent to a multi-layer perceptron MLP for perception to obtain a position weight coefficient corresponding to the position coordinate;

[0013] Step 2: Input the position coordinates contained in the 3D Gaussian point cloud to be reconstructed and the optimized original three-plane data into the optimized joint attention neural network model to obtain the corresponding position weight coefficient;

[0014] Step 3: Optimizing the parameters to be optimized for Gaussian sputtering three-dimensional reconstruction contained in the reconstructed 3D Gaussian point cloud using the position weight coefficient to obtain optimized parameters;

[0015] Step 4 uses the 3D Gaussian point cloud with optimized parameters to be optimized to perform Gaussian sputtering 3D reconstruction.

[0016] Furthermore, the 3D reconstructed scene includes multiple high-definition images taken from different angles and camera parameters corresponding to each high-definition image, which can construct a set of sparse SFM point clouds and are initialized as a set of 3D Gaussian point clouds;

[0017] Each Gaussian point cloud includes position coordinates, parameters to be optimized and a covariance matrix. The parameters to be optimized include rotation coefficients, scaling coefficients and opacity.

[0018] Furthermore, in the optimization process of the original three-plane data and the joint attention neural network model, the Gaussian sputtering 3D reconstruction obtained by the training set data is rendered as a final rendered image, and the parameters are optimized using a total loss function based on the difference between the final rendered image and the real image corresponding to the training set data;

[0019] The total loss function includes an image quality loss function and a mean absolute error between the final rendered image and the real image.

[0020] Furthermore, the total loss function is expressed as:

[0021] Loss = (1-λ)L1 + λ·L D-SSIM

[0022] Among them, Loss represents the total loss function; λ represents the weight coefficient, λ≤1, L D-SSIM represents the image quality loss function, and L1 is the mean absolute error.

[0023] Furthermore, in step 1, SENet is used as a channel attention network module to assign weights to the feature vectors of each position coordinate;

[0024] Swin-Transformer is used as the spatial attention network module to solve the position correlation between feature vectors of different position coordinates.

[0025] Furthermore, the so-called orthogonalization specifically includes: the three matrices in the three-plane data each form a gridded plane, and are arranged orthogonally in space to form a gridded square space. The orthogonalized three-plane data is represented as follows, that is, the three matrices in the three-plane data structure correspond to the gridded xy plane, yz plane and xz plane in the space, respectively, and the corresponding relationship remains consistent from the time of initialization.

[0026] Furthermore, the first and second attention feature vectors are the feature values ​​corresponding to the position coordinates in the grid space, that is, Tri_orth(x,y,z)=(ω xy ,ω yz ,ω zx ),ω xy ,ω yz ,ω zx They represent the eigenvectors corresponding to the projections of the position coordinates (x, y, z) on the xy plane, yz plane, and xz plane, respectively. The specific values ​​are obtained through double interpolation.

[0027] Furthermore, the process of calculating the position weight coefficient corresponding to the position coordinate is expressed as follows:

[0028] Tri1 = SE(Tri);

[0029] Tri 1 _orth=orth(Tri 1 );

[0030] F 1 =Tri 1 _orth(x,y,z);

[0031] Tri 2 =ST(Tri);

[0032] Tri 2 _orth=orth(Tri 2 );

[0033] F 2 =Tri 2 _orth(x,y,z);

[0034]

[0035] Among them, Tri represents the original three-plane data, SE represents the channel attention neural network in the channel attention network module, orth represents orthogonalization, and Tri 1 It represents the channel tri-plane data obtained by passing the original tri-plane data through the channel attention neural network. 1 _orth represents the three planes of the channel after orthogonalization, F 1 represents the first attention feature vector, ST represents the spatial attention neural network in the spatial attention network module, and Tri 2 It represents the spatial three-plane data obtained by passing the original three-plane data through the spatial attention neural network. 2 _orth represents the three planes of space after orthogonalization, F 2 represents the second attention feature vector, represents the vector concatenation operation, MLP represents the perception process of the multi-layer perceptron, and f θ (x,y,z) represents the position weight coefficient.

[0036] Furthermore, in step 4, the initial rotation coefficient, initial scaling coefficient and initial opacity of the 3D Gaussian point cloud are optimized by the position weight coefficient to obtain the corresponding optimized rotation coefficient, optimized scaling coefficient and optimized opacity. The specific formulas are as follows:

[0037] O′=O*f θ (x,y,z);

[0038] S′=S*f θ (x,y,z);

[0039] R′=R×f θ (x,y,z);

[0040] Among them, f θ (x, y, z) is the position weight coefficient, O represents the initial opacity, S represents the initial scaling factor, R represents the initial rotation factor, O′ represents the optimized opacity, S′ represents the optimized scaling factor, and R′ represents the optimized rotation factor.

[0041] Furthermore, in step 4, Gaussian adaptive density control is performed based on the position coordinates, optimized parameters and covariance matrix of the 3D Gaussian point cloud, and the 3D Gaussian point cloud obtained after density control is rendered to obtain a final rendered image.

[0042] The embodiments of the present application have the following beneficial effects:

[0043] The model and method proposed in the present invention combine the channel attention mechanism with the spatial attention mechanism, so that the model can discover and utilize the possible position correlation between different Gaussians, and optimize the different parameters of the three-dimensional Gaussian through position weights. This method achieved an average SSIM index of 0.852 and an average PSNR index of 24.02 in the scene in the Tanks&Temples dataset, improving the overall quality and local performance of the three-dimensional reconstructed image. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solution of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope of protection of the present invention. In each of the drawings, similar components are numbered similarly.

[0045] Figure 1 A schematic diagram of determining position weight coefficients based on a joint attention neural network model and position coordinates in a Gaussian sputtering 3D reconstruction method based on a joint attention mechanism proposed in an embodiment of the present application is shown.

[0046] Figure 2 A flow chart of a Gaussian sputtering 3D reconstruction method based on a joint attention mechanism proposed in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0047] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0048] In order to more clearly demonstrate the implementation steps and advantages of the invention, a specific implementation method is described below with reference to the accompanying drawings.

[0049] An embodiment of the present application proposes a Gaussian sputtering 3D reconstruction method based on a joint attention mechanism, which can discover and utilize possible positional correlations between a group of Gaussian point clouds, thereby improving the overall quality and local performance of the 3D reconstructed image.

[0050] The method in this example targets a set of 3D Gaussian point clouds constructed from the same 3D reconstruction scene. First, each 3D reconstruction scene contains multiple high-definition images taken from different angles and the camera parameters corresponding to each high-definition image. Based on this technology, a set of sparse SFM (Structure From Motion) point clouds is constructed, and then the constructed set of sparse SFM point clouds is initialized as a set of 3D Gaussian point clouds. Each Gaussian point cloud includes position coordinates, parameters to be optimized, and covariance matrix. The parameters to be optimized include several main parameters such as rotation coefficient, scaling coefficient, opacity, etc.

[0051] Typically, the specific implementation steps of this method are as follows:

[0052] First, establish Figure 1 The joint attention neural network model shown in the figure has two parts as its input, one of which is a preset three-plane data structure, which includes a trainable original three-plane data. The three-plane data structure is mainly used to store feature vectors representing Gaussian position coordinates. The feature vector is a trainable parameter, and its structure is three matrices of the same size (H, W, C), where H represents height, W represents width, and C represents depth, or the number of channels, which are generally obtained randomly in the initialization stage. For example, the trainable three-plane data is represented by a tensor Tri, and its size can be (3, 512, 24, 24), where 3 represents 3 planes, 24 represents width, 24 represents height, and 512 represents depth, or the number of channels.

[0053] like Figure 1As shown, in the model, the original three-plane data are respectively input into the channel attention network module for internal self-attention calculation and the spatial attention network module for spatial attention calculation. In this example, a joint attention neural network model is built based on SENet (Squeeze-and-Excitation Networks) and Swin-Transformer. Among them, SENet is a convolutional neural network structure for deep learning, which optimizes feature extraction by introducing a channel attention mechanism. In other words, SENet in this example is a channel attention network module, which is used to give weights to the feature vectors of each position coordinate, that is, for internal self-attention calculation, which is equivalent to the self-attention mechanism module of the feature. Such methods are well known to those skilled in the art and will not be described here. Swin-Transformer is a method for processing computer vision tasks using a transformer architecture. In this example, the Swin-Transformer network is used as a spatial attention network module to solve the position correlation between feature vectors of different position coordinates, so that the feature vectors of each position coordinate have global correlation, that is, for spatial attention calculation. Such methods are well known to those skilled in the art and will not be described here.

[0054] The three-plane data after the channel attention network module and the spatial attention network module are orthogonalized to obtain the channel three-plane and the spatial three-plane; the so-called orthogonalization is basically as follows Figure 1 As shown in , the three matrices in the three-plane data each form a gridded plane and are arranged orthogonally in space to form a gridded square space. The orthogonalized three-plane data is expressed as Tri_orth=(Tri xy ,Tri yz ,Tri xz ); it can be understood that the three matrices in the three-plane data structure correspond to the gridded xy plane, yz plane and xz plane in space, and the corresponding relationship is always the same at the time of initialization. Orthogonalization transforms a linearly independent vector system into an orthogonal system (that is, the vectors are orthogonal to each other, that is, their inner product is zero). Its purpose is to simplify the processing and analysis of complex problems by generating a new set of orthogonal vectors. These vectors maintain the original linear correlation, but their characteristics make the problem easier to solve.

[0055] At this time, the model introduces another part of the input, that is, the position coordinates contained in each Gaussian point cloud data in the aforementioned set of 3D Gaussian point clouds. By projecting the position coordinates onto the channel three planes and the spatial three planes respectively, the first attention feature vector corresponding to the channel three planes and the second attention feature vector corresponding to the spatial three planes are obtained. Specifically, the first and second attention feature vectors are the eigenvalues ​​corresponding to the position coordinates in the grid space, that is, Tri_orth(x,y,z)=(ω xy ,ω yz ,ω zx ),ω xy ,ω yz ,ω zx They represent the eigenvectors corresponding to the projections of the position coordinates (x, y, z) on the xy plane, yz plane, and xz plane, respectively. The specific values ​​are obtained through double interpolation.

[0056] The first attention feature vector and the second attention feature vector are concatenated. In this example, the concatenation is performed in the depth dimension of the attention feature vector. The concatenated feature vector is sent to a multi-layer perceptron (MLP) for perception to obtain the position weight coefficient corresponding to the position coordinate.

[0057] Exemplarily, the process of calculating the position weight coefficient corresponding to the position coordinates is shown as follows:

[0058] Tri 1 = SE(Tri);

[0059] Tri 1 _orth=orth(Tri 1 );

[0060] F 1 =Tri 1 _orth(x,y,z);

[0061] Tri 2 =ST(Tri);

[0062] Tri 2 _orth=orth(Tri 2 );

[0063] F 2 =Tri 2 _orth(x,y,z);

[0064]

[0065] Among them, Tri represents the original three-plane data, SE represents the channel attention neural network in the channel attention network module, orth represents orthogonalization, and Tri 1It represents the channel tri-plane data obtained by passing the original tri-plane data through the channel attention neural network. 1 _orth represents the three planes of the channel after orthogonalization, F 1 represents the first attention feature vector, ST represents the spatial attention neural network in the spatial attention network module, and Tri 2 It represents the spatial three-plane data obtained by passing the original three-plane data through the spatial attention neural network. 2 _orth represents the three planes of space after orthogonalization, F 2 represents the second attention feature vector, represents the vector concatenation operation, MLP represents the perception process of the multi-layer perceptron, and f θ (x,y,z) represents the position weight coefficient.

[0066] Then, if Figure 2 As shown, the method designed in this example is to optimize the parameters to be optimized through the position weight coefficient to obtain the optimized parameters; Gaussian adaptive density control is performed based on the position coordinates, the optimized parameters and the covariance matrix, and the three-dimensional Gaussian obtained after density control is rendered to obtain the final rendered image. Among them, the parameters to be optimized include the initial rotation coefficient, the initial scaling coefficient and the initial opacity.

[0067] Specifically, the initial rotation coefficient, initial scaling coefficient and initial opacity are optimized respectively through the position weight coefficient of the Gaussian point cloud to obtain the corresponding optimized rotation coefficient, optimized scaling coefficient and optimized opacity respectively.

[0068] The specific formula for the above optimization process is as follows:

[0069] O′O*f θ (x,y,z);

[0070] S′=S*f θ (x,y,z);

[0071] R′=R×f θ (x,y,z);

[0072] Among them, f θ (x, y, z) is the position weight coefficient, O represents the initial opacity, S represents the initial scaling factor, R represents the initial rotation factor, O′ represents the optimized opacity, S′ represents the optimized scaling factor, and R′ represents the optimized rotation factor.

[0073] Gaussian adaptive density control is performed based on the position coordinates, covariance matrix, optimized rotation coefficient, optimized scaling coefficient, and optimized opacity of the Gaussian point cloud, so that the total number of 3D Gaussian point clouds is guaranteed to be within a reasonable range. Among them, Gaussian adaptive density control is a technology that is usually introduced when processing 3D Gaussian Splatting or related 3D scene representation methods. Adaptive density control is a mechanism that dynamically adjusts the Gaussian density according to the geometric complexity of the scene, which helps to improve the compactness and efficiency of the 3D scene representation, and plays an important role in various 3D rendering technologies. Such methods are well known to those skilled in the art and will not be described in detail here.

[0074] The 3D Gaussian point cloud obtained after density control is sent to a renderer for rendering to obtain a final rendered image. Demonstratively, the 3D Gaussian point cloud can be rendered using fast differentiable rasterization to obtain a final rendered image.

[0075] For the training of the aforementioned joint attention neural network model, in this example, preferably but not limited to, parameter optimization is performed based on a loss function that characterizes the difference between the final rendered image and the real image to obtain an optimized joint attention neural network model.

[0076] In the exemplary training of this example, in each training, the final rendered image is obtained according to the method shown, and the weighted addition of the mean absolute error (L1 Loss) and the structural similarity index (SSIM) loss function between the final rendered image and the real image corresponding to the final rendered image is used as the total loss function Loss to optimize the joint attention neural network model. The joint attention neural network model is trained according to a pre-set number of training times until the joint attention neural network model converges to obtain an optimized joint attention neural network model.

[0077] Among them, the total loss function Loss formula of the model is:

[0078] Loss = (1-λ)L1 + λ·L D-SSIM

[0079] Among them, Loss represents the total loss function; λ represents the weight coefficient, λ≤1, L D-SSIM represents the image quality loss function, and L1 is the mean absolute error.

[0080] After obtaining the optimized joint attention neural network model, the test set in the 3D reconstruction data is input into the optimized joint attention neural network model to obtain the corresponding 3D reconstructed image.

[0081] The 3D reconstruction dataset used for training in this embodiment is taken from the Tanks & Temples dataset, which is divided into a training set and a test set. The Tanks & Temples dataset is a large multi-view dataset for 3D reconstruction research, released in 2017 by Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun of Intel Labs. Its training set contains 7 high-resolution videos of 7 scenes. The test set is divided into an intermediate group and an advanced group. Among them, the intermediate group contains scenes such as sculptures, large vehicles, and house buildings with exterior camera trajectories. These scenes are relatively simple and suitable for preliminary 3D reconstruction tests; the advanced group contains indoor scenes shot from inside and large outdoor scenes. These scenes have complex geometric layouts and camera trajectories, which put forward higher requirements for 3D reconstruction algorithms. This dataset has a wide range of applications and influence in the field of 3D reconstruction, and is used as a benchmark dataset by multiple research projects and papers.

[0082] The performance of the proposed method in this example on the Tanks&Temples dataset is compared with the current state-of-the-art results as shown in the following table:

[0083]

[0084]

[0085] Among them, PSNR (Peak signal-to-noise ratio) is an engineering term that represents the ratio of the maximum possible power of a signal to the destructive noise power that affects its representation accuracy. Since many signals have a very wide dynamic range, the peak signal-to-noise ratio is often expressed in logarithmic decibel units. The peak signal-to-noise ratio is often used as a measure of signal reconstruction quality in fields such as image compression, and it is often simply defined by the mean square error (MSE).

[0086] The model and method in this example combine the channel attention mechanism and the spatial attention mechanism, enabling the model to discover and utilize possible positional correlations between different Gaussians, and optimize different parameters of the three-dimensional Gaussian through position weights. In the scenarios in the Tanks&Temples dataset, an average SSIM index of 0.852 and an average PSNR index of 24.02 were achieved, improving the overall image quality and local detail performance of the three-dimensional reconstructed image.

[0087] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A Gaussian sputtering 3D reconstruction method based on joint attention mechanism, characterized in that: The following steps are involved: Step 1: Construct the joint attention neural network model for obtaining position weight coefficients according to the following strategy: The input of the joint attention neural network model includes two parts, one of which is a pre-set three-plane data structure, which includes a trainable original three-plane data. The three-plane data structure is mainly used to store a feature vector representing the Gaussian position coordinates, and the feature vector is a trainable parameter; The original three-plane data are respectively input into the channel attention network module for internal self-attention calculation and the spatial attention network module for spatial attention calculation; The three-plane data after the channel attention network module and the spatial attention network module are orthogonalized to obtain the channel three-plane and the spatial three-plane; Another part of the model input is the position coordinates of each Gaussian point cloud data in a set of 3D Gaussian point clouds constructed by the same 3D reconstruction scene; by projecting the position coordinates onto the channel three planes and the spatial three planes respectively, the first attention feature vector corresponding to the channel three planes and the second attention feature vector corresponding to the spatial three planes are obtained; The first attention feature vector and the second attention feature vector are concatenated, and the concatenated feature vector is sent to a multi-layer perceptron MLP for perception to obtain a position weight coefficient corresponding to the position coordinate; Step 2: Input the position coordinates contained in the 3D Gaussian point cloud to be reconstructed and the optimized original three-plane data into the optimized joint attention neural network model to obtain the corresponding position weight coefficient; Step 3: Optimizing the parameters to be optimized for Gaussian sputtering three-dimensional reconstruction contained in the reconstructed 3D Gaussian point cloud using the position weight coefficient to obtain optimized parameters; Step 4 uses the 3D Gaussian point cloud with optimized parameters to be optimized to perform Gaussian sputtering 3D reconstruction.

2. The method according to claim 1, characterized in that The 3D reconstructed scene includes multiple high-definition images taken from different angles and camera parameters corresponding to each high-definition image, which can construct a set of sparse SFM point clouds and are initialized as a set of 3D Gaussian point clouds; Each Gaussian point cloud includes position coordinates, parameters to be optimized and a covariance matrix. The parameters to be optimized include rotation coefficients, scaling coefficients and opacity.

3. The method according to claim 1, characterized in that In the optimization process of the original three-plane data and the joint attention neural network model, the Gaussian sputtering 3D reconstruction obtained by the training set data is rendered as the final rendered image, and the parameters are optimized using the total loss function based on the difference between the final rendered image and the real image corresponding to the training set data; The total loss function includes an image quality loss function and a mean absolute error between the final rendered image and the real image.

4. The method according to claim 1, characterized in that: The total loss function is expressed as: Loss=(1-λ)L1+λ·L D-SSIM Among them, Loss represents the total loss function; λ represents the weight coefficient, λ≤1, L D-SSIM represents the image quality loss function, and L1 is the mean absolute error.

5. The method according to claim 1, characterized in that In step 1, SENet is used as the channel attention network module to assign weights to the feature vectors of each position coordinate; Swin-Transformer is used as the spatial attention network module to solve the position correlation between feature vectors of different position coordinates.

6. The method according to claim 5, characterized in that The so-called orthogonalization specifically includes: the three matrices in the three-plane data each form a gridded plane, and are arranged orthogonally in space to form a gridded square space. The orthogonalized three-plane data is represented as, that is, the three matrices in the three-plane data structure correspond to the gridded xy plane, yz plane and xz plane in the space, and the corresponding relationship remains consistent from the time of initialization.

7. The method according to claim 6, characterized in that The first and second attention feature vectors are the feature values ​​corresponding to the position coordinates in the grid space, that is, Tri_orth(x,y,z)=(ω xy ,ω yz ,ω zx ),ω xy ,ω yz ,ω zx They represent the eigenvectors corresponding to the projections of the position coordinates (x, y, z) on the xy plane, yz plane, and xz plane, respectively. The specific values ​​are obtained through double interpolation.

8. The method according to claim 7, characterized in that The process of calculating the position weight coefficient corresponding to the position coordinate is expressed as follows: Tri1=SE(Tri); Tri1_orth=orth(Tri1); F1 = Tri1_orth(x,y,z); Tri2 = ST (Tri); Tri2_orth = orth(Tri2); F2 = Tri2_orth(x,y,z); Among them, Tri represents the original three-plane data, SE represents the channel attention neural network in the channel attention network module, orth represents orthogonalization, Tri1 represents the channel three-plane data obtained by passing the original three-plane data through the channel attention neural network, Tri1_orth represents the channel three-plane after orthogonalization, F1 represents the first attention feature vector, ST represents the spatial attention neural network in the spatial attention network module, Tri2 represents the spatial three-plane data obtained by passing the original three-plane data through the spatial attention neural network, Tri2_orth represents the spatial three-plane after orthogonalization, and F2 represents the second attention feature vector. represents the vector concatenation operation, MLP represents the perception process of the multi-layer perceptron, and f θ (x,y,z) represents the position weight coefficient.

9. The method according to claim 2, characterized in that: In step 4, the initial rotation coefficient, initial scaling coefficient and initial opacity of the 3D Gaussian point cloud are optimized by the position weight coefficient to obtain the corresponding optimized rotation coefficient, optimized scaling coefficient and optimized opacity. The specific formulas are as follows: O′=O*f θ (x,y,z); S′=S*f θ (x,y,z); R′=R×f θ (x,y,z); Among them, f θ (x, y, z) is the position weight coefficient, O represents the initial opacity, S represents the initial scaling factor, R represents the initial rotation factor, O′ represents the optimized opacity, S′ represents the optimized scaling factor, and R′ represents the optimized rotation factor.

10. The method according to claim 9, characterized in that In step 4, Gaussian adaptive density control is performed based on the position coordinates, optimized parameters and covariance matrix of the 3D Gaussian point cloud, and the 3D Gaussian point cloud obtained after density control is rendered to obtain a final rendered image.

Citation Information

Patent Citations

  • Three-dimensional reconstruction method and system based on implicit function

    CN117095132A

  • Automobile part point cloud multi-scale segmentation method based on multi-layer attention mechanism

    CN118429650A

  • Model training method, scene reconstruction method, device, equipment, medium and product

    CN118506322A

  • Scene reconstruction method and device, electronic equipment, storage medium and product

    CN118840478A

  • Adaptive Surface Splatting Method and Apparatus for 3 Dimension Rendering

    KR101284446B1

Cited By

  • Single composite image shadow generation method based on 3D perception

    CN121353508A