Three-dimensional Gaussian splashing method and device

Through multi-stage training optimization and compression processing, the target compressed data is generated and decoded and rendered, and the problem of large storage space occupied by three-dimensional Gaussian splattering technology is solved, achieving high fidelity and quality rendered images.

CN120070748APending Publication Date: 2025-05-30GREATER BAY AREA UNIV (IN PREPARATION)
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510122118.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The three-dimensional Gaussian splashing technology requires millions of neural Gauss, which makes the storage space consumed large amounts. The existing compression technology cannot effectively compress and ensure rendering fidelity.

Method used

By obtaining the target point cloud and performing initialization processing, the initial anchor point and its characteristic information are obtained, and multi-stage training and optimization are carried out to obtain the optimal anchor point and its characteristic information. This information is used for compression processing, the target compressed data is generated, and the rendered image is obtained through decoding and rendering processing.

Benefits of technology

Reduces the storage cost of 3D Gaussian splattering technology and improves the fidelity and quality of compressed rendered images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070748A_ABST
    Figure CN120070748A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional Gaussian splashing method and device which are applied to the technical field of three-dimensional reconstruction, and the method comprises the steps: obtaining a target point cloud and carrying out initialization processing, obtaining a plurality of initial anchor points and characteristic information thereof, the characteristic information comprising anchor point attributes and three-dimensional coordinates, and the anchor point attributes comprising offset coordinates, attribute characteristics and scale scaling factors; performing multi-stage training optimization according to the plurality of initial anchor points and the characteristic information thereof to obtain a target training result, a plurality of optimal anchor points and the characteristic information of each optimal anchor point, the target training result comprising a target upsampling module, a target prediction module, a compressed three-plane structure and a target mask; performing compression processing according to the target training result, the plurality of optimal anchor points and the characteristic information of each optimal anchor point to obtain target compressed data; and decoding and rendering the target compressed data to obtain a target rendered image. According to the invention, the storage cost of the three-dimensional Gaussian splashing technology can be reduced, and the rendering fidelity and the rendering quality can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of 3D reconstruction technology, and in particular to a 3D Gaussian splatting method and device. Background Art

[0002] Thanks to the Neural Radiance Field (NeRF) technology, the novel view synthesis task in 3D scene representation has entered a new era. The NeRF technology and its variants use a Multi-Layer Perceptron (MLP) to map coordinates and view directions to colors and densities, and through Volume Rendering (VP), they show highly realistic rendering results in the novel view synthesis task. Although the NeRF technology effectively improves the rendering quality and fidelity of 3D scenes, its implicit expression is not intuitive, with low controllability, and its training efficiency is low, making it inapplicable to high-resolution real-time rendering.

[0003] Recently, the 3D Gaussian Splatting (3DGS) technology has been proposed as an effective 3D scene representation technology, which fundamentally solves the pain points of the NeRF technology. The 3DGS technology uses a set of neural Gaussians generated by the Structure from Motion (SfM) technology to represent 3D scenes. These neural Gaussians have learnable attributes, shape, and appearance parameters, such as color, shape, and opacity, etc., and these neural Gaussians can be splatted onto a 2D plane and rasterized to achieve parallel, fast, and differentiable rendering.

[0004] Although the 3DGS technology has excellent rendering quality and real-time rendering speed, due to the representation of large scenes requiring millions of neural Gaussians, the explicit expression structure of the 3DGS technology usually contains millions of neural Gaussians, which means that the 3DGS technology requires a large amount of storage space. In response, related technologies propose to use existing compression technologies such as hash grids to compress the rendered images of the 3DGS technology, making it possible to be used in more lightweight and low-configuration scenarios. However, limited by the characteristics of neural Gaussians, the existing compression technologies are not applicable to the 3DGS technology, and problems such as parameter loss may occur during compression, unable to ensure subsequent rendering fidelity and quality. Summary of the Invention

[0005] Embodiments of this application provide a 3D Gaussian splatting method and device, which are used to reduce the storage cost of the 3DGS technology and improve the fidelity and quality of the compressed rendered images.

[0006] On the one hand, an embodiment of the present application provides a three-dimensional Gaussian splashing method, including the following steps: Obtain a target point cloud and perform initialization processing to obtain a number of initial anchor points and characteristic information of each of the initial anchor points; wherein, the characteristic information includes anchor point attributes and three-dimensional coordinates, and the anchor point attributes include offset coordinates, attribute features, and scale factors; Perform multi-stage training optimization based on the number of initial anchor points and the characteristic information of each initial anchor point to obtain a target training result, multiple optimal anchor points, and characteristic information of each optimal anchor point; wherein, the target training result includes a target upsampling module, a target prediction module, a compressed triplane structure, and a target mask; Perform compression processing based on the target training result, the multiple optimal anchor points, and the characteristic information of each optimal anchor point to obtain target compressed data; Perform decoding and rendering processing on the target compressed data to obtain a target rendered image.

[0007] On the other hand, an embodiment of the present application provides a three-dimensional Gaussian splashing device, including: A first processing module, configured to obtain a target point cloud and perform initialization processing to obtain a number of initial anchor points and characteristic information of each of the initial anchor points; wherein, the characteristic information includes anchor point attributes and three-dimensional coordinates, and the anchor point attributes include offset coordinates, attribute features, and scale factors; A second processing module, configured to perform multi-stage training optimization based on the number of initial anchor points and the characteristic information of each initial anchor point to obtain a target training result, multiple optimal anchor points, and characteristic information of each optimal anchor point; wherein, the target training result includes a target upsampling module, a target prediction module, a compressed triplane structure, and a target mask; A third processing module, configured to perform compression processing based on the target training result, the multiple optimal anchor points, and the characteristic information of each optimal anchor point to obtain target compressed data; A fourth processing module, configured to perform decoding and rendering processing on the target compressed data to obtain a target rendered image.

[0008] The beneficial effects of this application are as follows: A three-dimensional Gaussian splashing method and device are provided. First, the target point cloud is obtained and initialized to obtain a number of initial anchor points and their characteristic information. Among them, the characteristic information includes the anchor point attributes and three-dimensional coordinates, and the anchor point attributes include offset coordinates, attribute features, and scale factors. Then, multi-stage training optimization is performed according to the number of initial anchor points and their characteristic information to obtain the target training results, multiple optimal anchor points, and the characteristic information of each optimal anchor point. Among them, the target training results include a target upsampling module, a target prediction module, a compressed triplanar structure, and a target mask. After that, compression processing is performed according to the target training results, multiple optimal anchor points, and the characteristic information of each optimal anchor point to obtain the target compressed data. Finally, decoding and rendering processing are performed on the target compressed data to obtain the target rendered image. This application can reduce the storage cost of three-dimensional Gaussian splashing technology and improve the rendering fidelity and rendering quality. Description of the Drawings

[0009] Figure 1 is a flowchart of a three-dimensional Gaussian splashing method provided by this application; Figure 2 is a schematic diagram of a three-dimensional Gaussian splashing method provided by this application; Figure 3 is a schematic comparison diagram of the principles of a three-dimensional Gaussian splashing method provided by this application and two baseline methods; Figure 4 is a qualitative comparison diagram of a three-dimensional Gaussian splashing method provided by this application and two baseline methods in terms of quality and storage space; Figure 5 is a qualitative comparison diagram of a three-dimensional Gaussian splashing method provided by this application and one of the baseline methods in terms of rendering details. Detailed Embodiments

[0010] In order to make the objectives, technical solutions, and advantages of this application clearer, the following further elaborates on this application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0011] The following further describes this application in combination with the drawings of the specification and specific embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0012] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.

[0014] In view of the deficiencies in the related art, embodiments of this application provide a three-dimensional Gaussian splashing method and apparatus, aiming to reduce the storage cost of three-dimensional Gaussian splashing technology and improve the fidelity and quality of the compressed rendered image.

[0015] First, the specific implementation steps of a three-dimensional Gaussian splashing method provided by this application will be elaborated in detail below with reference to the accompanying drawings.

[0016] A three-dimensional Gaussian splashing method provided by this application can be applied to a terminal, a server, or software running on a terminal or a server. The terminal can be a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. In addition, the server can also be a node server in a blockchain network, but is not limited thereto. Among them, the blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms.

[0017] Referring to Figure 1 , the three-dimensional Gaussian splashing method provided by this application may include the following steps S100-S400.

[0018] S100, obtain a target point cloud and perform initialization processing to obtain a plurality of initial anchor points and the characteristic information of each initial anchor point.

[0019] It should be noted that the characteristic information may include, but is not limited to, anchor attributes and three-dimensional coordinates. Specifically, the three-dimensional coordinates refer to the coordinates of the anchor in the Normalized Device Coordinates (NDC) space obtained after being processed by the motion inference structure technology, which may include the horizontal axis coordinate, the vertical axis coordinate, and the depth axis coordinate. In addition, the anchor attributes can map the characteristics of the neural Gaussian associated with the anchor, which follows the concept in Scaffold-GS. Briefly, the anchor attributes may include offset coordinates, attribute features, and scale factors. Among them, the attribute features are the local context features in Scaffold-GS, and the offset coordinates refer to the offset coordinates of the anchor in the neural Gaussian space, which can be operated with the three-dimensional coordinates of the anchor itself to obtain the spatial coordinates of the neural Gaussian derived from the anchor.

[0020] It can be understood that the target point cloud refers to the point cloud data estimated from at least one target image using the motion inference structure technology, and the target image is the three-dimensional reconstruction object.

[0021] In this step, first, the point cloud data is estimated from at least one target image using the motion inference structure technology, and the point cloud data is determined as the target point cloud. Then, the target point cloud is initialized. After that, the anchor points and the neural Gaussian distribution are generated through the initialized target point cloud, so as to obtain a number of initial anchor points and their anchor attributes and three-dimensional coordinates. This step belongs to the prior art and will not be elaborated here. It should be noted that each anchor point is associated with a group of neural Gaussians, and each anchor point is equipped with corresponding anchor attributes, and the anchor attributes can reflect the characteristics of the neural Gaussian associated with the anchor point. For the sake of understanding, the anchor attributes of the anchor point can be expressed as , where represents the offset coordinates of the anchor point, represents the attribute features of the anchor point, represents the scale factor of the anchor point, and the three-dimensional coordinates of the anchor point can be expressed as .

[0022] Optionally, the number of initial anchor points can be set according to the actual situation, and this embodiment does not make specific limitations on this.

[0023] S200. Perform multi-stage training optimization according to a number of initial anchor points and the characteristic information of each initial anchor point to obtain the target training result, multiple optimal anchor points, and the characteristic information of each optimal anchor point.

[0024] It should be noted that the target training result refers to the data obtained through multi-stage training, which may include a target upsampling module, a target prediction module, a compressed tri-plane structure, and a target mask. Among them, the target mask is used to screen multiple optimal anchor points to obtain the final anchor points for compression; the compressed tri-plane structure refers to the tri-plane structure after compression; the target upsampling module is used to restore the compressed tri-plane structure into a tri-plane structure available for anchor point sampling; the target prediction module is used to predict the attribute distribution and quantization step size of the anchor points based on the sampling results and three-dimensional coordinates of the anchor points. The quantization step size is used to perform quantization processing on the anchor point attributes to ensure the stability of the anchor point attributes during the compression process, while the attribute distribution is used to combine with the anchor point attributes to jointly achieve the compression processing of the three-dimensional Gaussian splash parameters.

[0025] It is worth noting that the tri-plane structure is a continuous tri-plane data structure proposed in this application. It is used to project and sample anchor points to obtain the coordinate features of the anchor points on different two-dimensional planes. These coordinate features can generally be referred to as tri-plane features, and the tri-plane features are the sampling results of the anchor points. They are used to combine with the three-dimensional coordinates to predict the attribute distribution and quantization step size of the anchor points. Specifically, the tri-plane structure includes three two-dimensional planes: the first plane, the second plane, and the third plane. As Figure 2 shown, the first plane is a two-dimensional plane composed of the horizontal axis and the vertical axis, denoted as , the second plane is a two-dimensional plane composed of the vertical axis and the depth axis, denoted as , and the third plane is a two-dimensional plane composed of the horizontal axis and the depth axis, denoted as . These three two-dimensional planes together represent the three-dimensional space, that is, the tri-plane structure.

[0026] It can be understood that since the optimal anchor points need to be further screened during the compression stage, the optimal anchor points refer to the anchor points initially used for compression, rather than the final anchor points used for compression. In addition, the characteristic information of the optimal anchor points follows the concept of the characteristic information in the previous steps and will not be elaborated here.

[0027] In this step, during the multi-stage training and optimization stage, iterative training and optimization are carried out using these initial anchor points and their characteristic information, aiming to screen out the anchor points suitable for compression processing from several initial anchor points and continuously optimize their characteristic information. At the same time, data such as the upsampling module, prediction module, tri-plane structure, and mask are updated. These data are crucial for compression processing. After multi-stage training, the target training result, multiple optimal anchor points, and the characteristic information of each optimal anchor point can be obtained. The target training result is the target upsampling module, target prediction module, compressed tri-plane structure, and target mask, so as to achieve the compression processing of the three-dimensional Gaussian splash parameters in the subsequent steps.

[0028] S300. Perform compression processing based on the target training result, multiple optimal anchor points, and the characteristic information of each optimal anchor point to obtain target compressed data.

[0029] It should be noted that the target compressed data refers to the compression result of three-dimensional Gaussian splash parameters, and the three-dimensional Gaussian splash parameters are the parameters of the three-dimensional Gaussian splash technology. It is worth noting that in the embodiments of the present application, the three-dimensional Gaussian splash parameters refer to the anchor attributes of the anchor points.

[0030] In this step, in the compression stage, compression processing is performed based on the target training result, multiple optimal anchor points, and the characteristic information of each optimal anchor point to obtain target compressed data. In this way, targeted compression processing of the three-dimensional Gaussian splash parameters is achieved while fully considering the characteristics of neural Gaussian, effectively reducing the number of parameters of the three-dimensional Gaussian splash parameters.

[0031] S400. Perform decoding and rendering processing on the target compressed data to obtain a target rendered image.

[0032] In this step, in the rendering stage, the target compressed data is decoded by a preset decoder to obtain a decoding result, and then the decoding result is subjected to differentiable rasterization rendering to obtain a rendered image associated with the target point cloud, that is, the rendered image of the target point cloud, thereby realizing three-dimensional Gaussian splash.

[0033] In some embodiments, the specific implementation process of the above step S200 may include the following steps S210 - S230.

[0034] S210. Perform the first-stage training optimization processing based on a preset mask, several initial anchor points, and the characteristic information of each initial anchor point to obtain a first training result.

[0035] It should be noted that the first training result may include a first mask, multiple first target anchor points, and the characteristic information of each first target anchor point. Among them, the first target anchor point refers to the anchor point obtained after the first-stage training. The number of the first target masks is less than the number of the initial anchor points. The concept of the characteristic information of the first target anchor point follows the concept of the characteristic information in the foregoing embodiments and will not be elaborated here. It is worth noting that the mask is a learnable mask used to perform a masking operation on the anchor points. The preset mask refers to an untrained mask, and its masking object is the initial anchor point, while the first mask refers to the mask after the first-stage training, and its masking object is the first target anchor point.

[0036] In this step, the first-stage training is the initial training optimization stage. In each round of training optimization in the first stage, first, a preset mask is used to screen a number of initial anchor points to obtain the anchor points participating in the training. Then, the screened anchor points are rendered to obtain a preliminarily rendered image, and the loss function value is calculated using the preliminarily rendered image. After that, the preset mask, as well as the anchor point attributes and three-dimensional coordinates of the screened anchor points, are updated to optimize the mask and three-dimensional Gaussian splash parameters and ensure the accuracy of subsequent three-plane sampling. Finally, the next round of training is entered to achieve iterative optimization. After multiple rounds of training in the first stage, a first mask, multiple first target anchor points, and the characteristic information of each first target anchor point can be obtained as the first training result.

[0037] Optionally, the first-stage training refers to the 1st round of training to the th round of training in the multi-stage training optimization, where is a preset first threshold, which can be flexibly set according to the actual situation. For example, the first threshold can be 10,000, so the 1st to 10,000th rounds of training in the multi-stage training optimization are the first-stage training.

[0038] S220. Perform the second-stage training optimization process according to the preset three-plane structure, preset prediction module, and the first training result to obtain the second training result.

[0039] It should be noted that the second training result may include an initial three-plane structure, an initial prediction module, a second mask, multiple second target anchor points, and the characteristic information of each of the second target anchor points. Among them, the second target anchor points refer to the anchor points obtained after the second-stage training. The number of second target anchor points is less than the number of first target anchor points. The concept of the characteristic information of the second target anchor points follows the concept of the characteristic information in the foregoing embodiments and will not be elaborated here. In addition, the second mask refers to the mask after the second-stage training, and its masked object is the second target anchor points.

[0040] It is worth noting that the preset three-plane structure refers to the three-plane structure that has not been trained, while the initial three-plane structure refers to the three-plane structure after the second-stage training. Among them, the three-plane structure follows the concept of the three-plane structure in the foregoing embodiments and will not be elaborated here. In addition, the preset prediction module refers to the prediction module that has not been trained, while the initial prediction module refers to the prediction module after the second-stage training. It should be understood that the type of the prediction module can be flexibly set according to the actual situation. For example, the prediction module is a multi-layer perceptron, but other prediction modules such as convolutional neural networks can also be used. This embodiment does not make specific limitations in this regard.

[0041] In this step, the second-stage training is the intermediate training optimization stage. To improve the compression efficiency of the three-dimensional Gaussian splash parameters, reduce the possibility of parameter loss during compression, and enhance the rendering quality and rendering fidelity, a three-plane structure and a prediction module are introduced for the training optimization in the second stage based on the mask and the anchor points. Specifically, in each round of training optimization in the second stage, first, the first mask is used to screen a number of first target anchor points, aiming to select the anchor points participating in the training. Secondly, the preset three-plane structure is used to perform projection sampling on the selected anchor points, and the three-dimensional coordinates of the anchor points and the results of the projection sampling are input into the preset prediction module for prediction processing. Then, the prediction results are used to quantify the anchor point attributes of the selected anchor points and perform rendering to obtain a preliminarily rendered image, and the loss function value is calculated using the prediction results, the quantified anchor point attributes, and the preliminarily rendered image. After that, the loss function value is used to update the learnable parameters of the first mask, the preset three-plane structure, the preset prediction module, and the characteristic information of the selected anchor points, aiming to optimize the mask, the three-plane structure, the prediction module, and the three-dimensional Gaussian splash parameters, and ensure the accuracy of subsequent three-plane sampling. Finally, it enters the next round of training to achieve iterative optimization. After multiple rounds of training in the second stage, an initial three-plane structure, an initial prediction module, a second mask, multiple second target anchor points, and the characteristic information of each of the second target anchor points can be obtained as the second training result.

[0042] Optionally, the training in the second stage refers to the round of training to the round of training in the multi-stage training optimization. is a preset second threshold, which can be flexibly set according to the actual situation. For example, the first threshold can be 10,000 times, and the second threshold can be 25,000 times. Then, the 10,001st to 25,000th rounds of training in the multi-stage training optimization are the second-stage training.

[0043] S230. Perform the third-stage training optimization process according to the preset compression module and the second training result to obtain the target training result, multiple optimal anchor points, and the characteristic information of each optimal anchor point.

[0044] It should be noted that the preset compression module may include a preset upsampling module and a preset downsampling module. The preset downsampling module refers to a downsampling module that has not been trained and is used to perform a compression operation on the three-plane structure. The preset upsampling module refers to an upsampling module that has not been trained and is used to perform a restoration operation on the compressed three-plane structure. It should be understood that both the preset downsampling module and the preset upsampling module can be flexibly set according to the actual situation. For example, the preset upsampling module can be a deconvolutional neural network (DN), and the preset downsampling module can be a convolutional neural network (CNN), but it is not limited to this.

[0045] It can be understood that the training processing results of the third stage are the target upsampling module, the target prediction module, the compressed three-plane structure, the target mask, multiple optimal anchor points, and the characteristic information of each optimal anchor point. Among them, the optimal anchor point refers to the anchor point obtained after the third stage of training. The number of optimal anchor points is less than the number of second target anchor points. The target upsampling module refers to the upsampling module after the third stage of training. The target prediction module refers to the prediction module after the third stage of training. The compressed three-plane structure refers to the three-plane structure compressed in the third stage of training. The target mask refers to the mask after the third stage of training, and its masking object is the optimal anchor point.

[0046] In this step, the third-stage training is the final training optimization stage. Considering that in subsequent compression processing, three-plane projection sampling needs to be performed on the anchor points, and the parameters of the three-plane structure are large, which is not conducive to the compression processing of the three-dimensional Gaussian splash parameters. Therefore, on the basis of the mask, anchor points, three-plane structure, and prediction module, a compression module is introduced for the training optimization of the third stage. The compression module is used to perform compression operations and restoration operations on the three-plane structure. Specifically, in each round of training optimization in the third stage, first, a second mask is used to screen a number of second target anchor points, aiming to screen out the anchor points participating in the training. Secondly, the initial three-plane structure is used to perform projection sampling on the screened anchor points, and the three-dimensional coordinates of the anchor points and the results of the projection sampling are input into the initial prediction module for prediction processing. Then, the prediction results are used to quantify the anchor point attributes of the screened anchor points and perform rendering to obtain a preliminarily rendered image. At the same time, the initial three-plane structure is downsampled and upsampled through a preset compression module to obtain a compressed initial three-plane structure and a restored initial three-plane structure. After that, the loss function value is calculated using the prediction results, quantified anchor point attributes, preliminarily rendered image, initial three-plane structure, and restored initial three-plane structure, and the loss function value is used to update the learnable parameters of the second mask, initial three-plane structure, initial prediction module, learnable parameters of the preset compression module, and the characteristic information of the screened anchor points, aiming to optimize the mask, three-plane structure, prediction module, compression module, and three-dimensional Gaussian splash parameters, and ensure the accuracy of three-plane sampling. Finally, it enters the next round of training to achieve iterative optimization. After multiple rounds of training in the third stage, a target upsampling module, a target prediction module, a compressed three-plane structure, a target mask, multiple optimal anchor points, and the characteristic information of each optimal anchor point can be obtained as the target training results.

[0047] Optionally, the training in the third stage refers to the th round to the th round of training in the multi-stage training optimization. is a preset fourth threshold, which can be flexibly set according to the actual situation. For example, the second threshold can be 25,000 times, and the fourth threshold can be 30,000 times. Then, the 25,001st to 30,000th training in the multi-stage training optimization is the third-stage training.

[0048] Next, the training optimization from the first stage to the third stage will be further described in combination with Figure 2 . In Figure 2 , "GT" represents the original image, "Render Result" represents the rendered image, "MLPs" represents the prediction module, " " represents the rendering loss, " " represents the adaptive wavelet loss, " " represents the plane compression loss, " " represents the coding compression loss, "Codec" represents entropy coding, " " represents the feature of the anchor point relative to the first plane of, " " represents the feature of the anchor point relative to the second plane of, " " represents the feature of the anchor point relative to the third plane of.

[0049] (1) Training optimization in the first stage: In some embodiments, the specific implementation process of the above step S210 may include the following steps S211 - S214, where the following steps start from the beginning, that is, from the first round of training in the multi - stage training optimization.

[0050] S211, according to the preset mask and the set of anchor points in the round of training, obtain a plurality of first anchor points.

[0051] In this step, use the preset mask to perform a masking operation on the set of anchor points in the round of training, aiming to mask the useless anchor points, and screen from the set of anchor points in the round of training the final anchor points participating in the round of training to improve the training efficiency, and then obtain a plurality of first anchor points, which will participate in the round of training. Among them, if the round of training is the first round of training in the first stage, then the set of anchor points in the round of training includes several initial anchor points; if the round of training is other rounds of training in the first stage except the first round of training, then the set of anchor points in the round of training includes the plurality of first anchor points in the round of training.

[0052] S212, according to the anchor point attributes of each first anchor point, obtain the rendered image of the round of training.

[0053] In this step, after completing the masking operation, perform rendering processing according to the anchor point attributes of each first anchor point combined with the rendering pipeline method of Scaffold - GS to obtain the rendered image of the round of training. Briefly speaking, convert the first anchor point into the neural Gaussian set represented by the anchor point, and then render the neural Gaussian set on a two - dimensional plane by rasterization to obtain the rendered image. This process belongs to the prior art and will not be elaborated here.

[0054] S213, according to the The rendered images of the round of training update the preset mask and the characteristic information of each first anchor point.

[0055] In this step, after the rendering process is completed, the rendered image of the round of training and the original image are used to calculate the L1 loss function, obtaining the target loss of the round of training, and the target loss of the round of training is used for gradient backpropagation to update the preset mask and optimize the anchor attributes and three-dimensional coordinates of each first anchor point. It can be understood that by optimizing the three-dimensional coordinates of the anchor points, the accuracy of the three-plane sampling can be ensured during the training optimization in the second and third stages. In addition, the L1 loss function refers to the Mean Absolute Error (MAE) loss function.

[0056] S214, if , is the preset first threshold, then the anchor point set for the next round of training is determined based on multiple first anchor points and jumps to the next round of training; otherwise, the updated preset mask is determined as the first mask, the updated characteristic information of each first anchor point is determined as the new characteristic information of the first anchor point, and each first anchor point is determined as each first target anchor point, thereby obtaining the first training result.

[0057] In this step, after the parameter update is completed, it is judged whether holds. If so, it indicates that the training optimization in the first stage has not been completed. At this time, the updated preset mask is used as the new preset mask, the updated characteristic information of each first anchor point is used as the new characteristic information of each first anchor point, multiple first anchor points are integrated into the anchor point set for the next round of training, and jumps to the next round of training to achieve iterative training. If not, it indicates that all rounds of training optimization in the first stage are completed. At this time, the training optimization in the first stage is ended, the updated preset mask is determined as the first mask, the updated characteristic information of each first anchor point is determined as the new characteristic information of the first anchor point, and each first anchor point is determined as each first target anchor point, thus obtaining the first training result.

[0058] (2) Training optimization in the second stage: In some embodiments, the specific implementation process of the above step S220 may include the following steps S221 - S227, where the following steps start from beginning, that is, starting from the round of training in the multi-stage training optimization.

[0059] S221, according to the first mask and the anchor point set of the round of training, obtain multiple second anchor points.

[0060] In this step, a first mask is used to perform a masking operation on the set of anchors in the th round of training, aiming to mask the useless anchors and screen out the final anchors participating in the th round of training from the set of anchors in the th round of training, so as to improve the training efficiency, and then obtain multiple second anchors, which will participate in the th round of training. Among them, if the th round of training is the first round of training in the second stage, the set of anchors in the th round of training includes multiple first target anchors; if the th round of training is other rounds of training in the second stage except the first round of training, the set of anchors in the th round of training includes the multiple second anchors in the th round of training.

[0061] S222. Perform anchor sampling processing based on a preset three-plane structure to obtain the three-plane feature sets of each second anchor.

[0062] In this step, after the masking operation is completed, all second anchors are projected and sampled on the preset three-plane structure, so as to obtain the three-plane feature set of each second anchor, and its specific implementation will be described in the following embodiments. Among them, the three-plane feature sets of the second anchors will be used to combine with the anchor attributes to predict the attribute distribution and quantization step of the second anchors. In this way, by using the three-plane structure for anchor projection and sampling, a smoother transition can be obtained in a high-resolution scene, thereby improving the accuracy of the anchor feature information and reducing the possibility of loss of anchor feature information, which is beneficial to improving the rendering quality and rendering fidelity.

[0063] Optionally, before the start of the training optimization in the second stage, initialize the preset three-plane structure.

[0064] Optionally, in some embodiments, if all second anchors are projected and sampled in each round of training, it will lead to low training efficiency. To this end, before the anchor sampling processing, all second anchors are randomly sampled to only select some second anchors to participate in the subsequent anchor sampling processing. It is found through experiments that randomly sampling 10% of the anchors for training can achieve an effect similar to that of using all anchors for training.

[0065] S223. According to the three-plane feature sets and three-dimensional coordinates of each second anchor, combine with a preset prediction module to obtain the predicted attribute distribution and quantization step of each second anchor.

[0066] It should be noted that predicting the attribute distribution may include the predicted mean and predicted variance of the anchor attributes, that is, the predicted mean and predicted variance of the offset coordinates, the predicted mean and predicted variance of the attribute features, and the predicted mean and predicted variance of the scale scaling factor. In addition, the quantization step is used to quantize the anchor attributes to ensure the relative stability of the anchor attributes during the compression and decompression rendering processes.

[0067] In this step, after completing the three-plane projection and sampling, the three-plane feature sets and three-dimensional coordinates of all the second anchors are input into a preset prediction module, and prediction processing is performed through the preset prediction module, thereby obtaining the predicted attribute distribution and quantization step of each second anchor.

[0068] S224. Quantize the anchor attributes of each second anchor according to the quantization step of each second anchor to obtain the quantized anchor attributes of each second anchor as the new anchor attributes of each second anchor.

[0069] In this step, after completing the prediction processing, for each second anchor, use the predicted quantization step of the second anchor to quantize the anchor attributes of the second anchor to obtain the quantized anchor attributes of the second anchor, and determine the quantized anchor attributes of the second anchor as the new anchor attributes of the second anchor, thereby completing the quantization update of the anchor attributes. In this way, the relative stability of the anchor attributes during the compression and decompression rendering processes can be improved. Among them, the quantization processing in the multi-stage training optimization follows the following formula (1): , (1); In formula (1), represents the quantized anchor attributes of the anchor during the training optimization stage, represents the anchor attributes of the anchor, represents a uniform distribution, is a preset initial step size, is the quantization step of the anchor predicted by the prediction module.

[0070] Optionally, the initial step size can be set according to the actual situation, and this embodiment does not make specific limitations on this.

[0071] S225. Obtain the rendering image of the th round of training according to the anchor attributes of each second anchor.

[0072] In this step, after completing the quantization processing, perform rendering and rasterization processing according to the new anchor attributes of each second anchor in combination with the method of three-dimensional Gaussian splashing technology, thereby obtaining the Rendered images of each round of training. Briefly, the second anchor points are converted into the neural Gaussian sets they represent, and then the neural Gaussian sets are rasterized and rendered on a two-dimensional plane to obtain the rendered images. This process belongs to the prior art and will not be elaborated here. Among them, the method of the above three-dimensional Gaussian splashing technique can be the rendering pipeline method of Scaffold-GS or the rendering pipeline method of the traditional three-dimensional Gaussian splashing technique, but it is not limited to this.

[0073] S226. According to the rendered images of each round of training, combined with the predicted attribute distributions, anchor point attributes, and quantization steps of each second anchor point, update the first mask, the preset three-plane structure, the preset prediction module, and the characteristic information of each second anchor point.

[0074] In this step, first, based on the rendered images of each round of training, as well as the predicted attribute distributions, anchor point attributes, and quantization steps of each second anchor point, calculate the objective loss of each round of training. This process will be elaborated in the following embodiments. Then, use the objective loss of each round of training for gradient backpropagation to update and optimize the learnable parameters of the first mask, the preset three-plane structure, the preset prediction module, and the characteristic information of each second anchor point. It can be understood that the anchor point attributes involved in this step are the quantized anchor point attributes.

[0075] S227. If , then determine the anchor point set for the next round of training based on multiple second anchor points and jump to the next round of training. Otherwise, determine the updated preset prediction module as the initial prediction module, the updated preset three-plane structure as the initial three-plane structure, the updated first mask as the second mask, the updated characteristic information of each second anchor point as the new characteristic information of each second anchor point, and each second anchor point as each second target anchor point, thereby obtaining the second training result.

[0076] In this step, after completing the parameter update, judge Whether it holds. If so, it indicates that the training optimization in the second stage has not been completed. At this time, the updated first mask is determined as the new first mask, the updated preset prediction module is determined as the new preset prediction module, the updated preset three-plane structure is determined as the new preset three-plane structure, the updated feature information of each second anchor point is determined as the new feature information of each second anchor point, the multiple second anchor points are integrated into the anchor point set for the next round of training, and the process jumps to the next round of training to achieve iterative training. If not, it indicates that all rounds of training optimization in the second stage have been completed. At this time, the training optimization in the second stage is ended, the updated preset prediction module is determined as the initial prediction module, the updated preset three-plane structure is determined as the initial three-plane structure, the updated first mask is determined as the second mask, the updated feature information of each second anchor point is determined as the new feature information of each second anchor point, and each second anchor point is determined as each second target anchor point. In this way, the second training result can be obtained, thus completing the training optimization in the second stage.

[0077] In some embodiments, the specific implementation process of the above step S222 may include any one of the following steps S01 - S02.

[0078] S01, if , then sampling and copying processing is performed on each second anchor point based on the preset three-plane structure to obtain the three-plane feature set of each second anchor point.

[0079] In this step, when the -th round of training is between the -th round of training and the -th round of training in the multi-stage training optimization, it indicates that the -th round of training is in the early stage of the training optimization in the second stage. At this time, the capacity of the three-plane structure is sufficient. For each second anchor point, sampling is performed on the preset three-plane structure based on the three-dimensional coordinates of the second anchor point to obtain the three-plane feature of the second anchor point and copy it times to obtain three-plane features of the second anchor point as the three-plane feature set of the second anchor point. In other words, in the early stage of the training optimization in the second stage, the three-plane feature set of the second anchor point includes three-plane features of the second anchor point.

[0080] S02, if , then sampling processing is performed on each second anchor point and the set of neighboring anchor points of each second anchor point based on the preset three-plane structure to obtain the three-plane feature set of each second anchor point.

[0081] In this step, when the -th round of training is between the -th round of training and the -th round of training in the multi-stage training optimization, it indicates that the The round of training is in the late stage of training optimization in the second phase, at which time the capacity of the three-plane structure is insufficient. In order to enhance the capacity of the three-plane structure, this embodiment designs a three-plane decoding method based on a clustering algorithm by utilizing the inherent relationship between the continuous three-plane and the anchor cluster. This decoding method can decode the three-plane features of the anchor and the three-plane features of its neighboring anchors, which can effectively enhance the capacity of the three-plane structure, and its performance is superior to decoding the anchor attributes of a single anchor, that is, superior to the above-mentioned sampling and replication processing. Specifically, the clustering algorithm used is the K-Nearest Neighbor (KNN) algorithm. For a given anchor, first, find the nearest anchors to this anchor as neighboring anchors according to the spatial distance, and integrate these neighboring anchors into the neighboring anchor set of this anchor; then, sample on the three-plane structure based on the three-dimensional coordinates of this anchor to obtain the three-plane features of this anchor, and at the same time sample on the three-plane structure based on the three-dimensional coordinates of its neighboring anchors to obtain the three-plane features of these neighboring anchors of this anchor; finally, integrate the three-plane features of this anchor and the three-plane features of these neighboring anchors of this anchor into the three-plane feature set of this anchor. In other words, in the late stage of training optimization in the second phase, the three-plane feature set of the second anchor includes the three-plane features of the second anchor and the three-plane features of the neighboring anchors of the second anchor.

[0082] As a further embodiment, sampling the anchor on the three-plane structure to obtain the three-plane features of the anchor, and its specific implementation process is as follows: The three-plane structure includes the first plane , the second plane and the third plane these three two-dimensional planes, and each two-dimensional plane can be defined as , represents the resolution, represents the number of feature dimensions. First, decompose the original three-dimensional coordinates of the anchor into three coordinates , and , and then sample the anchor on the first plane , the second plane and the third plane based on these three coordinates to obtain the features of the anchor relative to the first plane , the features of the anchor relative to the second plane and the features of the anchor relative to the third plane characteristics , and finally the characteristics , characteristics and characteristics are tensor concatenated to obtain the three-plane characteristics of the anchor points.

[0083] It should be noted that in the second stage, if a clustering algorithm is introduced in the early stage of training optimization to implement anchor sampling processing, since the three-plane structure has not been fully trained in the early stage of training optimization, the introduction of the clustering algorithm will cause the training of the three-plane structure to be unstable, thus affecting the projection sampling effect in the compression stage. Therefore, the clustering algorithm is not introduced in the early stage of training optimization, and in order to align the input dimensions of the prediction module, the three-plane characteristics of the projected anchor points are copied times, so as to obtain the relevant characteristics of the anchor points; while the three-plane structure has been trained to a certain extent in the later stage of training optimization. At this time, the introduction of the clustering algorithm will not cause the training of the three-plane structure to be unstable. Therefore, the clustering algorithm is introduced in the later stage of training optimization to enhance the capacity of the three-plane structure, and in order to align the input dimensions of the prediction module, the three-plane characteristics of the anchor points and the three-plane characteristics of the nearest neighbor anchor points of the anchor points are obtained through the K-nearest neighbor clustering algorithm, so as to obtain the relevant characteristics of the anchor points.

[0084] Optionally, since the three-dimensional Gaussian is unbounded while the three-plane is bounded, it is easy to have a problem of mismatch between the upper and lower bounds during the three-plane projection. To solve this problem, in some embodiments, the present application uses a scaling function to fold the Gaussian point coordinates exceeding a certain boundary into the space with a definite upper bound, so that they can be evenly distributed in the space, thereby meeting the projection requirements of the three-plane. Among them, the scaling function satisfies the following formula (2): (2).

[0085] Optionally, the early stage of training optimization in the second stage refers to the th round to the th round of multi-stage training optimization. The later stage of training optimization in the second stage refers to the th round to the th round of multi-stage training optimization. is a preset third threshold, which can be flexibly set according to the actual situation. For example, the first threshold can be 10000, the second threshold can be 25000, and the third threshold can be 15000. Then, the 10001-15000th rounds of multi-stage training optimization are the early stage of training optimization in the second stage, and the 15001-25000th rounds of training are the later stage of training optimization in the second stage.

[0086] In some embodiments, the specific implementation process of the above step S226 may include the following steps S11 - S14.

[0087] S11, based on the rendered image and the original image of the th round of training, obtain the rendering loss and wavelet loss of the th round of training.

[0088] In this step, the rendering loss and wavelet loss of the th round of training jointly characterize the loss of the th round of training in the rendering and rasterization processes. For the rendering loss, calculate the L1 loss function based on the rendered image and the original image of the th round of training to obtain the rendering loss of the th round of training. Since the quantization operation is used in the foregoing steps and the quantization operation will also be used in the subsequent compression process, this quantization operation may cause the edges of the object to be blurred and floating, affecting the rendering quality. In this regard, this embodiment proposes an adaptive learning loss based on wavelet transform, simply referred to as the adaptive wavelet loss, which can enable the model to focus on low-frequency information in the early stage of training and high-frequency information in the later stage of training, thereby improving the rendering quality and compression quality. Specifically, for the wavelet loss, calculate the adaptive wavelet loss based on the rendered image and the original image of the th round of training to obtain the wavelet loss of the th round of training. Among them, the adaptive wavelet loss satisfies the following formula (3): (3); In formula (3), represents the adaptive wavelet loss; represents the original image; represents the rendered image; represents the low-frequency information obtained by performing wavelet transform on the original image; represents the low-frequency information obtained by performing wavelet transform on the rendered image; represents the high-frequency information obtained by performing wavelet transform on the original image; represents the high-frequency information obtained by performing wavelet transform on the rendered image; represents the adaptive weight of the low-frequency information; represents the adaptive weight of the high-frequency information; represents the calculation of the L1 loss function.

[0089] It should be noted that in the adaptive wavelet loss, the adaptive weights of the low-frequency information and the high-frequency information both change automatically according to the number of iterations. Specifically, when When it is greater than a preset wavelet threshold, it is regarded as the round of training is in the later stage of multi-stage training optimization. At this time, high-frequency information is focused on, and the adaptive weight of high-frequency information is greater than the adaptive weight of low-frequency information; when is less than the wavelet threshold, it is regarded as the round of training is in the early stage of multi-stage training optimization. At this time, low-frequency information is focused on, and the adaptive weight of low-frequency information is greater than the adaptive weight of high-frequency information; when is equal to the wavelet threshold, the two adaptive weights are equal. In this way, the adaptive wavelet loss proposed in this application can automatically change the ratio of the loss of high-frequency information during the training process, so as to ensure good performance in the rendering of high-frequency features (i.e., object edges) of the scene.

[0090] Optionally, the value range of the adaptive weight can be set according to the actual situation, and this embodiment does not make specific limitations in this regard. For example, the value of the adaptive weight can be set between 0.01 and 0.02.

[0091] Optionally, the value of the wavelet threshold can be set according to the actual situation, and this embodiment does not make specific limitations in this regard.

[0092] S12. According to the predicted attribute distribution, anchor point attributes, and quantization step of each second anchor point, obtain the round of training coding compression loss.

[0093] In this step, the coding compression loss of the round of training can reflect the loss of the round of training in entropy coding. This coding compression loss is the entropy loss mentioned in the Hash-grid Assisted Context for 3D Gaussian Splatting Compression (HAC), and will not be elaborated here.

[0094] Optionally, considering the issues of calculation time and computational complexity, in some embodiments, this application uses a random sampling method to select some second anchor points. The selected second anchor points will participate in the calculation of the entropy coding loss, while the unselected second anchor points will not participate in the calculation of the entropy coding loss. Through experiments, it is found that randomly sampling 10% of the anchor points for the calculation of the entropy coding loss can achieve an effect similar to that of using all anchor points for the calculation of the entropy coding loss.

[0095] S13. According to the rendering loss, wavelet loss, and coding compression loss of the round of training, obtain the round of training target loss.

[0096] In this step, the first weight is assigned to the rendering loss of the th round of training, the second weight is assigned to the wavelet loss of the th round of training, the third weight is assigned to the encoding compression loss of the th round of training, and then the rendering loss, wavelet loss, and encoding compression loss of the th round of training are weighted to obtain the target loss of the th round of training.

[0097] Optionally, the first weight, the second weight, and the third weight can all be set according to the actual situation, and this embodiment does not make specific limitations on this.

[0098] S14. Update the first mask, the preset three-plane structure, the preset prediction module, and the characteristic information of each second anchor point by using the target loss of the th round of training.

[0099] In this step, the target loss of the th round of training is used for gradient backpropagation to update and optimize the first mask, the preset three-plane structure, the preset prediction module, and the characteristic information of each second anchor point.

[0100] (3) Training optimization in the third stage: In some embodiments, the specific implementation process of the above step S230 may include the following steps S231 - S238, where the following steps start from i.e., start from the th round of training in the multi-stage training optimization.

[0101] S231. Obtain a plurality of third anchor points according to the second mask and the anchor point set of the th round of training.

[0102] In this step, the second mask is used to perform a masking operation on the anchor point set of the th round of training, aiming to mask the useless anchor points and screen out the anchor points participating in the th round of training from the anchor point set of the th round of training to improve the training efficiency, and then obtain a plurality of third anchor points, which will participate in the th round of training. Among them, if the th round of training is the first round of training in the third stage, the anchor point set of the th round of training includes a plurality of second target anchor points; if the th round of training is other rounds of training in the third stage except the first round of training, the anchor point set of the th round of training includes the plurality of third anchor points of the th round of training.

[0103] S232. Based on the initial three - plane structure, sample each third anchor point and the set of neighboring anchor points of each third anchor point to obtain the three - plane feature set of each third anchor point.

[0104] In this step, after completing the masking operation, project and sample all third anchor points on the initial three - plane structure, thereby obtaining the three - plane feature set of each third anchor point. It should be noted that the method of performing projection sampling on the third anchor points in this step follows the implementation method of the above - mentioned step S02. Among them, the three - plane feature set of the third anchor points will be used in combination with the anchor point attributes to predict the attribute distribution and quantization step size of the third anchor points. In this way, by using the three - plane structure for anchor point projection and sampling, a smoother transition can be obtained in a high - resolution scenario, thereby improving the accuracy of anchor point feature information and reducing the possibility of loss of anchor point feature information, which is beneficial to improving the rendering quality and rendering fidelity.

[0105] Optionally, in some embodiments, if all third anchor points are projected and sampled in each round of training, it will lead to low training efficiency. To this end, before performing anchor point sampling processing, perform random sampling on all third anchor points to only select some third anchor points to participate in the subsequent anchor point sampling processing. Through experiments, it is found that randomly sampling 10% of the anchor points for training can achieve an effect similar to that of training with all anchor points.

[0106] S233. According to the three - plane feature set and three - dimensional coordinates of each third anchor point, in combination with the initial prediction module, obtain the predicted attribute distribution and quantization step size of each third anchor point.

[0107] In this step, after completing the three - plane projection and sampling, input the three - plane feature set and three - dimensional coordinates of all third anchor points into the initial prediction module, and perform prediction processing through the initial prediction module, thereby obtaining the predicted attribute distribution and quantization step size of each third anchor point.

[0108] S234. Quantize the anchor point attributes of each third anchor point according to the quantization step size of each third anchor point to obtain the quantized anchor point attributes of each third anchor point as the new anchor point attributes of each third anchor point.

[0109] In this step, after completing the prediction processing, for each third anchor point, use the predicted quantization step size of the third anchor point to quantize the anchor point attributes of the third anchor point, obtain the quantized anchor point attributes of the third anchor point, and determine the quantized anchor point attributes of the third anchor point as the new anchor point attributes of the third anchor point, thereby completing the quantization update of the anchor point attributes. In this way, the stability of the anchor point attributes during the compression and decompression rendering process can be improved. Among them, the quantization processing follows the above - mentioned formula (1).

[0110] S235. Obtain the rendering image of the n-th round of training according to the anchor point attributes of each third anchor point.

[0111] In this step, after the quantization process, rendering and rasterization processing are performed according to the new anchor point attributes of each third anchor point in combination with the method of three-dimensional Gaussian splashing technology, so as to obtain the rendering image of the n-th round of training. Briefly speaking, the third anchor point is transformed into the neural Gaussian set represented by the anchor point, and then the neural Gaussian set is rendered on a two-dimensional plane through rasterization to obtain the rendering image. This process belongs to the prior art and will not be elaborated here. Among them, the method of the above three-dimensional Gaussian splashing technology can be the rendering pipeline method of Scaffold-GS, or the rendering pipeline method of the traditional three-dimensional Gaussian splashing technology, but it is not limited to this.

[0112] S236. Obtain the compressed initial three-plane structure and the restored initial three-plane structure according to the initial three-plane structure and the preset compression module.

[0113] In this step, considering that in the subsequent compression process, projection sampling of the three planes needs to be performed on the anchor points, and the parameters of the three-plane structure are large, which is not conducive to the compression process of the three-dimensional Gaussian splashing parameters. Therefore, on the basis of the mask, anchor points, three-plane structure and prediction module, a compression module is introduced for the training optimization of the third stage. The compression module is used to perform compression operations and restoration operations on the three-plane structure. It should be noted that in this application, the compression module is selected to be introduced in the training optimization of the third stage, which is to allow the three-plane structure to be fully trained in the training optimization of the second stage and then compressed, so as to improve the integrity of the potential features of the three-plane structure, which is beneficial to improving the rendering fidelity and rendering quality. Specifically, while performing rendering, first input the initial three-plane structure into the preset downsampling module, and perform a compression operation on the initial three-plane structure through the preset downsampling module, aiming to downsample the initial three-plane structure into latent features, which are used to indicate rich information in the three planes, so as to obtain the compressed initial three-plane structure. Then, input the compressed initial three-plane structure into the preset upsampling module, and perform a restoration operation on the compressed initial three-plane structure through the preset upsampling module to obtain the restored initial three-plane structure, and the restored initial three-plane structure will be used for loss calculation.

[0114] S237. Update the initial three-plane structure, the initial prediction module, the second mask, the preset compression module and the characteristic information of each third anchor point according to the initial three-plane structure, the restored initial three-plane structure and the rendering image of the n-th round of training, in combination with the predicted attribute distribution, anchor point attributes and quantization step of each third anchor point.

[0115] In this step, first, based on the initial three-plane structure, the restored initial three-plane structure, the rendered images of the round of training, as well as the predicted attribute distributions, anchor attributes, and quantization steps of each third anchor point, the objective loss for the round of training is calculated. This process will be described in the following embodiments; then, the objective loss for the round of training is used for gradient backpropagation to update and optimize the initial three-plane structure, the learnable parameters of the initial prediction module, the second mask, the learnable parameters of the preset compression module, and the characteristic information of each third anchor point.

[0116] S238, if , then determine the anchor set for the next round of training based on multiple third anchor points and jump to the next round of training; otherwise, determine the target upsampling module based on the updated preset compression module, determine the updated initial prediction module as the target prediction module, determine the compressed initial three-plane structure as the compressed three-plane structure, determine the updated second mask as the target mask, determine the updated characteristic information of each third anchor point as the new characteristic information of each third anchor point, determine each third anchor point as each optimal anchor point, and thus obtain the target training result.

[0117] In this step, after completing the parameter update, it is judged whether holds. If so, it means that the training optimization in the third stage has not been completed. At this time, determine the updated initial three-plane structure as the new initial three-plane structure, determine the updated initial prediction module as the new initial prediction module, determine the updated second mask as the new second mask, determine the updated preset compression module as the new preset compression module, determine the updated characteristic information of each third anchor point as the new characteristic information of each third anchor point, integrate multiple third anchor points into the anchor set for the next round of training, and jump to the next round of training to achieve iterative training. If not, it means that all rounds of training optimization in the third stage have been completed. At this time, end the training optimization in the third stage, select the updated preset upsampling module from the updated preset compression module as the target upsampling module, determine the updated initial prediction module as the target prediction module, determine the compressed initial three-plane structure as the compressed three-plane structure, determine the updated second mask as the target mask, determine the updated characteristic information of each third anchor point as the new characteristic information of each third anchor point, determine each third anchor point as each optimal anchor point, and thus obtain the target training result, thereby completing the multi-stage training optimization.

[0118] In some embodiments, the specific implementation process of the above step S237 may include the following steps S21 - S25.

[0119] S21, according to the Rendered images and original images for each round of training to obtain the rendering loss and wavelet loss for the

[0120] round of training. In this step, the rendering loss and wavelet loss for the round of training jointly characterize the loss during the rendering and rasterization processes for the round of training. For the rendering loss, the L1 loss function is calculated based on the rendered image and the original image for the round of training to obtain the rendering loss for the round of training. Since the quantization operation is used in the foregoing steps and will also be used in subsequent compression processing, this quantization operation may cause the edges of objects to be blurred and floating, affecting the rendering quality. Therefore, in this embodiment, the adaptive wavelet loss is calculated based on the rendered image and the original image for the round of training to obtain the wavelet loss for the

[0121] round of training. Among them, the adaptive wavelet loss is as shown in the above formula (3). S22. According to the predicted attribute distribution, anchor point attributes, and quantization step size of each third anchor point, obtain the

[0122] encoding compression loss for the round of training. In this step, the encoding compression loss for the

[0123] round of training can reflect the loss of the

[0124] round of training in entropy coding. This encoding compression loss is the entropy loss mentioned in HAC and will not be elaborated here. Optionally, considering the issues of calculation time and computational complexity, in some embodiments, the present application adopts a random sampling method to select some of the third anchor points. The selected third anchor points will participate in the calculation of the entropy coding loss, while the unselected third anchor points will not participate in the calculation of the entropy coding loss. Through experiments, it is found that randomly sampling 10% of the anchor points for the calculation of the entropy coding loss can achieve an effect similar to that of calculating the entropy coding loss using all anchor points.

[0125] S23. According to the initial three-plane structure and the restored initial three-plane structure, obtain the plane compression loss for the round of training. Plane compression loss for round training.

[0126] S24. Based on the rendering loss, wavelet loss, coding compression loss, and plane compression loss for the round of training, obtain the objective loss for the round of training.

[0127] In this step, assign the fourth weight to the rendering loss for the round of training, assign the fifth weight to the wavelet loss for the round of training, assign the sixth weight to the coding compression loss for the round of training, assign the seventh weight to the plane compression loss for the round of training. Then, perform a weighted operation on the rendering loss, wavelet loss, coding compression loss, and plane compression loss for the round of training to obtain the objective loss for the round of training.

[0128] Optionally, the fourth weight, fifth weight, sixth weight, and seventh weight can all be set according to the actual situation, and this embodiment does not make specific limitations on this.

[0129] S25. Use the objective loss for the round of training to update the characteristics information of the initial three-plane structure, initial prediction module, second mask, preset compression module, and each third anchor point.

[0130] In this step, use the objective loss for the round of training to perform gradient backpropagation to update and optimize the characteristics information of the initial three-plane structure, initial prediction module, second mask, preset compression module, and each third anchor point.

[0131] The above compression stage will be further described below.

[0132] In some embodiments, the specific implementation process of the above step S300 may include the following steps S310 - S360.

[0133] S310. Based on the target mask and multiple optimal anchor points, obtain multiple anchor points to be compressed.

[0134] It should be noted that the anchor points to be compressed refer to the final anchor points participating in the compression process.

[0135] In this step, perform a masking operation on the multiple optimal anchor points using the target mask, aiming to further mask the useless anchor points and select the final anchor points suitable for the compression process to improve the compression efficiency, thereby obtaining multiple anchor points to be compressed, and these anchor points will participate in the compression process.

[0136] S320. Obtain the target three - plane structure according to the compressed three - plane structure and the target up - sampling module.

[0137] In this step, in order to reduce the storage space occupied by the three - plane structure, a compression operation is performed on the three - plane structure during the training optimization in the third stage to obtain the compressed three - plane structure. And in the compression stage, it is necessary to restore the compressed three - plane structure to the original form of the three - plane structure to facilitate projection sampling of the anchor points, thereby improving the compression efficiency and being beneficial to enhancing the rendering fidelity and rendering quality. For this reason, in this step, the compressed three - plane structure is input into the target up - sampling module, and the target up - sampling module performs a restoration operation on the compressed three - plane structure to obtain the restored three - plane structure, that is, the target three - plane structure.

[0138] S330. Perform sampling and replication processing on each anchor point to be compressed based on the target three - plane structure to obtain the three - plane feature set of each anchor point to be compressed.

[0139] In this step, after completing the restoration operation of the three - plane structure, all the anchor points to be compressed are subjected to projection sampling processing on the target three - plane structure, so as to obtain the three - plane features of each anchor point to be compressed as the three - plane feature set of each anchor point to be compressed. It should be noted that the way of performing projection sampling processing on the anchor points in the compression stage follows the implementation method of the above - mentioned step S01, aiming to ensure the consistency of the anchor point dimensions in the multi - stage training and the compression stage.

[0140] S340. According to the three - plane feature set and three - dimensional coordinates of each anchor point to be compressed, combined with the target prediction module, obtain the predicted attribute distribution and quantization step size of each anchor point to be compressed.

[0141] In this step, after completing the anchor point sampling processing, the three - plane feature set and three - dimensional coordinates of all the anchor points to be compressed are input into the target prediction module, and prediction processing is performed through the target prediction module, so as to obtain the predicted attribute distribution and quantization step size of each anchor point to be compressed.

[0142] S350. Quantize the anchor point attributes of each anchor point to be compressed according to the quantization step size of each anchor point to be compressed to obtain the quantized anchor point attributes of each anchor point to be compressed as the new anchor point attributes of each anchor point to be compressed.

[0143] In this step, after the prediction process is completed, for each anchor point to be compressed, the quantization step of the anchor point to be compressed obtained by prediction is used to perform quantization processing on the anchor point attributes of the anchor point to be compressed, obtaining the quantized anchor point attributes of the anchor point to be compressed, and determining the quantized anchor point attributes of the anchor point to be compressed as the new anchor point attributes of the anchor point to be compressed, thereby completing the quantization update of the anchor point attributes. In this way, the stability of the anchor point attributes during the compression and decompression rendering processes can be improved. Among them, the quantization processing in the compression stage follows the following formula (4): , (4); In formula (4), represents the quantized anchor point attributes of the anchor point in the compression stage; represents the anchor point attributes of the anchor point; represents rounding the floating point number or integer; represents the preset initial step size, represents the quantization step of the anchor point predicted by the prediction module.

[0144] Optionally, the initial step size can be set according to the actual situation, and this embodiment does not make specific limitations on this.

[0145] S360. According to the predicted attribute distribution and anchor point attributes of each anchor point to be compressed, obtain the bitstream data of each anchor point to be compressed as the target compression data.

[0146] It should be noted that the target compression data includes the bitstream data of each anchor point to be compressed.

[0147] It can be understood that since the anchor point attributes follow a Gaussian distribution, entropy coding can be used to compress the anchor point attributes, thereby realizing the compression processing of the three-dimensional Gaussian sputtering parameters.

[0148] In this step, after the quantization operation is completed, arithmetic coding is performed on the predicted attribute distribution and anchor point attributes of each anchor point to be compressed. Through arithmetic coding, the entropy coding of the anchor point attributes of each anchor point to be compressed can be completed, and finally the data saved in the form of a bitstream is obtained as the target compression data. In this way, the compression processing of the three-dimensional Gaussian splash parameters is realized, which can effectively reduce the storage space required by the three-dimensional Gaussian splash technology, thereby reducing the storage cost of the three-dimensional Gaussian splash technology. At the same time, the possibility of key parameter loss is reduced, and the subsequent rendering quality and rendering fidelity are improved.

[0149] For the convenience of understanding the above three-dimensional Gaussian splash method in the embodiments of the present application, the following will use an application scenario and combine Figure 2 to illustrate the principle of the above three-dimensional Gaussian splash method in the embodiments of the present application.

[0150] S101. First, use the SfM technique to estimate point cloud data from at least one target image, and determine this point cloud data as the target point cloud. Then, initialize the target point cloud. After that, generate anchor points and neural Gaussian distributions through the initialized target point cloud, so as to obtain a number of initial anchor points and their anchor point attributes and three-dimensional coordinates. The anchor point attributes include offset coordinates, attribute features, and scale factors.

[0151] S102. Training optimization in the 1st to 10000th rounds, that is, the training optimization in the first stage: In each round of training optimization in the first stage, first use a preset mask to perform a masking operation on the anchor point set of this round of training to obtain multiple first anchor points participating in this round of training. Then, perform rendering processing according to the anchor point attributes of each first anchor point in combination with the rendering pipeline method of Scaffold-GS to obtain the rendered image of this round of training. After that, calculate the target loss of this round of training by using the rendered image and the original image of this round of training, and use the target loss of this round of training for gradient backpropagation to update the preset mask and optimize the feature information of each first anchor point. Finally, integrate the multiple first anchor points of this round of training into the anchor point set of the next round of training, and jump to the next round of training until the 10000th round of training optimization is completed, ending the training optimization in the first stage. Through multiple rounds of training optimization in the first stage, a first mask, multiple first target anchor points, and the feature information of each first target anchor point can be obtained as the first training result.

[0152] S103. Training optimization in the 10001st to 25000th rounds, that is, the training optimization in the second stage: The training optimization in the 10001st to 15000th rounds is the early stage of the second stage.

[0153] First, perform adaptive mask screening and triplane projection: Use the first mask to perform a masking operation on the anchor point set of this round of training to obtain multiple second anchor points participating in this round of training. For each second anchor point, sample on a preset triplane structure based on the three-dimensional coordinates of the second anchor point to obtain the triplane features of the second anchor point and replicate times to obtain triplane features of the second anchor point as the triplane feature set of the second anchor point.

[0154] Secondly, perform prediction: Input the triplane feature sets and three-dimensional coordinates of all second anchor points into a preset prediction module, and perform prediction processing through the preset prediction module to obtain the predicted attribute distribution and quantization step size of each second anchor point. Among them, the preset prediction module is an MLP.

[0155] Then, a quantization operation is performed: the anchor attributes of each second anchor point are quantized according to the quantization step of each second anchor point, which follows the above formula (1), so as to obtain the quantized anchor attributes of each second anchor point as the new anchor attributes of each second anchor point.

[0156] After that, rendering and rasterization are performed: rendering and rasterization processing are carried out according to the new anchor attributes of each second anchor point in combination with the method of three-dimensional Gaussian splashing technology, so as to obtain the rendered image of this round of training.

[0157] Finally, gradient backpropagation is performed: the L1 loss function is calculated based on the rendered image and the original image of this round of training to obtain the rendering loss of this round of training, and the adaptive wavelet loss is calculated based on the rendered image and the original image of this round of training, which follows the above formula (3), to obtain the wavelet loss of this round of training, and the coding compression loss of this round of training is obtained according to the predicted attribute distribution, anchor attributes and quantization step of each second anchor point, which follows the entropy loss in HAC. Then, the rendering loss, wavelet loss and coding compression loss of this round of training are weighted and calculated to obtain the target loss of this round of training, and the target loss of this round of training is used for gradient backpropagation to update and optimize the feature information of the first mask, the preset three-plane structure, the preset prediction module and each second anchor point. After the gradient backpropagation is completed, the multiple second anchor points of this round of training are integrated into the anchor point set of the next round of training, and jump to the next round of training until the 15000th round of training optimization is completed, and the preliminary training optimization of the second stage is ended.

[0158] The training optimization from the 15001st to the 25000th round is the later stage of the second stage. The process of the later stage training optimization of the second stage is roughly the same as that of the preliminary training optimization of the second stage, and the difference lies in the three-plane projection. Specifically, in the three-plane projection of the later stage of the second stage, first, the anchor points with the shortest distance to this anchor point are found according to the spatial distance as the neighboring anchor points, and these neighboring anchor points are integrated into the neighboring anchor point set of this anchor point; then, sampling is performed on the three-plane structure based on the three-dimensional coordinates of this anchor point to obtain the three-plane feature of this anchor point, and at the same time, sampling is performed on the three-plane structure based on the three-dimensional coordinates of its neighboring anchor points to obtain the three-plane features of the neighboring anchor points of this anchor point; finally, the three-plane feature of this anchor point and the three-plane features of the neighboring anchor points of this anchor point are integrated into the three-plane feature set of this anchor point. When the 25000th round of training optimization is completed, the later stage training optimization of the second stage is ended.

[0159] Through multiple rounds of training optimization in the second stage, an initial three-plane structure, an initial prediction module, a second mask, multiple second target anchor points, and the feature information of each of the second target anchor points can be obtained as the second training result.

[0160] S104, training optimization from the 25001st to the 30000th round, that is, the training optimization in the third stage. The process of the training optimization in the third stage is roughly the same as that of the late training optimization in the second stage, and the differences are as follows: First, after the quantization operation and before the gradient backpropagation, in the third stage, the initial three-plane structure needs to be input into a preset downsampling module. The initial three-plane structure is compressed by the preset downsampling module to obtain a compressed initial three-plane structure, and the compressed initial three-plane structure is input into a preset upsampling module. The compressed initial three-plane structure is restored by the preset upsampling module to obtain a restored initial three-plane structure, so as to facilitate the calculation of the plane compression loss. Among them, the upsampling module is a transposed convolutional neural network, and the downsampling module is a convolutional neural network.

[0161] Second, in the gradient backpropagation, in addition to calculating the wavelet loss, rendering loss, and coding compression loss of this round of training, the third stage also needs to calculate the plane compression loss of this round of training. Specifically, based on the initial three-plane structure and the restored initial three-plane structure, the L1 loss function is calculated to obtain the plane compression loss of this round of training. Then, the rendering loss, wavelet loss, coding compression loss, and plane compression loss of this round of training are weighted and calculated to obtain the target loss of this round of training. In addition, in addition to performing gradient backpropagation updates on the initial three-plane structure, the second mask, the initial prediction module, and the feature information of the third anchor points, the third stage also needs to perform gradient backpropagation updates on the preset compression module.

[0162] When the 30000th round of training optimization is completed, the training optimization in the third stage ends. Through multiple rounds of training optimization in the third stage, a target upsampling module, a target prediction module, a compressed three-plane structure, a target mask, multiple optimal anchor points, and the feature information of each optimal anchor point can be obtained as the target training result.

[0163] S105, compression stage: First, use the target mask to perform a masking operation on multiple optimal anchor points to obtain multiple anchor points to be compressed. At the same time, input the compressed three-plane structure into the target upsampling module, and the compressed three-plane structure is restored by the target upsampling module to obtain a restored three-plane structure. For each anchor point to be compressed, sampling is performed on the target three-plane structure based on the three-dimensional coordinates of the anchor point to be compressed to obtain the three-plane feature of the anchor point to be compressed and duplicate times to obtain the A set of three-plane features of three-plane features as anchor points to be compressed; then, input the set of three-plane features and three-dimensional coordinates of all anchor points to be compressed into the target prediction module, and perform prediction processing through the target prediction module to obtain the predicted attribute distribution and quantization step size of each anchor point to be compressed; after that, for each anchor point to be compressed, use the quantization step size of the anchor point to be compressed obtained by prediction to perform quantization processing on the anchor point attributes of the anchor point to be compressed, and obtain the quantized anchor point attributes of the anchor point to be compressed. The quantization process follows the above formula (4); finally, perform arithmetic coding on the predicted attribute distribution and anchor point attributes of each anchor point to be compressed. Through arithmetic coding, the entropy coding of the anchor point attributes of each anchor point to be compressed can be completed, and finally the data saved in the form of a bitstream is obtained as the target compressed data, so as to realize the compression processing of the three-dimensional Gaussian splash parameters.

[0164] S106, decompression and rendering: Perform decoding processing on the target compressed data through a preset decoder to obtain a decoding result, and then perform differentiable rasterization rendering on the decoding result to obtain a rendered image of the target point cloud, so as to realize three-dimensional Gaussian splash.

[0165] The effects of the embodiments of the present application will be verified through the following Embodiment 1 to Embodiment 5.

[0166] Embodiment 1: Refer to Figure 3, both 3DGS and Scaffold-GS are three-dimensional Gaussian splashing techniques. Scaffold-GS is a variant of 3DGS, and neither of them has a compression stage. In 3DGS, after the SfM point cloud is initialized, it becomes three-dimensional Gaussians (3DGaussians), and then a rendered image can be obtained through a differentiable rasterization renderer. This rendered image includes seven million three-dimensional Gaussians (7 Million Gaussians) and occupies 1.1 gigabytes (GB) of space. In the variant Scaffold-GS of 3DGS, after the SfM point cloud is initialized, it becomes anchors. Each anchor has corresponding attributes and represents a group of Gaussians. Then a rendered image can be obtained through a differentiable rasterization renderer. This rendered image includes 0.88 million anchors (0.88 Million Anchors) and occupies 268 megabytes (MB). In this application, after the SfM point cloud is initialized, it becomes anchors. Each anchor has corresponding attributes and represents a group of Gaussians, which is the same as the processing method of Scaffold-GS. Then the anchors are projected onto a three-plane structure and distribution prediction is performed. Then, entropy coding is implemented using the anchor attributes and their distributions to obtain compressed data in the form of a bitstream. Finally, the compressed data in the form of a bitstream is decoded and differentiable rasterization rendering processing is performed to obtain a rendered image. This rendered image includes 0.74 million anchors (0.74 Million Anchors) and occupies 19.3 megabytes (MB). It can be seen that this application can significantly reduce the storage space occupied by the three-dimensional Gaussian splashing technique.

[0167] Example 2: Based on the publicly available datasets Tank&Temple and DeepBlending, this application is compared with 3DGS and Scaffold-GS. 3DGS and Scaffold-GS do not have a compression stage. Refer to Figure 4 , Figure 4 where "GroundTruth" represents the original image, "3DGS" represents the rendered image after being processed by 3DGS, "Scaffold-GS" represents the rendered image after being processed by Scaffold-GS, and "Ours" represents the rendered image after being processed by this application. By comparing the peak signal-to-noise ratio (PSNR) of each method and the size of the rendered image, it can be known that this application can significantly reduce the storage space occupied by the three-dimensional Gaussian splashing technique and at the same time achieve comparable or even higher rendering quality.

[0168] Example 3: The present application is compared with the prior art, and the comparison results are shown in Table 1 below. In Table 1, "Datasets" represents the dataset, which includes the public datasets Synthetic-NeRF, Tank&Temple, and DeepBlending; "Methods" represents the methods, where 3DGS and Scaffold-GS are existing three-dimensional compression splash techniques that do not have a compression stage, and the remaining methods are existing compression methods, while "TC-GS (Ours)" represents the method proposed in the present application; the metrics used are peak signal-to-noise ratio, structural similarity (Structural Similarity Index Measure, SSIM), learned perceptual image patch similarity (LPIPS), and the size of the rendered image. The larger the first two, the better, and the smaller the last two, the better. As can be seen from Table 1, on the public dataset Synthetic-NeRF, the present application achieves the best value in terms of SIZE; in the public dataset Tank&Temple, the present application achieves the best values in LPIPS and SIZE, and the second-best value in PSNR; in the public dataset DeepBlending, the present application achieves the best values in LPIPS and SIZE, which indicates that the present application can significantly reduce the storage space occupied by the three-dimensional Gaussian splash technique, and at the same time can improve the quality and fidelity of the rendered image.

[0169] Table 1: Comparison Table of Metrics between Existing Methods and the Method of the Present Application

[0170] Example 4: A detailed qualitative comparison is made between the present application and Scaffold-GS in terms of rendering details, as Figure 5 shown, Figure 5 where "Error Map" represents the error map. As can be seen from Figure 5 , the present application performs better in the processing of rendering details.

[0171] Example 5: This application is divided into four innovative parts: a three-plane structure, a compression module, an adaptive wavelet loss, and a mask. Ablation experiments were conducted on these four innovative parts, and the results of the ablation experiments are shown in Table 2 below. In Table 2, "full" represents the method containing the above four innovative parts (i.e., the overall method of this application), "w / o tri-plane" represents the method of this application with only the three-plane structure removed, "w / o tri-sompression" represents the method of this application with only the compression module removed, "w / o wavelet" represents the method of this application with only the adaptive wavelet loss removed, and "w / o anchor mask" represents the method of this application with only the mask removed. As can be seen from Table 2: The overall method of this application achieved the best values in four aspects: PSNR, SSIM, LPIPS, and SIZE. This indicates that the above four innovative parts in this application can jointly reduce the storage space occupied by the 3D Gaussian splashing technique significantly, while improving the quality and fidelity of the rendered images.

[0172] Table 2: Ablation experiment table of this application

[0173] The embodiments of this application found that the method of deriving neural Gaussians through anchor points can significantly reduce the information required to be stored, and the anchor point attributes exhibit a normal distribution, conforming to the Gaussian distribution. This means that it can be encoded in the form of a bitstream to achieve compression processing of the 3D Gaussian splashing parameters. In addition, due to the spatial correlation between neural Gaussians, the embodiments of this application proposed a three-plane structure composed of three two-dimensional planes with continuity. This three-plane structure has a strong internal relationship with the anchor cluster, and can provide a smoother transition in high-resolution scenarios, obtaining a high-fidelity rendering quality while compressing.

[0174] Based on this idea, the embodiments of this application proposed a 3D Gaussian splashing framework based on the three-plane structure and implicit-explicit hybrid representation. Through the three-dimensional coordinates of the anchor points, the three-plane structure, and key parts (such as masks, prediction modules, quantization operations, and entropy coding, etc.), the 3D Gaussian splashing parameters (which are anchor point attributes, and each anchor point is associated with a group of neural Gaussians, and the anchor point attributes map the characteristics of the neural Gaussians associated with the anchor points) are specifically compressed, and then the compressed 3D Gaussian splashing parameters are decompressed and rendered to achieve 3D Gaussian splashing.

[0175] It can be seen that the embodiment of the present application realizes targeted compression processing of three-dimensional Gaussian splash parameters while fully considering the characteristics of neural Gaussians, and then performs three-dimensional Gaussian splash based on the compressed three-dimensional Gaussian splash parameters, thereby obtaining a rendered image associated with the target point cloud. In this way, not only can the parameters of the saved scene be effectively reduced, thereby reducing the storage cost of the three-dimensional Gaussian splash technology, but also the possibility of parameter loss during compression can be effectively reduced, thereby improving the fidelity and quality of the rendered image.

[0176] Secondly, since the edges of the rendered object are extremely prone to blurring and artifacts, which affect the rendering quality, the embodiment of the present application also provides an adaptive wavelet loss, which can enable the three-dimensional Gaussian splash framework to focus on features of different frequencies in the training optimization at different stages, thereby enhancing the learning of high-frequency information such as edges by the three-dimensional Gaussian splash framework and improving the rendering quality and rendering fidelity of the three-dimensional Gaussian splash technology.

[0177] Furthermore, in order to improve the relative stability of the anchor point attributes during the compression and decompression rendering processes, thereby ensuring the rendering quality and rendering fidelity of the three-dimensional Gaussian splash technology, the embodiment of the present application also provides a quantization operation to enable the anchor point attributes to be quantized before rendering, thereby maintaining relative stability.

[0178] Finally, in order to ensure that all compressed anchor points are highly effective, the embodiment of the present application also provides a learnable mask, which is used to mask invalid anchor points in the multi-stage training optimization, thereby improving the training optimization efficiency, and at the same time masking invalid anchor points in the compression stage, thereby improving the compression efficiency.

[0179] In addition, the embodiment of the present application also provides a three-dimensional Gaussian splash device, which may include: a first processing module, a second processing module, a third processing module, and a fourth processing module. The first processing module is used to obtain the target point cloud and perform initialization processing to obtain a number of initial anchor points and the feature information of each initial anchor point. The characteristic information includes anchor point attributes and three-dimensional coordinates. The anchor point attributes include offset coordinates, attribute features, three-dimensional coordinates, and scale scaling factors. The second processing module is used to perform multi-stage training optimization according to the number of initial anchor points and the characteristic information of each initial anchor point to obtain a target training result, a plurality of optimal anchor points, and the characteristic information of each optimal anchor point. Among them, the target training result includes a target upsampling module, a target prediction module, a compressed three-plane structure, and a target mask. The third processing module is used to perform compression processing according to the target training result, a plurality of optimal anchor points, and the characteristic information of each optimal anchor point to obtain target compressed data. The fourth processing module is used to perform decoding and rendering processing on the target compressed data to obtain a target rendered image.

[0180] The content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented in the device embodiments of the present application are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0181] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application. The scope of the present application is defined by the claims and their equivalents.

[0182] The above is a specific description of the preferred embodiments of the present application, but the present application is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included in the scope defined by the claims of the present application.

Claims

1. A three-dimensional Gaussian splashing method, characterized in that: The method comprises the following steps: Acquire the target point cloud and perform initialization processing to obtain a number of initial anchor points and characteristic information of each of the initial anchor points; wherein the characteristic information includes anchor point attributes and three-dimensional coordinates, and the anchor point attributes include offset coordinates, attribute features, and scale scaling factors; Performing multi-stage training optimization according to the initial anchor points and the characteristic information of each of the initial anchor points to obtain a target training result, multiple optimal anchor points and the characteristic information of each of the optimal anchor points; wherein the target training result includes a target upsampling module, a target prediction module, a compressed three-plane structure and a target mask; Perform compression processing according to the target training result, the multiple optimal anchor points and the characteristic information of each optimal anchor point to obtain target compressed data; The target compressed data is decoded and rendered to obtain a target rendered image.

2. The three-dimensional Gaussian splashing method according to claim 1, characterized in that: The multi-stage training optimization is performed according to the plurality of initial anchor points and the characteristic information of each of the initial anchor points to obtain a target training result, a plurality of optimal anchor points and the characteristic information of each of the optimal anchor points, including: Performing a first-stage training optimization process according to a preset mask, a plurality of the initial anchor points, and characteristic information of each of the initial anchor points to obtain a first training result, wherein the first training result includes a first mask, a plurality of first target anchor points, and characteristic information of each of the first target anchor points; Performing a second-stage training optimization process according to the preset three-plane structure, the preset prediction module, and the first training result to obtain a second training result, wherein the second training result includes an initial three-plane structure, an initial prediction module, a second mask, a plurality of second target anchor points, and characteristic information of each of the second target anchor points; The third stage of training optimization processing is performed according to the preset compression module and the second training result to obtain the target training result, the multiple optimal anchor points and the characteristic information of each optimal anchor point.

3. The three-dimensional Gaussian splashing method according to claim 2, characterized in that: The first stage of training optimization processing is performed according to the preset mask, the plurality of initial anchor points and the characteristic information of each initial anchor point to obtain a first training result, including: According to the preset mask and the The anchor point set of round training is used to obtain multiple first anchor points; According to the anchor point attributes of each of the first anchor points, the first Rendered images of round training; According to updating the preset mask and characteristic information of each of the first anchor points using a rendered image of a round of training; like , is a preset first threshold, then based on the multiple first anchor points, an anchor point set for the next round of training is determined and the training is jumped to the next round; otherwise, the updated preset mask is determined as the first mask, the updated characteristic information of each first anchor point is determined as the new characteristic information of the first anchor point, and each first anchor point is determined as each first target anchor point, so as to obtain the first training result; Among them, the above steps are start.

4. The three-dimensional Gaussian splashing method according to claim 2, characterized in that: The second stage of training optimization processing is performed according to the preset three-plane structure, the preset prediction module and the first training result to obtain the second training result, including: According to the first mask and the A set of anchor points for round training is used to obtain multiple second anchor points; Perform anchor point sampling processing based on the preset three-plane structure to obtain a three-plane feature set of each second anchor point; According to the three-plane feature set and the three-dimensional coordinates of each of the second anchor points, in combination with the preset prediction module, a prediction attribute distribution and a quantization step size of each of the second anchor points are obtained; quantizing the anchor point attributes of each second anchor point according to the quantization step size of each second anchor point, and obtaining the quantized anchor point attributes of each second anchor point as new anchor point attributes of each second anchor point; According to the anchor point attributes of each of the second anchor points, the first Rendered image of the training round; According to The first mask, the preset three-plane structure, the preset prediction module and the characteristic information of each second anchor point are updated based on the rendered image of the round training, combining the predicted attribute distribution, the anchor point attribute and the quantization step size of each second anchor point; like , is a preset second threshold, then based on the multiple second anchor points, an anchor point set for the next round of training is determined and the next round of training is jumped to; otherwise, the updated preset prediction module is determined as the initial prediction module, the updated preset three-plane structure is determined as the initial three-plane structure, the updated first mask is determined as the second mask, the updated characteristic information of each second anchor point is determined as the new characteristic information of each second anchor point, and each second anchor point is determined as each second target anchor point, so as to obtain the second training result; Among them, the above steps are start.

5. The three-dimensional Gaussian splashing method according to claim 4, characterized in that: The anchor point sampling process is performed based on the preset three-plane structure to obtain a three-plane feature set of each second anchor point, including: like , is a preset third threshold, sampling and copying processing is performed on each of the second anchor points based on the preset three-plane structure to obtain a three-plane feature set of each of the second anchor points; Or, if , then, based on the preset three-plane structure, sampling processing is performed on each of the second anchor points and the neighboring anchor point sets of each of the second anchor points to obtain a three-plane feature set of each of the second anchor points.

6. The three-dimensional Gaussian splashing method according to claim 4, characterized in that: According to the The first mask, the preset three-plane structure, the preset prediction module and the characteristic information of each second anchor point are updated by combining the predicted attribute distribution, the anchor point attribute and the quantization step size of each second anchor point, including: According to The rendered images and original images of the first round of training are obtained. Rendering loss and wavelet loss for round training; According to the predicted attribute distribution, anchor point attribute and quantization step size of each second anchor point, the first Encoding compression loss for round training; According to The rendering loss, wavelet loss and coding compression loss of the first round of training are obtained. Target loss for round training; Use the The target loss of the round training updates the characteristic information of the first mask, the preset three-plane structure, the preset prediction module and each of the second anchor points.

7. The three-dimensional Gaussian splashing method according to claim 2, characterized in that: The third stage of training optimization processing is performed according to the preset compression module and the second training result to obtain the target training result, the multiple optimal anchor points and the characteristic information of each optimal anchor point, including: According to the second mask and the The anchor point set of round training is used to obtain multiple third anchor points; Based on the initial three-plane structure, sampling is performed on each of the third anchor points and a set of neighboring anchor points of each of the third anchor points to obtain a three-plane feature set of each of the third anchor points; According to the three-plane feature set and the three-dimensional coordinates of each of the third anchor points, combined with the initial prediction module, the prediction attribute distribution and the quantization step size of each of the third anchor points are obtained; quantizing the anchor point attributes of each of the third anchor points according to the quantization step size of each of the third anchor points, and obtaining the quantized anchor point attributes of each of the third anchor points as new anchor point attributes of each of the third anchor points; According to the anchor point attributes of each of the third anchor points, the first Rendered images of round training; According to the initial three-plane structure and the preset compression module, obtaining the compressed initial three-plane structure and the restored initial three-plane structure; According to the initial three-plane structure, the restored initial three-plane structure and the The rendered image of the round training is combined with the prediction attribute distribution, anchor point attribute and quantization step size of each of the third anchor points to update the initial three-plane structure, the initial prediction module, the second mask, the preset compression module and the characteristic information of each of the third anchor points; like , is a preset fourth threshold, then based on the multiple third anchor points, an anchor point set for the next round of training is determined and jumps to the next round of training; otherwise, based on the updated preset compression module, the target upsampling module is determined, the updated initial prediction module is determined as the target prediction module, the compressed initial three-plane structure is determined as the compressed three-plane structure, the updated second mask is determined as the target mask, the updated characteristic information of each of the third anchor points is determined as the new characteristic information of each of the third anchor points, and each of the third anchor points is determined as each of the optimal anchor points, so as to obtain the target training result; Among them, the above steps are start.

8. The three-dimensional Gaussian splashing method according to claim 7, characterized in that: the initial three-plane structure, the restored initial three-plane structure and the The method comprises: updating the initial three-plane structure, the initial prediction module, the second mask, the preset compression module and the characteristic information of each of the third anchor points in combination with the predicted attribute distribution, the anchor point attribute and the quantization step size of each of the third anchor points, including: According to The rendered images and original images of the first round of training are obtained. Rendering loss and wavelet loss for round training; According to the predicted attribute distribution, anchor point attribute and quantization step size of each of the third anchor points, the first Encoding compression loss for round training; According to the initial three-plane structure and the restored initial three-plane structure, the first Plane compression loss for round training; According to The rendering loss, wavelet loss, coding compression loss and plane compression loss of the first round of training are obtained. Target loss for round training; Use the The target loss of the round training updates the characteristic information of the initial three-plane structure, the initial prediction module, the second mask, the preset compression module and each of the third anchor points.

9. The three-dimensional Gaussian splashing method according to claim 1, characterized in that: The step of performing compression processing according to the target training result, the multiple optimal anchor points and the characteristic information of each optimal anchor point to obtain target compressed data includes: Obtaining a plurality of anchor points to be compressed according to the target mask and the plurality of optimal anchor points; Obtaining a target three-plane structure according to the compressed three-plane structure and the target upsampling module; Based on the target three-plane structure, each of the anchor points to be compressed is sampled and copied to obtain a three-plane feature set of each of the anchor points to be compressed; According to the three-plane feature set and the three-dimensional coordinates of each anchor point to be compressed, in combination with the target prediction module, the predicted attribute distribution and the quantization step size of each anchor point to be compressed are obtained; Quantizing the anchor point attributes of each anchor point to be compressed according to the quantization step size of each anchor point to be compressed, and obtaining the quantized anchor point attributes of each anchor point to be compressed as new anchor point attributes of each anchor point to be compressed; According to the predicted attribute distribution and anchor point attributes of each anchor point to be compressed, the bit stream data of each anchor point to be compressed is obtained as the target compressed data.

10. A three-dimensional Gaussian splashing device, characterized in that: include: A first processing module is used to obtain a target point cloud and perform initialization processing to obtain a plurality of initial anchor points and characteristic information of each of the initial anchor points; wherein the characteristic information includes anchor point attributes and three-dimensional coordinates, and the anchor point attributes include offset coordinates, attribute features, and scale scaling factors; A second processing module is used to perform multi-stage training optimization according to the initial anchor points and the characteristic information of each of the initial anchor points to obtain a target training result, multiple optimal anchor points and the characteristic information of each of the optimal anchor points; wherein the target training result includes a target upsampling module, a target prediction module, a compressed three-plane structure and a target mask; A third processing module is used to perform compression processing according to the target training result, the multiple optimal anchor points and the characteristic information of each optimal anchor point to obtain target compressed data; The fourth processing module is used to decode and render the target compressed data to obtain a target rendered image.

Citation Information

Cited By

  • Implicit neural Gaussian splashing method and device based on multiple states and medium

    CN121280580A

  • Static scene three-dimensional reconstruction method based on three-dimensional Gaussian splashing

    CN121414978A