A method for quickly constructing volume rendering Gaussian representation based on visual-body data alignment
By employing a Gaussian representation method for volume rendering aligned with visual-volume data, and utilizing a feedforward neural network and a volume geometry forcing mechanism, the method directly predicts the 3D Gaussian representation from volume data and its multi-view information. This solves the problem of low rendering efficiency in volume data visualization and achieves efficient and accurate volume data visualization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH OF CHINA
- Filing Date
- 2026-04-10
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies suffer from low rendering efficiency, high preprocessing costs, and difficulty in representing the internal structure of the data in volumetric visualization. They also lack efficient and accurate 3D representation conversion methods.
We adopt a Gaussian representation method for volume rendering based on vision-volume data alignment. We predict the 3D Gaussian representation directly from volume data and its multi-view information through a feedforward neural network, and combine it with a volume geometry forcing mechanism for alignment and fusion to avoid scene-by-scene optimization.
Significantly improves rendering efficiency, ensures consistency across multiple views and accurate representation of volume structure, and is suitable for real-time visualization and new perspective rendering of large-scale volume data.
Smart Images

Figure CN122115738A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of volume data visualization technology, and specifically to a method for rapidly constructing Gaussian representations of volume drawing based on visual-volume data alignment. Background Technology
[0002] Volumetric data visualization is a long-standing and important research area in computer graphics and scientific visualization. This type of problem typically takes three-dimensional volumetric data (such as medical images and scientific simulation data) as input and converts it into two-dimensional images through visualization techniques to assist users in analysis and understanding.
[0003] Direct Volume Rendering (DVR) is the most classic and widely used technique. DVR methods typically perform dense sampling of volume data along the viewpoint and accumulate the color and opacity of the sampled points to generate the final rendered image. However, the computational complexity of DVR is highly dependent on the volume data resolution and sampling density. As the volume data size increases, its rendering efficiency significantly decreases, making it difficult to meet the demands of real-time interaction and large-scale data visualization.
[0004] To improve rendering efficiency, recent research has focused on 3D scene representation methods based on explicit radiative representation. Among these, 3D Gaussian Splatting (3DGS) has achieved good results in novel perspective synthesis and real-time rendering by modeling 3D scenes using a set of sparse Gaussian primitives. Some studies attempt to convert volumetric data into multi-view images and then estimate the corresponding 3D Gaussian parameters through optimization algorithms, thereby achieving 3DGS-based volumetric data visualization. These methods reduce the computational overhead of the rendering stage to some extent and improve display performance.
[0005] On the other hand, feedforward 3D representation learning methods have made significant progress in surface reconstruction and novel viewpoint synthesis in recent years. The VGGT model can predict explicit 3D representations with a single network forward inference, thus avoiding scene-by-scene optimization. However, most existing feedforward methods are based on the assumption of a one-to-one correspondence between surface geometry and pixels to 3D points, making them difficult to apply directly to volumetric data visualization scenarios. Since a single pixel in volumetric rendering often corresponds to the cumulative contribution of multiple spatial locations along the viewing direction, the imaging mechanism of volumetric data differs fundamentally from that of surface scenes, making it difficult for existing feedforward methods to accurately represent volumetric structures and ensure multi-view accuracy. Figure 1 To the point of being responsive.
[0006] Therefore, existing technologies in volumetric data visualization still generally suffer from problems such as limited rendering efficiency, high preprocessing costs, and difficulty in representing the internal structure of the volume. There is still a lack of a technical solution that can efficiently and accurately convert volumetric data directly into a real-time renderable 3D representation without scene-by-scene optimization. Summary of the Invention
[0007] To address the aforementioned technical issues, this invention provides a method for rapidly constructing Gaussian representations of volume rendering based on vision-volume data alignment. By combining multi-view information with geometric constraints of volume data, it achieves a direct mapping from volume data to a 3D Gaussian sputtering representation, avoiding the scene-by-scene optimization process.
[0008] First, the framework directly predicts 3D Gaussian representations from volume data and its multi-view information through a feedforward neural network structure, thereby significantly reducing computational overhead and improving construction efficiency, solving the problem of low efficiency in traditional optimization methods.
[0009] Secondly, the aforementioned conversion framework from volume data to 3D Gaussian representation introduces a volume geometry enforcement mechanism. Through an epipolar constraint-based alignment of 2D and 3D features, multi-view appearance information is effectively internalized into the 3D Gaussian representation. This allows the introduction of 3D geometric priors from the volume data without relying on surface assumptions, ensuring multi-view... Figure 1 It accurately expresses volume rendering characteristics and solves the problem that existing feedforward methods have difficulty handling volume structures.
[0010] Extensive experiments were conducted in various volume visualization scenarios, and the present invention outperforms existing direct volume rendering methods and optimized 3D Gaussian representation methods in terms of rendering efficiency and visual quality.
[0011] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: A fast method for constructing Gaussian representations of volume rendering based on vision-volume data alignment includes: Obtain the 3D volume data to be visualized and the corresponding multi-view images; The three-dimensional volume data and multi-view images are input into a feedforward neural network. The three-dimensional Gaussian sputtering representation of the three-dimensional volume data is obtained by direct mapping through the feedforward neural network. The three-dimensional Gaussian sputtering representation is composed of parameters of multiple anisotropic three-dimensional Gaussian elements. The feedforward neural network includes: a dual Transformer network for jointly modeling the two-dimensional appearance information of multi-view images and the three-dimensional geometric information of three-dimensional volume data; and a volume geometry enforcement mechanism for aligning the two-dimensional appearance information and the three-dimensional geometric information based on epipolar geometric constraints and injecting them into the three-dimensional Gaussian splash representation, so that the three-dimensional Gaussian splash representation internalizes multi-view geometric consistency.
[0012] In one embodiment, the step of inputting the 3D volume data and multi-view images into a feedforward neural network, and directly mapping the 3D Gaussian sputtering representation of the 3D volume data through the feedforward neural network, specifically includes: 3D volume data and A collection of images for each view The input is fed into the feedforward neural network. The output is from A three-dimensional Gaussian spatter representation composed of the parameters of anisotropic three-dimensional Gaussian elements: ; in, Indicates the first The Gaussian center position, opacity, rotation, scale, and color of a 3D Gaussian element.
[0013] In one embodiment, the dual Transformer network, used to jointly model the two-dimensional appearance information of multi-view images and the three-dimensional geometric information of three-dimensional volume data, specifically includes: The dual Transformer network includes an appearance Transformer and a geometry Transformer; Appearance Transformer: Images for each viewpoint were processed using the DINOv2 encoder. Extract image features and combine them with learnable auxiliary camera features. After concatenation, the input is the appearance Transformer, which consists of a multi-layer attention network. The appearance Transformer outputs multi-scale two-dimensional appearance features at different layers to obtain multi-layer, multi-scale two-dimensional appearance features. Camera parameters Explicitly inject the appearance Transformer using the camera encoder. Generate embeddings and initialize convolutions with zero. injection: ; To enhance camera features, they are injected into the appearance Transformer to guide cross-view appearance feature aggregation; Geometric Transformer: Explicitly initialize a set of volume Gaussians for the 3D volume data that can collectively accumulate contributions along the rays. : ; Indicates the first An initial Gaussian body, Indicates the first The initialization parameters for a 3D Gaussian element are: Gaussian center position, opacity, rotation, scale, and color. To initialize the number of Gaussians in the volume; PTV3 point cloud encoder adopted , initialize the Gaussian body The data is treated as a point cloud and encoded to obtain 3D geometric features, which are then processed by a Gaussian output head. Decode 3D geometric features into 3D Gaussian meta-parameters: ; The Gaussian center position prediction residual, opacity, rotation, scale, and color of the k-th 3D Gaussian element are represented. With the initial Gaussian center The summation yields the final Gaussian center position.
[0014] In one embodiment, the volume geometry forcing mechanism is used to align two-dimensional appearance information with three-dimensional geometric information based on epipolar geometry constraints and inject it into the three-dimensional Gaussian sputtering representation, specifically including: Multi-layer, multi-scale two-dimensional appearance features extracted based on appearance Transformer; For the 3D geometric features output by the geometry Transformer, the 3D position of the 3D geometric features is projected onto all views using known camera parameters, and the corresponding 2D features are sampled from multi-layer, multi-scale 2D appearance features at the projection points. The sampling two-dimensional features are injected into the corresponding three-dimensional geometric features through the epipolar cross-attention mechanism to achieve alignment between two-dimensional appearance information and three-dimensional geometric information.
[0015] In one embodiment, the extraction of multi-layer, multi-scale two-dimensional appearance features from multiple levels of the appearance Transformer specifically includes: Two-dimensional appearance features are obtained from the outputs of multiple layers of the appearance Transformer, and multi-scale features are generated through a multi-scale DPT feature decoding head. , Indicates scale level. This indicates the attention layer index.
[0016] In one embodiment, the step of injecting the sampled two-dimensional features into the corresponding three-dimensional geometric features through an epipolar cross-attention mechanism to achieve alignment between two-dimensional appearance information and three-dimensional geometric information specifically includes: ; in, This represents the aligned 3D geometric features. Indicates the number of multi-view images. The number of scales representing two-dimensional appearance features. The number of layers representing two-dimensional appearance features. Representing feature dimension, For the softmax function, Key features representing three-dimensional geometric features The key features and value features of the sampled two-dimensional features; Attention bias based on geometric distance: ; Representing three-dimensional geometric features and the first Geometric distance measurement between views This represents hyperparameters.
[0017] In one embodiment, the 3D Gaussian spatter representation output by the feedforward neural network is rendered using a 3D Gaussian spatter renderer to obtain a multi-view image; when training the feedforward neural network based on the rendered multi-view image and the original multi-view image, a combination of pixel-level and perceptual-level loss is used. Supervise the rendering results: ; The loss is used to constrain pixel error. For the loss function used to constrain structural similarity, This is the loss used to constrain perceived consistency.
[0018] Compared with the prior art, the beneficial technical effects of the present invention are: 1. This invention proposes a forward Gaussian generative framework (VVGT) based on volume geometry forcing, which constructs a dual Transformer network to simultaneously process multiple views. Figure 2 This method jointly models 3D appearance information and volume data geometric priors, and directly predicts sparse 3D Gaussian representations using volume data and a limited number of multi-view images as input. This avoids the problems of traditional volume rendering and 3D Gaussian splashing methods that rely on scene-by-scene iterative optimization or long training times. This method can significantly improve inference efficiency while ensuring high-quality volume rendering results, and is suitable for interactive and explorable volume scene visualization tasks, solving the shortcomings of existing methods in terms of efficiency and scalability.
[0019] 2. This invention proposes a Volume Geometry Forcing (VGF) mechanism to explicitly guarantee the geometric consistency of multiple views. Unlike existing methods that simply aggregate 2D multi-view features or rely on voxels and dense 3D representations as underlying geometric constraints, this application utilizes sparse 3D Gaussian representations as the underlying 3D structure and introduces a cross-attention mechanism based on epipolar geometric constraints to ensure geometric consistency across multiple views. Figure 2 The method forces the alignment of 3D appearance features and injects them into the 3D Gaussian representation, enabling the network to explicitly perform multi-view geometric consistency modeling. While ensuring geometric consistency, this method is memory-friendly and computationally efficient, capable of handling large-scale volumetric data scenes, and applicable to volumetric visualization and new perspective rendering tasks under different numbers of viewpoints and resolutions. It effectively solves the problem of simultaneously ensuring geometric and appearance consistency in multi-view volumetric scenes. Attached Figure Description
[0020] Figure 1 This is a structural diagram of the Transformer feedforward neural network of the present invention.
[0021] Figure 2 A qualitative comparison chart of the rendering quality of volume data converted to 3D Gaussian using different methods. Detailed Implementation
[0022] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.
[0023] This invention proposes a VisualVolume-Grounded Transformer (VVGT) with geometric constraints on visual-volume data. This framework can directly map input volume data and its multi-view information into a real-time renderable 3D Gaussian representation without scene-by-scene optimization, making it suitable for various volume data visualization scenarios.
[0024] Specifically, the technical solution of this application includes the following aspects: First, a feedforward 3D representation prediction framework is constructed, which directly predicts 3D Gaussian splash representation from volume data and its multi-view observations through neural networks, thereby avoiding the traditional scene-by-scene optimization process, significantly reducing computational costs and improving construction efficiency.
[0025] Secondly, a volume geometry constraint mechanism is introduced to address the multi-view... Figure 2 The appearance information is aligned and fused with the 3D Gaussian representation, so that the multi-view information can be effectively internalized into the 3D representation. The 3D structural prior of the volume data is introduced without relying on surface assumptions, thereby ensuring the consistency of multi-view rendering results and the accurate expression of volume structure.
[0026] Finally, through joint modeling of the overall framework, while ensuring the realism of the rendering, structural consistency and efficiency requirements are also taken into account, so as to achieve high-quality, scalable and interactive visualization of volume data.
[0027] The overall network structure of this application is as follows: Figure 1 As shown. The purpose of this application is to train a feedforward volume data visualization network model that, through a single network forward inference, directly maps input volume data and its multi-view image information into a real-time renderable 3D Gaussian spatter (3DGS) representation, thereby achieving high-quality, interactive volume data visualization. This application proposes a Visual Volume-Grounded Transformer (VVGT) feedforward neural network with visual-volume data geometric constraints, which mainly includes two core components: Dual-Transformer Network (DTN): Dual-Transformer Networks are used to jointly model two-dimensional multi-view appearance information and three-dimensional volume geometry information.
[0028] Volume Geometry Forcing (VGF): The volume geometry forcing mechanism is used to forge multiple views. Figure 2 The dimensional information is internalized into the 3D Gaussian representation to enable the initial Gaussian to learn accurate properties and improve visualization quality.
[0029] Among them, attention network, DINOv2 encoder, DPT (Dense Prediction Transformer) feature decoding head, PTV3 point cloud encoder and 3D Gaussian spatter can all be regarded as existing technology modules, so they will not be described in detail.
[0030] This application combines the above modules to construct a feedforward mapping from volume data to 3D Gaussian sputtering and its 2D-3D alignment mechanism. It does not require improvements to the internal structure of the existing modules, so its internal implementation details will not be elaborated here.
[0031] The key method modules of this application are described below: 1. Problem definition and output parameter format.
[0032] Given volume data and A set of input images corresponding to a calibrated camera This application trains a feedforward neural network. Output A set of parameters for anisotropic 3D Gaussian primitives is used to represent the scene geometry and appearance, and their mapping relationship is as follows: ; in, The Gaussian center position, opacity, rotation, scale, and color of the k-th 3D Gaussian element are represented. The output set of 3D Gaussian elements can be directly used for high-quality volumetric scene rendering from a new perspective without the need for scene-by-scene iterative optimization.
[0033] 2. Dual Transformer Network (DTN).
[0034] (1) Appearance Transformer: This invention provides images from each viewpoint. Image features were extracted using a DINOv2 encoder and combined with learnable auxiliary camera features. After concatenation, the input is the appearance Transformer, which consists of a multi-layer attention network. The appearance Transformer outputs multi-scale two-dimensional appearance features at different layers to obtain multi-layer, multi-scale two-dimensional appearance features. Camera parameters Explicitly inject the appearance Transformer using the camera encoder. Generate embeddings and initialize convolutions with zero. injection: ; here, To enhance camera features, they are injected into the appearance Transformer to guide cross-view appearance feature aggregation.
[0035] (2) Geometric Transformer: To overcome the surface assumptions commonly found in existing forward 3D Gaussian sputtering methods, this application explicitly initializes a set of volume Gaussian sputtering data that can collectively accumulate contributions along the rays. : .
[0036] Indicates the first An initial Gaussian body, Indicates the first The initialization parameters for a 3D Gaussian element are: Gaussian center position, opacity, rotation, scale, and color. To initialize the number of Gaussians in the volume.
[0037] In one embodiment, an initial volume Gaussian set can be generated based on a structured sampling strategy for volume data (e.g., sampling candidate primitives in the wavelet transform domain or multi-scale domain). This initialization provides reliable spatial locations, but its properties still need to be refined through network prediction, thus avoiding the high cost of traditional scene-by-scene optimization.
[0038] Secondly, a PTV3 point cloud encoder is used. Initialize each Gaussian It is treated as a point primitive and its three-dimensional geometric features are obtained. Then, it is processed through a Gaussian output head. Decode the three-dimensional geometric features into three-dimensional Gaussian meta-parameters: ; The Gaussian center position prediction residual, opacity, rotation, scale, and color of the k-th 3D Gaussian element are represented. With the initial Gaussian center The summation yields the final Gaussian center position.
[0039] 3. Volumetric Geometric Forcing Mechanism (VGF).
[0040] To enable multi-view Figure 2 To effectively internalize 2D appearance information into sparse volumetric Gaussian representation, this application proposes a volumetric geometry forcing mechanism. Through cross-attention of epipolar geometric constraints, 2D features are aligned and injected into 3D geometric features, thereby improving the accuracy of Gaussian attribute learning and rendering fidelity.
[0041] (1) Multi-scale two-dimensional feature extraction: Two-dimensional appearance features are obtained from the outputs of multiple layers of the appearance Transformer, and multi-scale features are generated through a multi-scale DPT feature decoding head. ,in, Indicates scale level. This indicates the attention layer index.
[0042] In one embodiment, the following settings are provided: Each scale The system consists of several levels, forming a feature pyramid that simultaneously contains deep semantic information and shallow detail information, in order to more accurately constrain the volume geometry.
[0043] (2) Epipolar Cross-Attention Mechanism: For each 3D geometric feature output by the 3D Geometry Transformer The system projects its 3D position onto all input views using known camera parameters, samples corresponding features from multi-scale 2D feature maps at the projection points, and constructs cross-attention mechanisms to achieve 2D-3D aligned injection. ; in, This represents the aligned 3D geometric features. Indicates the number of multi-view images. The number of scales representing two-dimensional appearance features. The number of layers representing two-dimensional appearance features. Representing feature dimension, For the softmax function, Key features representing three-dimensional geometric features The key features and value features of the sampled two-dimensional features.
[0044] Attention bias based on geometric distance: ; Representing three-dimensional geometric features and the first A geometric distance metric between views is used to suppress contributions from geometrically irrelevant or distant views, directing attention more towards visually and spatially relevant observations. This represents hyperparameters.
[0045] In one embodiment, to reduce computational and storage overhead, the epipolar attention is performed only at a predetermined layer (e.g., the second layer) of the 3D geometric Transformer, and subsequent layers continue to fuse the injected information through the propagation and aggregation capabilities of the Transformer.
[0046] 4. Loss function.
[0047] Multiview Injection via Geometric Forced Mechanism Figure 2 After obtaining the appearance information, the feedforward neural network outputs a refined 3DGS set, which is then rendered using a 3DGS renderer to obtain a multi-view image. During training, a combination of pixel-level and perceptual-level loss is used to supervise the rendering results. The overall objective function is: ; in, Used to constrain pixel errors Used to constrain structural similarity, Used to constrain perceived consistency; , where is the weighting coefficient. This combined supervision takes into account photometric accuracy, structural consistency, and perceptual fidelity, thereby guiding the network to learn accurate volume geometry and appearance under forward inference conditions.
[0048] To verify the effectiveness of the method (VVGT) in volumetric scene visualization and novel perspective rendering tasks, this embodiment uses different types of volumetric datasets to construct training and testing scenarios, and performs quantitative evaluation on multi-view rendering metrics. This embodiment selects two types of large-scale temperature field or turbulence field volumetric data generated by direct numerical simulation (DNS): 1. Temperature field of rotating stratified turbulence, with a resolution of [resolution value missing]. ; 2. Isotropic turbulence simulation data, resolution: .
[0049] Multiple data points were obtained by cropping from the above volume data. The sub-scenes are used as independent samples. Specifically, 100 sub-scenes are selected for training and 10 sub-scenes are used for testing. For each sub-scene, 10 basic transfer functions (TFs) are designed to highlight different scalar value ranges and internal structures, and to adjust the resolution. The resulting multi-view images are then rendered and used as supervisory data. Rendering tools can include ParaView and NVIDIA IndeX.
[0050] To evaluate the rendering quality of the new perspective, this embodiment calculates the following metrics between the rendered image and the real image: PSNR (Peak Signal-to-Noise Ratio), SSIM (Structural Similarity Index), and LPIPS (Learned Perceptual Similarity). PSNR and SSIM are used to measure photometric or structural similarity, while LPIPS measures perceptual quality.
[0051] For a comprehensive evaluation, this embodiment compares the method of this application with two representative methods: (1) Methods based on reverse iterative optimization: including 3DGS and iVRGS, which require iterative optimization for each scene; (2) Feedforward neural network-based method: AnySplat[3], direct forward prediction Gaussian representation.
[0052] The VVGT method in this application belongs to the forward prediction framework and further provides optional post-optimization steps to further improve rendering quality while maintaining efficiency advantages.
[0053] Table 1. Quantitative comparison results of volume data to 3D Gaussian rendering quality
[0054] On the zero-sample test dataset, evaluations were conducted with 24-view and 36-view settings, and the results are shown in Table 1. As can be seen from Table 1, compared to other forward methods (such as NoPoSplat and AnySplat), the method of this invention (VVGT) significantly outperforms other methods in terms of PSNR, SSIM, and LPIPS, indicating a clear advantage in volumetric scene rendering quality. While maintaining a second-level inference time, VVGT achieves rendering quality close to that of optimization-based methods (such as 3DGS and iVRGS). Performing a small amount of post-optimization on top of VVGT, such as 100 steps (VVGT_100) or 200 steps (VVGT_200), can further improve rendering quality, reaching or exceeding optimization-based methods in terms of PSNR, SSIM, and LPIPS, while maintaining a significantly lower overall time consumption than traditional scene-by-scene long-duration optimization processes. Figure 2 A qualitative comparison of the rendering quality of volume data converted to 3D Gaussian using different methods.
[0055] The present invention presents a method for rapid construction of Gaussian representations for volume rendering based on visual-volume data alignment. This method avoids scene-by-scene optimization and significantly reduces the preprocessing cost of volume data visualization. It replaces traditional direct volume rendering methods at the representation level, improving rendering efficiency and scalability. By introducing a dual Transformer network and a volume data geometric constraint mechanism, it efficiently utilizes the inherent structural information of the volume data and prior knowledge of the volume rendering line integral, effectively achieving representation-level transformation. This results in fewer constructed Gaussian spheres and higher rendering quality. Consequently, it enables high-quality, real-time, and interactive volume data visualization, suitable for various applications such as medical imaging and scientific simulation. Qualitative and quantitative experiments both verify that the proposed method has significant advantages over existing technologies in terms of rendering quality and efficiency.
[0056] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0057] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple steps or stages, which are not necessarily completed at the same time, but may be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but may be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0058] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0059] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.
[0060] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A fast method for constructing Gaussian representations of volume rendering based on visual-volume data alignment, characterized in that, include: Obtain the 3D volume data to be visualized and the corresponding multi-view images; The three-dimensional volume data and multi-view images are input into a feedforward neural network. The three-dimensional Gaussian sputtering representation of the three-dimensional volume data is obtained by direct mapping through the feedforward neural network. The three-dimensional Gaussian sputtering representation is composed of parameters of multiple anisotropic three-dimensional Gaussian elements. The feedforward neural network includes: a dual Transformer network for jointly modeling the two-dimensional appearance information of multi-view images and the three-dimensional geometric information of three-dimensional volume data; and a volume geometry enforcement mechanism for aligning the two-dimensional appearance information and the three-dimensional geometric information based on epipolar geometric constraints and injecting them into the three-dimensional Gaussian splash representation, so that the three-dimensional Gaussian splash representation internalizes multi-view geometric consistency.
2. The method for fast construction of Gaussian representation of volume rendering based on visual-volume data alignment according to claim 1, characterized in that, The process of inputting 3D volume data and multi-view images into a feedforward neural network, and directly mapping the feedforward neural network to obtain a 3D Gaussian sputtering representation for the 3D volume data, specifically includes: 3D volume data and A collection of images for each view The input is fed into the feedforward neural network. The output is from A three-dimensional Gaussian spatter representation composed of the parameters of anisotropic three-dimensional Gaussian elements: ; in, Indicates the first The Gaussian center position, opacity, rotation, scale, and color of a 3D Gaussian element.
3. The method for fast construction of Gaussian representation of volume rendering based on visual-volume data alignment according to claim 1, characterized in that, The dual Transformer network is used to jointly model the two-dimensional appearance information of multi-view images and the three-dimensional geometric information of three-dimensional volume data, specifically including: The dual Transformer network includes an appearance Transformer and a geometry Transformer; Appearance Transformer: Images for each viewpoint were processed using the DINOv2 encoder. Extract image features and combine them with learnable auxiliary camera features. After concatenation, the input is the appearance Transformer, which consists of a multi-layer attention network. The appearance Transformer outputs multi-scale two-dimensional appearance features at different layers to obtain multi-layer, multi-scale two-dimensional appearance features. Camera parameters Explicitly inject the appearance Transformer using the camera encoder. Generate embeddings and initialize convolutions with zero. injection: ; To enhance camera features, they are injected into the appearance Transformer to guide cross-view appearance feature aggregation; Geometric Transformer: Explicitly initialize a set of volume Gaussians for the 3D volume data that can collectively accumulate contributions along the rays. : ; Indicates the first An initial Gaussian body, Indicates the first The initialization parameters for a 3D Gaussian element are: Gaussian center position, opacity, rotation, scale, and color. To initialize the number of Gaussians in the volume; PTV3 point cloud encoder adopted , initialize the Gaussian body The data is treated as a point cloud and encoded to obtain 3D geometric features, which are then processed by a Gaussian output head. Decode 3D geometric features into 3D Gaussian meta-parameters: ; The Gaussian center position prediction residual, opacity, rotation, scale, and color of the k-th 3D Gaussian element are represented. With the initial Gaussian center The summation yields the final Gaussian center position.
4. The method for fast construction of Gaussian representation of volume rendering based on visual-volume data alignment according to claim 1, characterized in that, The volume geometry forcing mechanism is used to align two-dimensional appearance information with three-dimensional geometric information based on epipolar geometric constraints and inject it into the three-dimensional Gaussian sputtering representation. Specifically, it includes: Multi-layer, multi-scale two-dimensional appearance features extracted based on appearance Transformer; For the 3D geometric features output by the geometry Transformer, the 3D position of the 3D geometric features is projected onto all views using known camera parameters, and the corresponding 2D features are sampled from multi-layer, multi-scale 2D appearance features at the projection points. The sampling two-dimensional features are injected into the corresponding three-dimensional geometric features through the epipolar cross-attention mechanism to achieve alignment between two-dimensional appearance information and three-dimensional geometric information.
5. The method for fast construction of Gaussian representation of volume rendering based on visual-volume data alignment according to claim 4, characterized in that, The extraction of multi-layer, multi-scale two-dimensional appearance features from multiple levels of the appearance Transformer specifically includes: Two-dimensional appearance features are obtained from the outputs of multiple layers of the appearance Transformer, and multi-scale features are generated through a multi-scale DPT feature decoding head. , Indicates scale level. This indicates the attention layer index.
6. The method for fast construction of Gaussian representation of volume rendering based on visual-volume data alignment according to claim 4, characterized in that, The process of injecting the sampled two-dimensional features into the corresponding three-dimensional geometric features through an epipolar cross-attention mechanism to align two-dimensional appearance information with three-dimensional geometric information specifically includes: ; in, This represents the aligned 3D geometric features. Indicates the number of multi-view images. The number of scales representing two-dimensional appearance features. The number of layers representing two-dimensional appearance features. Representing feature dimension, For the softmax function, Key features representing three-dimensional geometric features The key features and value features of the sampled two-dimensional features; Attention bias based on geometric distance: ; Representing three-dimensional geometric features and the first Geometric distance measurement between views This represents hyperparameters.
7. The method for fast construction of Gaussian representation of volume rendering based on visual-volume data alignment according to claim 1, characterized in that, The 3D Gaussian splash representation output by the feedforward neural network is rendered using a 3D Gaussian splash renderer to obtain multi-view images. When training the feedforward neural network based on the rendered multi-view images and the original multi-view images, a combination of pixel-level and perceptual-level loss is employed. Supervise the rendering results: ; The loss is used to constrain pixel error. For the loss function used to constrain structural similarity, This is the loss used to constrain perceived consistency.