Embedded 3d gaussian scene reconstruction method based on multi-view saliency supervision
By introducing saliency supervision information into 3D Gaussian scene reconstruction, the mapping problem of saliency attributes in 3D Gaussian scene representation is solved, and joint optimization of saliency attributes with geometric and appearance attributes is achieved, thereby improving the perception quality of 3D scenes and the efficiency of downstream applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI UNIV OF TECH
- Filing Date
- 2026-06-25
- Publication Date
- 2026-07-24
AI Technical Summary
Existing 3D Gaussian scene reconstruction methods lack technical solutions for effectively mapping 2D saliency supervision information to 3D Gaussian primitives and participating in joint optimization, resulting in limitations in perceptual optimization tasks.
An embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision is adopted. By acquiring multi-view images and saliency supervision distribution maps, a joint optimization objective of saliency attributes, geometric attributes, and appearance attributes is constructed. Saliency rendering and appearance rendering are then performed to achieve direct embedding of saliency attributes in the 3D Gaussian scene representation.
It improves the perceptual quality and downstream application efficiency of 3D Gaussian scene representation, optimizes transmission and rendering under bandwidth-constrained conditions, simplifies the subsequent saliency estimation process, and enhances the stability and continuity of visual reconstruction.
Smart Images

Figure CN122454070A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision, 3D scene reconstruction and immersive media processing technology, and specifically to an embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision. Background Technology
[0002] With the development of immersive applications such as virtual reality, augmented reality, and mixed reality, high-quality modeling and real-time rendering of 3D scenes have gradually become an important research direction in related fields. Volumetric video, as a media form capable of representing the 3D structure and dynamic content of real scenes, has broad application prospects in scenarios such as digital humans, immersive communication, remote interaction, and digital twins. Compared to traditional 3D representation methods such as dynamic meshes, point clouds, and neural radiation fields, 3D Gaussian Splatting (3DGS) uses a set of Gaussian primitives to represent the scene, balancing high reconstruction quality, fast rendering speed, and good real-time performance, thus gradually becoming an important technical route for 3D scene representation and reconstruction.
[0003] In existing technologies, 3DGS typically uses Gaussian primitives to model scenes based on attributes such as position, scaling, rotation, opacity, and color, and achieves fitting and reconstruction of multi-view images through differentiable rendering. However, with the increasing application of 3D scenes in tasks such as compression, transmission, progressive loading, and perceptually optimized rendering, relying solely on geometric and appearance information is no longer sufficient to optimize system performance based on viewer behavior. In real-world user viewing experience scenarios, different areas of a scene attract different levels of visual attention. Therefore, directly introducing salient information into the 3D scene representation can help prioritize the preservation and transmission of more important content under conditions of limited bitrate or resources, thereby improving perceptual quality.
[0004] Some existing methods attempt to indirectly estimate the importance of Gaussian elements through rendering results from sampled perspectives, such as assessing their importance based on their visibility, contribution, or pruning effect across multiple views. While these methods can support compression or transmission optimization to some extent, they are typically post-processing importance estimations, and their saliency or importance is not considered an intrinsic property of the scene representation itself during the reconstruction process. Furthermore, these methods largely rely on manually designed heuristics, making them susceptible to the influence of sampling perspectives, estimation strategies, and approximate assumptions. They struggle to accurately reflect the distribution of real user visual attention, thus often exhibiting limitations in perceptual optimization tasks.
[0005] Furthermore, since 2D image saliency prediction technology is relatively mature, it is possible to obtain saliency distribution maps corresponding to multi-view images. Considering that 3DGS reconstruction is essentially a process of optimizing 3D Gaussian primitives based on multi-view 2D image supervision, the saliency distribution map can be introduced as additional supervision information into the reconstruction stage to directly construct a 3D Gaussian scene representation containing saliency attributes. However, existing technologies still lack a technical solution for effectively mapping 2D saliency supervision information to 3D Gaussian primitives and participating in joint optimization. In particular, existing technologies are still insufficient in areas such as the mapping of saliency supervision to Gaussian primitives, the parameterized expression of saliency attributes, the collaborative rendering of color images and saliency distribution maps, and the joint optimization of saliency attributes with geometric and appearance attributes.
[0006] Therefore, it is necessary to provide a saliency-embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision, so as to explicitly introduce saliency attributes into the 3D Gaussian scene representation, so that saliency information can be optimized together with other attributes in the scene reconstruction stage, thereby providing a more effective 3D representation basis for subsequent perceptual optimization compression, transmission and rendering. Summary of the Invention
[0007] The embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision proposed in this invention can at least solve one of the technical problems in the background art.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: An embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision includes the following steps: S100. Acquire reconstruction scene data and establish a unified data foundation; S200. Based on a unified data foundation, construct Gaussian primitive representations with salient attributes; S300, based on Gaussian primitive representation, performs appearance rendering and saliency rendering; S400, based on appearance rendering and saliency rendering, constructs a joint optimization objective of saliency attributes, geometric attributes, and appearance attributes, and iteratively reconstructs the scene.
[0009] Furthermore, the method for establishing a unified data foundation in step S100 of the present invention includes: Obtain multi-view image sequences corresponding to the scene to be reconstructed Camera parameters and the saliency supervision distribution map corresponding to each viewpoint image. ; in, Represents the number of viewpoints. Indicates the first Color images from various perspectives This represents the saliency supervision distribution plot corresponding to this perspective. The camera parameters representing this viewpoint; The camera parameters include an intrinsic parameter matrix. and extrinsic parameter matrix ,in Represents the rotation matrix. Represents the translation vector; The first The projection model of each viewpoint is represented as:
[0010] in, For the first The three-dimensional center coordinates of each Gaussian element. For its in the The projection center on the image plane of each viewpoint.
[0011] Furthermore, the Gaussian meta-representation method with saliency attributes in this invention includes: Represent the scene as A set of Gaussian elements , among which, the Gao Siyuan Defined as:
[0012] in, For position parameters, For scale parameters, For rotation parameters, For opacity parameters, For appearance parameters, For saliency coding parameters; Saliency coding parameters ,in Preset dimensions; No. The covariance matrix of the Gaussian elements in three-dimensional space is represented as:
[0013] in, Indicated by rotation parameters The corresponding rotation matrix, These represent the scale parameters of the Gaussian element in the x-axis, y-axis, and z-axis directions, respectively. In the From each viewpoint, the two-dimensional covariance matrix of the Gaussian element projected onto the image plane is expressed as:
[0014] in, This represents the Jacobian matrix of the 3D-to-2D projection near the center of Gauss.
[0015] Furthermore, the appearance rendering and salience rendering methods in step S300 of the present invention include: For any pixel on the image plane , No. The first Gaussian Yuan in the The two-dimensional response of this pixel from each viewpoint is represented as follows:
[0016] In the formula, Let Gausky's elements be the two-dimensional covariance matrix projected onto the image plane. For the first Projection model from multiple perspectives; Combined with opacity parameter Its effect on pixels The pixel-level contribution is represented as:
[0017] If the Gaussian elements participating in the rendering of the same pixel are sorted according to depth from near to far, then the... The cumulative transmittance of each Gaussian pixel at this pixel location is expressed as:
[0018] Therefore, the first The color image rendering results from each viewpoint are represented as follows:
[0019] in, Indicates the first The first Gaussian Yuan in the Color output from various perspectives, determined by its appearance parameters Calculated by combining the viewpoint direction; A saliency rendering process is constructed to achieve the micro-projection of saliency attributes from 3D Gaussian primitives to the 2D image plane. For saliency rendering, saliency is first encoded. The mapping is to a saliency output, which is represented in the following form: ; in, The saliency mapping function can be a fully connected linear classifier. ; The rendering result of the significance distribution plot is then expressed as follows: .
[0020] Furthermore, the method for constructing the joint optimization objective of saliency attributes, geometric attributes, and appearance attributes in step S400 of the present invention includes: Image reconstruction loss is expressed as:
[0021] in, The image reconstruction error function is represented in the following form:
[0022] in, and For balance coefficient, express Norm, Indicators representing structural similarity; Significance-based supervised loss is expressed as:
[0023] in, and These represent the image height and width, respectively. Indicates the significance level number. Represents cross-entropy loss; set up Indicates the first A high-ranking official Nearest neighbor set Indicates the first The significance distribution corresponding to each Gaussian element is then expressed as:
[0024] in, This represents the number of sampled Gaussian elements involved in the calculation of three-dimensional consistency constraints. Indicates KL divergence; The overall optimization objective is expressed as:
[0025] in, and These are the weight coefficients for the significance supervision loss term and the three-dimensional neighborhood consistency constraint term, respectively.
[0026] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0027] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0028] As can be seen from the above technical solution, the embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision of the present invention introduces saliency encoding parameters, saliency rendering branches and joint optimization objectives in the 3D Gaussian scene reconstruction stage, so that the obtained 3D Gaussian scene representation can directly carry stable saliency attributes; these saliency attributes can be further used for priority scheduling under bandwidth-constrained conditions and significantly improve the perception reconstruction quality of the region within the viewport. Attached Figure Description
[0029] Figure 1 This is a flowchart of the embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision of the present invention; Figure 2 This is an overall flowchart of the saliency-embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision, which is an embodiment example of the present invention. Figure 3 This is a schematic diagram of the perceptual optimization transmission experiment of the 3DGS with embedded saliency attributes constructed in Example 1 of this invention to verify the present invention. Figure 4 This is a comparison diagram of the actual application effects of Example 1 of the present invention. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0031] like Figure 1 As shown in this embodiment, the embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision includes the following steps: S100. Acquire reconstruction scene data and establish a unified data foundation; S200. Based on a unified data foundation, construct Gaussian primitive representations with salient attributes; S300, based on Gaussian primitive representation, performs appearance rendering and saliency rendering; S400, based on appearance rendering and saliency rendering, constructs a joint optimization objective of saliency attributes, geometric attributes, and appearance attributes, and iteratively reconstructs the scene.
[0032] The following provides a detailed explanation of each step: S100. Acquire reconstruction scene data and establish a unified data foundation; Obtain multi-view image sequences corresponding to the scene to be reconstructed Camera parameters and the saliency supervision distribution map corresponding to each viewpoint image. .
[0033] in, Represents the number of viewpoints. Indicates the first Color images from various perspectives This represents the saliency supervision distribution plot corresponding to this perspective. The camera parameters represent the viewpoint. These camera parameters include an intrinsic parameter matrix. and extrinsic parameter matrix etc., among which Represents the rotation matrix. This represents the translation vector.
[0034] The first The projection model of each viewpoint is represented as:
[0035] in, For the first The three-dimensional center coordinates of each Gaussian element. For its in the The projection center on the image plane of each viewpoint.
[0036] By constructing the above input data, a unified data foundation can be established for subsequent geometric projection, image rendering, and saliency supervision of 3D Gaussian primitives.
[0037] S200. Based on a unified data foundation, construct Gaussian primitive representations with salient attributes; Represent the scene as A set of Gaussian elements , among which, the Gao Siyuan Defined as:
[0038] in, For position parameters, For scale parameters, For rotation parameters, For opacity parameters, For appearance parameters, For saliency coding parameters.
[0039] Saliency coding parameters ,in This is a preset dimension, which can be set according to specific application requirements. In one implementation, the saliency encoding parameter... Instead of being used directly as the final significance value, it is used as a trainable latent representation, which is transformed into a significance output through a significance mapping function, thereby enhancing the expressive power and optimization flexibility of significance modeling.
[0040] No. The covariance matrix of the Gaussian elements in three-dimensional space is represented as:
[0041] in, Indicated by rotation parameters The corresponding rotation matrix, These represent the scale parameters of the Gaussian element in the x-axis, y-axis, and z-axis directions, respectively.
[0042] In the From each viewpoint, the two-dimensional covariance matrix of the Gaussian element projected onto the image plane can be expressed as:
[0043] in, This represents the Jacobian matrix of the 3D-to-2D projection near the center of Gauss.
[0044] By introducing saliency coding parameters on the basis of traditional three-dimensional Gaussian representation This invention enables each Gaussian unit to not only represent the geometric structure and appearance information of the scene, but also to explicitly carry visual saliency information, providing a representational basis for subsequent saliency rendering and joint optimization.
[0045] S300, based on Gaussian primitive representation, performs appearance rendering and saliency rendering; For any pixel on the image plane , No. The first Gaussian Yuan in the The two-dimensional response of this pixel from each viewpoint is represented as follows:
[0046] Combined with opacity parameter Its effect on pixels The pixel-level contribution is represented as:
[0047] If the Gaussian elements participating in the rendering of the same pixel are sorted according to depth from near to far, then the... The cumulative transmittance of each Gaussian pixel at this pixel location is expressed as:
[0048] Therefore, the first The color image rendering results from each viewpoint are represented as follows:
[0049] in, Indicates the first The first Gaussian Yuan in the Color output from each viewing angle can be determined by its appearance parameters. It was calculated by combining the perspective and direction.
[0050] This invention further constructs a saliency rendering process, thereby realizing the differentiable projection of saliency attributes from three-dimensional Gaussian units to a two-dimensional image plane. For saliency rendering, saliency is first encoded. The mapping is to the saliency output. It is represented in the following form:
[0051] in, The saliency mapping function can be a fully connected linear classifier. .
[0052] The rendering result of the significance distribution plot can be expressed as:
[0053] This allows us to construct a direct relationship from 3D Gaussian units to 2D saliency distributions, enabling the 2D saliency distribution map to directly supervise the generation of 3D Gaussian unit saliency attributes, thus achieving differentiable training and reconstruction of saliency attributes.
[0054] S400: Based on appearance rendering and saliency rendering, construct a joint optimization objective of saliency attributes, geometric attributes, and appearance attributes, and iteratively reconstruct the scene; To achieve joint optimization of saliency attributes, geometric attributes, and appearance attributes, this invention constructs an overall objective function consisting of image reconstruction loss, saliency supervision loss, and 3D consistency constraints, and completes saliency-embedded 3D Gaussian scene reconstruction by iteratively updating the parameters of each Gaussian meta-parameter.
[0055] (1) Image reconstruction loss Image reconstruction loss can be expressed as:
[0056] in, This represents the image reconstruction error function. In one implementation, it can take the following form:
[0057] in, and For balance coefficient, express Norm, This represents the structural similarity index. This loss term is used to ensure that the 3D Gaussian scene representation has good appearance reconstruction capabilities, and to avoid the introduction of saliency attributes from ruining the visual reconstruction quality of the original scene.
[0058] (2) Significance of monitoring loss The saliency-supervised loss is used to constrain the difference between the rendered saliency distribution map and the saliency-supervised distribution map, thereby driving the learning of Gaussian saliency attributes. The saliency-supervised loss can be expressed as:
[0059] in, and These represent the image height and width, respectively. Indicates the significance level number. This represents the cross-entropy loss.
[0060] (3) Three-dimensional neighborhood consistency constraint To improve the smoothness and stability of saliency attributes in three-dimensional space, a three-dimensional neighborhood consistency constraint is imposed on adjacent Gaussian elements. Let... Indicates the first A high-ranking official Nearest neighbor set Indicates the first The significance distribution corresponding to each Gaussian element can then be expressed as:
[0061] in, This represents the number of sampled Gaussian elements involved in the calculation of three-dimensional consistency constraints. This represents the KL divergence. This constraint allows spatially adjacent and geometrically correlated Gaussian elements to maintain a certain degree of consistency in their saliency attributes, thereby reducing local noise and oscillations in saliency attributes, and improving the continuity and robustness of the final reconstruction results in three-dimensional space.
[0062] (4) Overall objective function Accordingly, the overall optimization objective of the present invention can be expressed as:
[0063] in, and These are the weight coefficients for the significance supervision loss term and the three-dimensional neighborhood consistency constraint term, respectively.
[0064] During the optimization process, the geometric parameters, appearance parameters, opacity parameters, and saliency encoding parameters of all Gaussian elements are iteratively updated to make the color rendering results and saliency rendering results simultaneously approach the corresponding supervised targets, ultimately achieving joint learning and reconstruction of Gaussian element appearance attributes and visual saliency attributes.
[0065] In the 3D Gaussian scene representation reconstructed through the above steps, each Gaussian primitive not only possesses basic attributes such as position, shape, transparency, and appearance, but also a saliency attribute that directly characterizes its visual importance. This saliency attribute can directly serve downstream applications such as compression, transmission, rendering acceleration, and perceptual quality optimization, without requiring additional complex saliency estimation, importance ranking, or post-processing analysis, thereby improving overall processing efficiency and application practicality.
[0066] This embodiment provides a saliency-embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision. During the 3D Gaussian scene reconstruction process, saliency supervision information from the 2D image space is stably introduced into the Gaussian primitive parameter optimization process. This ensures that each Gaussian primitive possesses geometric and appearance attributes, as well as saliency attributes that are trainable, renderable, and directly applicable to downstream tasks. Through this method, a 3D Gaussian scene representation containing geometric, appearance, and saliency information can be obtained, providing a unified data foundation for subsequent applications such as perception-driven compression, saliency-aware transmission, viewpoint-first rendering, and resource allocation for important regions.
[0067] like Figure 2 As shown, in this embodiment, a scene with multi-view images and corresponding saliency supervision distribution maps is selected as input data. The scene is obtained by synchronously capturing images using a surround camera array and includes... Synchronized images from multiple perspectives, with an image resolution of [missing information]. Each viewpoint corresponds to a set of camera parameters. Including intrinsic parameter matrix and external references ,in Represents the camera rotation matrix. This represents the camera translation vector. Additionally, each viewpoint image also corresponds to a saliency supervision distribution map. In this embodiment, the saliency supervision distribution map can be generated by existing saliency detection methods, or obtained by manual annotation, eye-tracking fixation statistics, or other visual attention estimation methods. This embodiment preferably uses a saliency supervision map that strictly corresponds to the original image in spatial resolution to ensure the accuracy of pixel-level supervision.
[0068] The specific steps of this embodiment are as follows: S100. Acquire reconstruction scene data and establish a unified data foundation; Acquire multi-view image sequences of the scene to be reconstructed Camera parameter set and a set of significance-supervised distribution maps .
[0069] In this embodiment, to facilitate subsequent supervised training, the saliency supervision distribution maps for each viewpoint are first preprocessed. Specifically, the original saliency maps are divided into three saliency levels according to a preset threshold, so that each pixel only takes... A certain value is provided, where 0 represents a background or non-interested region, 1 represents a low-significance region, and 2 represents a high-significance region. Through the above discretization process, the continuous saliency distribution can be transformed into discrete category labels, thus facilitating subsequent stability supervision using the cross-entropy loss function.
[0070] Furthermore, to avoid inconsistencies in the scale of supervised distributions across different perspectives, the original saliency map can be normalized before discretization, for example, by normalizing the saliency values to a certain level. The intervals are then categorized into levels based on a set threshold. Simultaneously, the saliency map can be lightly smoothed to reduce the interference of local noise and isolated erroneous responses on the training process, thereby improving the robustness of the supervision signal.
[0071] After saliency map preprocessing, the scene is geometrically initialized using camera parameters. Specifically, the sparse point cloud and camera pose of the scene can be estimated first using 3D reconstruction methods such as structured bundle adjustment, SFM, or COLMAP. Based on this, parameters such as the center position, scale, rotation, color, and opacity of Gaussian units are initialized. In other words, this embodiment can directly reuse existing 3D Gaussian scene initialization processes, only adding a saliency-related branch, thereby avoiding excessive modification of the original reconstruction framework and exhibiting good compatibility and feasibility.
[0072] S200. Based on a unified data foundation, construct Gaussian primitive representations with salient attributes; The entire scene is represented as A set of three-dimensional Gaussian elements In this embodiment, the initial number of Gaussian elements can be set to... This value is for illustrative purposes only and can be adjusted in practice based on scene complexity, image resolution, and available computing resources. Each Gaussian unit is defined as:
[0073] in, For position parameters, For scale parameters, For rotation parameters, For opacity parameters, For appearance parameters, This is the saliency encoding parameter. In this embodiment, the dimension of the saliency encoding parameter is set to... ,Right now .
[0074] The saliency encoding parameters, serving as trainable implicit features for each Gaussian primitive (GRP), do not directly correspond to a fixed saliency value. Instead, they are automatically learned during training through a mapping function, learning their relationship with saliency supervision. Compared to directly assigning a single saliency scalar to each GRP, this approach has stronger expressive power and can more fully model the complex differences in spatial location, appearance structure, and saliency semantics among different GRPs. The three-dimensional covariance matrix of Gaussian elements is represented as:
[0075] in, Indicated by rotation parameters The corresponding rotation matrix, These represent the scales of the Gaussian element along the three principal axes. The covariance matrix describes the shape and orientation of the Gaussian element in three-dimensional space. For any viewpoint... Its projection center on the image plane is:
[0076] Accordingly, the two-dimensional covariance matrix of this Gaussian element on the image plane is:
[0077] in, Let be the Jacobian matrix of the projection mapping at the center of the Gaussian. Through this Jacobian transformation, the ellipsoidal Gaussian distribution in three-dimensional space can be approximately mapped to an elliptical Gaussian distribution on the two-dimensional image plane, thus obtaining a two-dimensional parameter representation that can be used for screen-space rasterization.
[0078] Furthermore, in this embodiment, appearance parameters Modeling can be performed using spherical harmonic coefficients to obtain direction-dependent color output based on the viewing angle. In other words, each Gaussian primitive carries not only geometric distribution information and appearance reflection information, but also additional saliency encoding information, thus forming a more complete 3D scene representation.
[0079] S300, based on Gaussian primitive representation, performs appearance rendering and saliency rendering; For the Any pixel in a viewpoint image , No. The two-dimensional Gaussian response of a pixel to a given Gaussian element is represented as follows:
[0080] The corresponding pixel-level transparency contribution is:
[0081] After sorting the Gaussian elements participating in the rendering of the same pixel from near to far according to depth, the first... Each Gaussian element in a pixel The cumulative transmittance at this location is:
[0082] Therefore, the first The color image rendering results from each viewpoint are as follows:
[0083] in, Indicates the first The color output of each Gaussian unit at the current viewpoint. Through the above forward mixing process, the color reconstruction result corresponding to the input image can be obtained, such as... Figure 3 As shown.
[0084] To further render the class probabilities of the 3D Gaussian elements into pixel-level saliency prediction results and facilitate the direct use of saliency attributes in subsequent downstream tasks, scalar saliency values of the Gaussian elements can be calculated based on the class probability distribution:
[0085] in, The saliency mapping function can be a fully connected linear classifier. Furthermore, the pixel-level scalar saliency map can be further represented as:
[0086] Thus, this embodiment obtains a pixel-level saliency distribution for supervised training and visualization. This design ensures both the rigor of the training objectives and the ease of use of the saliency attribute in practical applications.
[0087] In the above manner, the saliency attributes in the three-dimensional Gaussian primitives can be mapped to the two-dimensional image plane through the same projection and blending mechanism as color rendering, thereby establishing a direct connection between two-dimensional saliency supervision and three-dimensional saliency attributes, so that the two-dimensional supervision information can stably and inversely affect the three-dimensional Gaussian primitive parameter update process.
[0088] S400: Based on appearance rendering and saliency rendering, construct a joint optimization objective of saliency attributes, geometric attributes, and appearance attributes, and iteratively reconstruct the scene; To achieve joint optimization of appearance and saliency attributes, in this embodiment, the overall optimization objective consists of three parts: image reconstruction loss, saliency supervision loss, and three-dimensional neighborhood consistency constraint.
[0089] (1) The image reconstruction loss is defined as:
[0090] in:
[0091] In this embodiment, , .in, This term is used to constrain pixel-level color errors. The term is used to enhance structural similarity constraints, so that the reconstruction results have better overall visual quality while preserving detailed textures.
[0092] (2) The significance of supervised loss is defined as:
[0093] in, Represents the cross-entropy loss function. This represents the number of saliency levels. Using this loss term, the saliency encoding learning of the three-dimensional Gaussian units can be directly supervised using two-dimensional discrete saliency labels, enabling the Gaussian units to gradually acquire saliency attributes consistent with the true attention distribution.
[0094] (3) To improve the local smoothness and stability of salient attributes in three-dimensional space, a three-dimensional neighborhood consistency constraint is introduced:
[0095] in, This represents the number of sampled Gaussian elements participating in the consistency constraint calculation. Indicates the first A Gaussian element in three-dimensional space Nearest neighbor set This represents the Kullback-Leibler divergence. In this embodiment, random sampling... Each Gaussian element participates in the constraint calculation, and each Gaussian element selects the nearest element in its space. Each nearest neighbor is used as a neighborhood set. This consistency term can promote a certain consistency in the saliency distribution of Gaussian elements that are spatially close and have strong local structural correlations, thereby reducing isolated noise and unstable oscillations in saliency attributes.
[0096] Therefore, the overall loss function is expressed as:
[0097] In this embodiment, Of course, the weighting coefficients can be adjusted appropriately according to different scenarios and training stages. For example, in the early stage of training, the image reconstruction term can be appropriately strengthened to ensure convergence of basic geometry and appearance, and in the later stage of training, the role of the saliency term and the consistency term can be increased to further enhance the saliency embedding effect.
[0098] In this example, gradient descent is used to jointly optimize all Gaussian meta-parameters. The optimized parameters include Gaussian meta-parameter location parameters. Scale parameters Rotation parameters Opacity parameter Appearance parameters Significance coding parameters Significance mapping function The network parameters. In other words, this embodiment does not add an additional saliency prediction module after reconstruction, but rather allows the three types of attributes—geometric, appearance, and saliency—to participate in training and co-update within a unified differentiable framework.
[0099] In this embodiment, the optimizer is Adam, and the initial learning rate is set to... The total number of training iterations was set to 30,000. During training, a viewpoint was randomly selected from all viewpoints for rendering each time, and the corresponding image reconstruction loss, saliency supervision loss, and 3D consistency loss were calculated. Then, the parameters were updated through backpropagation. Although only one viewpoint was selected for training each time, the training process was set up so that all samples participated in one training round, thus ensuring that all viewpoints were accessed at least once in each training round. Through this training strategy, multi-view supervision information can be gradually passed to all Gaussian units in multiple training rounds while controlling memory usage and computational complexity.
[0100] Furthermore, during training, conventional density control strategies used in 3D Gaussian scene reconstruction can be incorporated to split, duplicate, or eliminate Gaussian primitives. For example, Gaussian primitives with large gradient responses and wide coverage areas can be appropriately split to improve the accuracy of local geometry and saliency representation; Gaussian primitives with small contributions or low long-term transparency can be eliminated to control model size and improve training efficiency. Since the newly added saliency encoding parameters are bound to the Gaussian primitives along with the original geometric appearance parameters, the aforementioned dynamic adjustment process of Gaussian primitives can also be naturally extended to the updating and inheritance of saliency attributes.
[0101] As training iteratively progresses, the geometric structure of Gaussian elements gradually conforms to the real shape of the scene, and the color appearance gradually approximates the input image. Meanwhile, the saliency encoding parameters gradually converge to a stable state under multi-view saliency supervision constraints and 3D neighborhood consistency constraints. Ultimately, each Gaussian element not only accurately represents the local geometry and appearance content of the scene but also obtains saliency attributes corresponding to the visual importance of that local region, thus forming a saliency-embedded 3D Gaussian scene representation.
[0102] After training, a saliency-embedded 3D Gaussian scene representation is obtained. Each Gaussian element contains attributes for position, scale, rotation, opacity, appearance, and saliency. This represents the final number of Gaussian units obtained after splitting, eliminating, and optimizing during the training process; it does not necessarily have to be exactly the same as the initial number of Gaussian units.
[0103] For the reconstructed scene, under any given viewpoint, not only can the corresponding color image be rendered, but also a saliency probability distribution map or a scalar saliency map can be rendered simultaneously. Thus, each Gaussian unit in the scene is no longer just a color-carrying unit in the traditional sense, but becomes a unified representation unit with geometric meaning, visual appearance meaning, and saliency importance meaning.
[0104] Furthermore, in downstream tasks such as compression or transmission, the saliency attribute of Gaussian primitives can be directly used to prioritize different Gaussian primitives. For example, in bandwidth-constrained transmission scenarios, Gaussian primitives with higher saliency values can be transmitted first, while the transmission of Gaussian primitives with lower saliency values can be delayed or reduced. In storage compression scenarios, higher compression ratios can be applied to low-saliency regions, while more detailed geometric and appearance information can be preserved for high-saliency regions. In viewpoint rendering scenarios, more sampling and rendering resources can be allocated to high-saliency regions, thereby improving the user's subjective perception quality within a limited computational budget. Since this saliency attribute is learned directly in the 3D representation layer, there is no need to separately perform complex 2D saliency estimation, cross-view fusion, or 3D importance backpropagation processes in subsequent application stages, which can significantly simplify the system process, reduce additional computational overhead, and improve overall processing efficiency.
[0105] This embodiment demonstrates that the present invention can directly embed saliency attributes during the 3D Gaussian scene reconstruction stage, enabling saliency attributes to be jointly learned and optimized with geometric and appearance attributes under a unified framework. This avoids the complex process of additional saliency estimation and importance analysis after scene reconstruction in traditional methods, improving the completeness of 3D scene representation, the sufficiency of supervision information utilization, and the convenience of downstream applications, and has good prospects for engineering applications.
[0106] To verify the effectiveness of the proposed saliency-embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision, experimental verification was conducted using data with multi-view videos and saliency annotations. In the experiment, the scene was first trained according to the method in Example 1 to obtain a 3D Gaussian scene representation carrying saliency attributes. Then, the obtained saliency attributes were further used for transmission scheduling under bandwidth-constrained conditions to verify the practical application effect of these saliency attributes in downstream tasks.
[0107] During transmission verification, Gaussian elements in the scene are divided into spatial blocks according to their three-dimensional spatial locations. The transmission priority of each spatial block is determined based on the number of significant Gaussian elements in each block, with priority given to transmitting spatial blocks containing more significant Gaussian elements. The experiment uses in-viewport image quality as the evaluation object, and Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Learned Perceptual Image Patch Similarity Index (LPIPS) as evaluation metrics.
[0108] Table 1
[0109] Furthermore, the method of the present invention was further verified under a fluctuating 2000Mbps home network environment with 10fps conditions. The experimental results are shown in Table 1, and the visualization results are as follows. Figure 4 As shown in the figure. The results show that, using the 3DGS with embedded saliency proposed in this invention for perception-optimized transmission, the image quality within the viewport can achieve an SSIM of 0.9750, a PSNR of 31.8178 dB, and an LPIPS of 0.0301; while the default transmission strategy corresponds to an SSIM of 0.9187, a PSNR of 19.0027 dB, and an LPIPS of 0.0950. Compared with the default transmission strategy, the method of this invention improves SSIM by at least 0.0563, PSNR by at least 12.8151 dB, and LPIPS by at least 0.0649, indicating that the saliency attribute constructed in this invention can effectively improve the perception reconstruction quality of the viewport region under bandwidth-constrained conditions.
[0110] The above results show that by introducing saliency coding parameters, saliency rendering branches, and joint optimization objectives in the 3D Gaussian scene reconstruction stage, the present invention enables the obtained 3D Gaussian scene representation to directly carry stable saliency attributes. These saliency attributes can be further used for priority scheduling under bandwidth-constrained conditions and significantly improve the perceptual reconstruction quality of the viewport region.
[0111] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0112] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0113] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the embedded 3D Gaussian scene reconstruction methods based on multi-view saliency supervision in the above embodiments.
[0114] It is understood that the systems, devices, and storage media provided in the embodiments of the present invention correspond to the methods provided in the embodiments of the present invention, and the explanations, examples, and beneficial effects of the relevant content can be referred to the corresponding parts of the above methods.
[0115] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0116] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0117] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0118] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision, characterized in that, Includes the following steps: S100. Acquire reconstruction scene data and establish a unified data foundation; S200. Based on a unified data foundation, construct Gaussian primitive representations with salient attributes; S300, based on Gaussian primitive representation, performs appearance rendering and saliency rendering; S400, based on appearance rendering and saliency rendering, constructs a joint optimization objective of saliency attributes, geometric attributes, and appearance attributes, and iteratively reconstructs the scene.
2. The embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision according to claim 1, characterized in that, The method for establishing a unified data foundation in step S100 includes: Obtain multi-view image sequences corresponding to the scene to be reconstructed Camera parameters and the saliency supervision distribution map corresponding to each viewpoint image. ; in, Represents the number of viewpoints. Indicates the first Color images from various perspectives This represents the saliency supervision distribution plot corresponding to this perspective. The camera parameters representing this viewpoint; The camera parameters include an intrinsic parameter matrix. and extrinsic parameter matrix ,in Represents the rotation matrix. Represents the translation vector; The first The projection model of each viewpoint is represented as: in, For the first The three-dimensional center coordinates of each Gaussian element. For its in the The projection center on the image plane from each viewpoint.
3. The embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision according to claim 1, characterized in that, Methods for constructing Gaussian meta-representations with salient attributes include: Represent the scene as A set of Gaussian elements , among which, the Gao Siyuan Defined as: in, For position parameters, For scale parameters, For rotation parameters, For opacity parameters, For appearance parameters, For saliency coding parameters; Saliency coding parameters ,in Preset dimensions; No. The covariance matrix of the Gaussian elements in three-dimensional space is represented as: in, Indicated by rotation parameters The corresponding rotation matrix, These represent the scale parameters of the Gaussian element in the x-axis, y-axis, and z-axis directions, respectively. In the From each viewpoint, the two-dimensional covariance matrix of the Gaussian element projected onto the image plane is expressed as: in, This represents the Jacobian matrix of the 3D-to-2D projection near the center of Gauss.
4. The embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision according to claim 1, characterized in that, The appearance rendering and saliency rendering methods in step S300 include: For any pixel on the image plane , No. The first Gaussian Yuan in the The two-dimensional response of this pixel from each viewpoint is represented as follows: In the formula, Let Gausky's elements be the two-dimensional covariance matrix projected onto the image plane. For the first Projection model from multiple perspectives; Combined with opacity parameter Its effect on pixels The pixel-level contribution is represented as: If the Gaussian elements participating in the rendering of the same pixel are sorted according to depth from near to far, then the... The cumulative transmittance of each Gaussian pixel at this pixel location is expressed as: Therefore, the first The color image rendering results from each viewpoint are represented as follows: in, Indicates the first The first Gaussian Yuan in the Color output from various perspectives, determined by its appearance parameters Calculated by combining the viewpoint direction; A saliency rendering process is constructed to achieve the micro-projection of saliency attributes from 3D Gaussian primitives to the 2D image plane. For saliency rendering, saliency is first encoded. The mapping is to a saliency output, which is represented in the following form: ; in, The saliency mapping function can be a fully connected linear classifier. ; The rendering result of the significance distribution plot is then expressed as follows: .
5. The embedded 3D Gaussian scene reconstruction method based on multi-view saliency supervision according to claim 1, characterized in that, The method for constructing the joint optimization objective of saliency attributes, geometric attributes, and appearance attributes in step S400 includes: Image reconstruction loss is expressed as: in, The image reconstruction error function is represented in the following form: in, and For balance coefficient, express Norm, Indicators representing structural similarity; Significance-based supervised loss is expressed as: in, and These represent the image height and width, respectively. Indicates the significance level number. Represents cross-entropy loss; set up Indicates the first A high-ranking official Nearest neighbor set Indicates the first The significance distribution corresponding to each Gaussian element is then expressed as: in, This represents the number of sampled Gaussian elements involved in the calculation of three-dimensional consistency constraints. Indicates KL divergence; The overall optimization objective is expressed as: in, and These are the weight coefficients for the significance supervision loss term and the three-dimensional neighborhood consistency constraint term, respectively.