Image rendering method and system based on cosine modulation gaussian kernel representation
By using the cosine-modulated Gaussian kernel representation method, the split logical Gaussian primitive is divided into multiple standard Gaussian sub-primitives, which solves the problem of high-frequency detail blurring in 3D Gaussian sputtering, improves the rendering effect and maintains compatibility.
Patent Information
- Application Number
- CN202511157253.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Existing 3D Gaussian sputtering technology is prone to blurring or loss of detail when representing high-frequency details in the human body and clothing.
The cosine-modulated Gaussian kernel representation method is adopted. By splitting the logical Gaussian primitive into multiple standard Gaussian sub-primitives according to the wave vector, and using an adaptive splitting and merging mechanism to control the primitive density, the high-frequency detail representation capability is improved.
It significantly enhances the ability to express high-frequency details while maintaining good rendering effects in the low-frequency range, and is compatible with the existing 3DGS framework, enabling adaptive high-frequency mode control.
Smart Images

Figure CN120747325B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of three-dimensional reconstruction, in particular to an image rendering method and system based on cosine modulation Gaussian kernel representation. BACKGROUND
[0002] Three-dimensional Gaussian Splatting (3DGS) represents a scene by using a large number of three-dimensional Gaussian functions, which achieves excellent reconstruction quality while maintaining real-time rendering speed. The standard three-dimensional Gaussian function G(x, μ, Σ) is defined as:
[0003] ;
[0004] where x is a point in three-dimensional space, μ is the Gaussian center position, and Σ is the covariance matrix. Its color is usually represented by Spherical Harmonics (SH), and the opacity is In the frequency domain, the Fourier transform of the Gaussian function is still a Gaussian function:
[0005] ;
[0006] where ω is the frequency vector. The form of this Fourier transform shows that the Gaussian function is essentially a low-pass filter, whose spectral energy is mainly concentrated near zero frequency. This shows that the standard 3DGS often has problems of blurring or loss of details when representing fine textures, sharp edges, and other high-frequency details in the human body and clothes. In order to enable the Gaussian primitive to express high-frequency information, it is necessary to modulate its spectrum and move its energy from the low-frequency region to the high-frequency region. SUMMARY
[0007] In order to solve the problems existing in the prior art, the purpose of the present application is to provide an image rendering method and system based on cosine modulation Gaussian kernel representation, which can improve the high-frequency detail expression ability of three-dimensional Gaussian Splatting.
[0008] To achieve the above purpose, the present application provides the following scheme:
[0009] An image rendering method based on cosine modulation Gaussian kernel representation, comprising:
[0010] Collecting a target image, inputting the target image into a Splatting model to obtain a rendered image under a target camera view; the Splatting model is obtained by training a training set; the training set includes original images and corresponding camera parameters thereof;
[0011] The acquiring the rendering image under the target camera view comprises: establishing a corresponding cosine modulation Gaussian kernel, i.e. a logical Gaussian cell, for each sparse point in the target image, splitting each logical Gaussian cell into a plurality of standard Gaussian sub-cells according to a wave vector, and rendering the target image under the target camera view.
[0012] Optionally, the splitting each logical Gaussian cell into a plurality of standard Gaussian sub-cells according to a wave vector comprises:
[0013] According to the behavior of each cosine modulation Gaussian kernel in a cosine maximum value interval, when the inner product of a wave vector and a position deviation When in a first threshold interval, the cosine modulation Gaussian kernel is expanded by using a cosine function, and the expanded cosine modulation Gaussian kernel is substituted into the cosine modulation Gaussian;
[0014] When the square distance of the phase deviation from the kth cosine peak value When in a second threshold interval, the substituted cosine modulation Gaussian is adjusted, and an index term is reorganized based on a matching method to obtain a logical Gaussian cell in a standard Gaussian form;
[0015] The splitting each logical Gaussian cell in a standard Gaussian form into a plurality of standard Gaussian sub-cells according to a wave vector comprises:
[0016] ;
[0017] wherein, is an upper bound of the number of splits, is a weight of a standard Gaussian sub-cell, is an i-th Gaussian cell, is a cosine modulation function of an i-th standard Gaussian, is a wave vector of an i-th standard Gaussian, is a position of a point in a three-dimensional space, is a mean value of an i-th standard Gaussian, is a sub-Gaussian index, is a kth sub-Gaussian center of an i-th cell, is an inverse of a covariance multiplied by a point position minus a Gaussian mean value.
[0018] Optionally, the obtaining the logical Gaussian cell in a standard Gaussian form comprises:
[0019] An offset vector is set, and a quadratic term and a second term of an index term in the adjusted cosine modulation Gaussian kernel are expanded by using the offset vector, so as to determine a new center and a covariance matrix:
[0020] ;
[0021] wherein, is the position vector x and the k-th sub-Gaussian center transpose of the difference, is the position vector x and the original Gaussian center transpose of the difference, is the product of the inverse of the covariance matrix and the position bias, is the inner product of the wave vector transpose and the position bias, which represents the phase of the cosine modulation, is the index of the sub-Gaussian, is the constant term;
[0022] Based on the new center and the covariance matrix, a logical Gaussian primitive in the standard Gaussian form is obtained.
[0023] Optionally, determining the weight of the standard Gaussian sub-primitive includes:
[0024] ;
[0025] wherein, is the weight of the center sub-Gaussian, is the index of the sub-Gaussian, is the transpose of the wave vector, is the wave vector.
[0026] Optionally, determining the number of the standard Gaussian sub-primitives includes:
[0027] ;
[0028] wherein, is the number of the standard Gaussian sub-primitives.
[0029] Optionally, the splitting criterion of the logical Gaussian primitive is:
[0030] Based on the adaptive splitting and merging mechanism, the density of the logical Gaussian primitive is controlled, the density of the logical Gaussian primitive is increased in the high-frequency region, the density of the logical Gaussian primitive is reduced in the low-frequency region, and when the logical Gaussian primitive cannot represent local high-frequency details, the logical Gaussian primitive is split into a plurality of standard Gaussian sub-primitives.
[0031] Optionally, rendering the target image under the target camera perspective includes:
[0032] projecting the standard Gaussian sub-primitive to the image plane, calculating the 2D Gaussian parameters of the standard Gaussian sub-primitive;
[0033] sorting the standard Gaussian sub-primitives containing 2D Gaussian parameters according to depth, and mixing the accumulated color values through opacity rendering the target image under the target camera perspective.
[0034] Optionally, training the sputtering model by using the training set comprises:
[0035] initializing the logical Gaussian basis by using the sparse point cloud; wherein the center of the logical Gaussian basis is initialized as the three-dimensional coordinates of the center point, the basic covariance is initialized as the distance between the center point and its nearest neighbor point, the basic opacity is initialized as the first target value, the basic spherical harmonic color coefficient is initialized as the color of the center point, and the learnable parameter is initialized as the second target value;
[0036] iteratively training the initialized logical Gaussian basis by using the training set, and in each iteration, generating an original rendered image for the camera view of the original image by using all the logical Gaussian bases, and in the process of generating the original rendered image, decomposing each logical Gaussian basis into a plurality of standard Gaussian sub-bases by using the high-frequency modulation vector decoded from the learnable parameter of each logical Gaussian basis and the modulation vector;
[0037] calculating the L1 loss and the structural similarity D-SSIM loss between the original rendered image and the original image, performing weighted summation on the L1 loss and the structural similarity D-SSIM loss, and updating the parameters of all the logical Gaussian bases until the model converges.
[0038] To achieve the above object, the application further provides an image rendering system based on cosine modulation Gaussian kernel representation, comprising:
[0039] an image acquisition module for acquiring a target image;
[0040] an image rendering module for inputting the target image into a sputtering model to obtain a rendered image under a target camera view; the sputtering model is obtained by training a training set; the training set comprises an original image and corresponding camera parameters thereof;
[0041] obtaining the rendered image under the target camera view comprises: establishing a corresponding cosine modulation Gaussian kernel, i.e., a logical Gaussian basis, for each sparse point in the target image, splitting each logical Gaussian basis into a plurality of standard Gaussian sub-bases according to a wave vector, and rendering the target image under the target camera view.
[0042] The application has the following beneficial effects:
[0043] The application introduces the modulation and splitting mechanism of the Gaussian kernel, so that the scene model reconstructed from the real image can effectively encode and reconstruct the high-frequency information in the scene, and can significantly enhance the high-frequency detail expression capability.
[0044] The sputtering model based on the cosine modulation Gaussian can keep the good properties of the standard Gaussian function fitting data, and still has good rendering effect in the low-frequency part.
[0045] The learnable wave vector parameter in the application can adaptively adjust the high-frequency mode of the Gaussian kernel, thereby realizing adaptive high-frequency mode control.
[0046] The cosine modulation Gaussian kernel proposed in the application has good portability, can be used immediately, is compatible with the existing 3DGS framework, and can be directly embedded into the 3DGS rendering pipeline for use. BRIEF DESCRIPTION OF DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description only constitute some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0048] Figure 1 A framework diagram of an image rendering method based on a cosine modulation Gaussian kernel representation according to an embodiment of the application;
[0049] Figure 2 A visualization diagram of cosine modulation Gaussian functions under different modulation frequencies according to an embodiment of the application;
[0050] Figure 3 An effect diagram of a cosine modulation Gaussian sputtering model on scene reconstruction according to an embodiment of the application; (a) is a real image of a scene, (b) is a rendering image effect diagram of the real image of the scene under the cosine modulation Gaussian sputtering model, and (c) is an effect diagram of the real image of the scene under the cosine modulation Gaussian sputtering model on scene reconstruction. DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments only constitute some of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.
[0052] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the application will be further described in detail below with reference to the drawings and specific embodiments.
[0053] As Figure 1As shown, the embodiment discloses an image rendering method based on cosine modulation Gaussian kernel representation, comprising: collecting a target image, inputting the target image into a sputtering model to obtain a rendering image under a target camera view angle; the sputtering model is obtained by training a training set; the training set includes: original images and corresponding camera parameters; obtaining a rendering image under a target camera view angle includes: establishing a corresponding cosine modulation Gaussian kernel, i.e. a logical Gaussian primitive, for each sparse point in the target image, splitting each logical Gaussian primitive into multiple standard Gaussian sub-primitives according to wave vectors, and rendering the target image under the target camera view angle.
[0054] Further, splitting each logical Gaussian primitive into multiple standard Gaussian sub-primitives according to wave vectors includes: according to the behavior of each cosine modulation Gaussian kernel in the cosine maximum value interval, when in the first threshold interval, the cosine modulation Gaussian kernel is expanded by using the cosine function, and the expanded cosine modulation Gaussian kernel is substituted into the cosine modulation Gaussian; when in the second threshold interval, the substituted cosine modulation Gaussian is adjusted, and the index term is reorganized based on the method to obtain a logical Gaussian primitive in the standard Gaussian form; each logical Gaussian primitive in the standard Gaussian form is split into multiple standard Gaussian sub-primitives according to wave vectors.
[0055] obtaining the logical Gaussian primitive in the standard Gaussian form includes: setting an offset vector, expanding the quadratic term and the second term of the index term in the adjusted cosine modulation Gaussian kernel by using the offset vector, thereby determining a new center and a covariance matrix: based on the new center and the covariance matrix, the logical Gaussian primitive in the standard Gaussian form is obtained.
[0056] rendering the target image under the target camera view angle includes: projecting the standard Gaussian sub-primitive to the image plane, calculating the 2D Gaussian parameters of the standard Gaussian sub-primitive; sorting the standard Gaussian sub-primitives containing 2D Gaussian parameters according to depth, and rendering the target image under the target camera view angle by mixing and accumulating color values according to opacity
[0057] Specifically, the cosine modulation Gaussian kernel: according to the theory of cosine modulation Gaussian decomposition, the cosine modulation Gaussian function is defined as:
[0058] ;
[0059] wherein, is the wave vector, controls the frequency and direction of modulation, and encodes the frequency and direction of modulation, is a standard Gaussian function, is a cosine modulation function, is a point in three-dimensional space, is the center position of the Gaussian, is the covariance matrix, is the phase of the cosine modulation function, is the specific value of the cosine modulation Gaussian function. From the form of the cosine modulation Gaussian function, its isosurface has a unique periodic structure. Let the isosurface of the cosine modulation Gaussian function be where 0 < c < 1, then:
[0060] ;
[0061] where, is a natural number, is the cosine modulation function, is the inverse of the covariance multiplied by the point position minus the Gaussian mean.
[0062] This requires , is the cosine modulation function, i.e., the function only has positive values in the region where the cosine is positive. The cosine function divides the three-dimensional space into parallel layers with a layer spacing of In each layer, , is the cosine modulation function, the region of forms a "blob" that is approximately an ellipsoid. These blobs are periodically arranged along the wave vector direction, forming a "bead-like" isosurface structure.
[0063] Parameter calculation of Gaussian splitting: The cosine-modulated Gaussian function can be approximately represented as a weighted sum of multiple standard Gaussian functions, i.e.,
[0064] ;
[0065] where, is the upper bound of the number of sub-Gaussians, is the weight of the kth sub-Gaussian, and are the center and covariance matrix of the kth sub-Gaussian, respectively.
[0066] The splitting calculation of the cosine-modulated Gaussian function is proved as follows:
[0067] Step 1: Local approximation of the cosine function: Consider the behavior of the cosine-modulated Gaussian function near the kth cosine maximum. When , the cosine function approaches its maximum value 1. Within this neighborhood, use the periodicity and Taylor expansion of the cosine function to obtain:
[0068] ;
[0069] where, is the sub-Gaussian index, is a high-order infinitesimal. In the case of (wT Within a small region of (x-μ)-2kπ), higher-order terms can be ignored.
[0070] Step 2: Local form of the modulated Gaussian function: Substituting the Taylor expansion into the original cosine-modulated Gaussian function, we get:
[0071] ;
[0072] when When the value is sufficiently small, the approximation (1-ε)≈e is used. -ε ε≪1, we get:
[0073] ;
[0074] Where ε is a small quantity, e -ε It is an approximate exponential function.
[0075] This local form transforms the complex expression containing the cosine function into a pure exponential form, which can be represented as a standard Gaussian function. This transformation is the core of the entire splitting theory, allowing the cosine-modulated Gaussian function to be approximated by multiple standard Gaussian functions. In constructing the sputtering model, the significance of this splitting representation lies in avoiding the enormous computational overhead of directly calculating the cosine function in the rendering pipeline, while simultaneously enabling the cosine-modulated Gaussian kernel to be seamlessly integrated into the existing 3DGS rendering framework, since the 3DGS rendering pipeline itself is designed for handling standard Gaussian functions.
[0076] Step 3: Completing the Square and Rearranging the Exponent: To convert the above equation into standard Gaussian form, the exponent needs to be completed. Define the offset vector δ. k Make w T δ k =2kπ, and let μ k =μ+δ k Expanding the quadratic terms in the exponent term, we get:
[0077] ;
[0078] Expanding the second term, we get:
[0079] ;
[0080] Therefore, the overall quadratic form can be written as:
[0081] ;
[0082] Step 4: Determine the sub-Gaussian parameters: Find the new center μ using the completing the square method. k The covariance matrix Σ k , so that:
[0083] ;
[0084] where, is a constant independent of x. Through matrix algebraic operations, it can be proved that the sub-Gaussian mean is:
[0085] ;
[0086] This formula shows that the center of the sub-Gaussian is shifted along the direction, and the shift is proportional to k. The inverse of the sub-Gaussian covariance is:
[0087] ;
[0088] Using the Sherman-Morrison matrix inversion formula, we get:
[0089] ;
[0090] where, is the wave vector, is the transpose of the wave vector, is the covariance of the kth sub-Gaussian, is the covariance of the standard Gaussian, is the multiplication of two matrices. This shows that the sub-Gaussian is "flattened" in the direction of the wave vector , while it remains unchanged in the direction perpendicular to .
[0091] Step 5: Determine the weight coefficient: the constant term generated in the formulation process determines the weight of the sub-Gaussian , ; where, is the weight of the kth sub-Gaussian, is the weight of the central sub-Gaussian (usually set to 1). This weight reflects the spatial decay of the original Gaussian function, and the weight decreases exponentially with the increase of |k|, k being the index of the sub-Gaussian. This reflects that the cosine-modulated Gaussian function is a cosine function with a Gaussian envelope in the spatial domain. From the isosurface, the contribution of the sub-Gaussian far from the center gradually decreases, which is consistent with the spatial locality of the original Gaussian function, and the overall energy remains unchanged, only more concentrated in the central region.
[0092] Step 6: Determine the number of sub-Gaussians: The isosurface of the cosine-modulated Gaussian function presents a "bead-like" structure arranged periodically along the wave vector direction, theoretically containing an infinite number of small Gaussian spheres. In order to realize this representation in practical calculations, a finite number of sub-Gaussians are needed to fit this complex isosurface. However, directly using an infinite number of sub-Gaussians is not feasible in calculations, so a reasonable truncation criterion must be found.
[0093] The present application adopts the 3σ principle of Gaussian function to determine the truncation range. By multiplying with cosine function, the original Gaussian function is split into N s sub-Gaussians , . The schematic diagram of cosine-modulated Gaussian function split into sub-Gaussians is shown in Figure 2 . The number of sub-Gaussians is determined by and the spatial range of Gaussian, .
[0094] Although Gaussian function has non-zero value in the whole space in theory, its energy is mainly concentrated in the limited area near the center. According to the 3σ principle of Gaussian function in probability theory, in one-dimensional case, 68.3% of the probability mass is concentrated in the range of [μ-σ, μ+σ], 95.4% of the probability mass is concentrated in the range of [μ-2σ, μ+2σ], and 99.7% of the probability mass is concentrated in the range of [μ-3σ, μ+3σ]. For three-dimensional Gaussian function, the present application defines its effective spatial range as the area satisfying the Mahalanobis distance (x-μ) T Σ -1 (x-μ)≤9, which corresponds to the 3σ range in each principal axis direction. Outside this range, the value of Gaussian function is less than 0.01% of its peak value, which can be ignored. Based on this, the number of sub-Gaussians is determined as:
[0095] ;
[0096] wherein σ is the standard deviation and μ is the center position of Gaussian. Within the effective support range of Gaussian, i.e. 3σ distance along the wave vector direction, enough sub-Gaussians need to be placed to capture each period of cosine function. Therefore, it is necessary to ensure that within the effective spatial range of Gaussian, the oscillation of cosine function can be fully sampled, so that within the 3σ range, each period of cosine function has enough Gaussian elements to represent. Because cosine function itself has infinite periods, it is impossible to create infinite Gaussian balls, so the 3σ principle is adopted to truncate the iso-surface, but only the part containing 0.01% of the energy of the original function is discarded, so that the error caused by approximation can be finally reduced.
[0097] Sputtering model based on cosine-modulated Gaussian kernel: The sputtering model based on cosine-modulated Gaussian kernel is essentially a scene representation system composed of a large number of cosine-modulated Gaussian kernels. The traditional 3D Gaussian sputtering uses standard Gaussian function as the basic unit to represent the scene, while the present application uses cosine-modulated Gaussian kernel as the basic unit. The whole three-dimensional scene is represented as:
[0098] ;
[0099] wherein opacity weight of the i-th logical Gaussian cell, i-th standard Gaussian cell function, cosine modulation function of the i-th cell, inner product of the wave vector and position offset of the i-th cell (modulation phase), color function of the i-th cell (spherical harmonic representation).
[0100] The present application proposes a cosine-modulated Gaussian kernel based sphererendering model, which takes multi-view images of a scene as input and outputs a 3D model that can be used to render high-quality images of arbitrary new views, as shown in Figure 1 .
[0101] Model construction process: First, for each sparse point a cosine-modulated Gaussian kernel is created to form the initial representation of the scene. These cosine-modulated Gaussian kernels are implemented as logical Gaussian cells in the system, and each logical cell stores its complete parameter set 、 、 、 、 }. mean of the i-th cell, base covariance matrix of the i-th cell, base opacity of the i-th cell, base color function of the i-th cell, wave vector of the i-th cell.
[0102] Secondly, a dynamic splitting rendering mechanism is established. Since the direct rendering of cosine-modulated functions has huge computational overhead, the present application adopts a dynamic splitting strategy: during rendering, each cosine-modulated Gaussian kernel is split into multiple standard Gaussian sub-cells in real time according to its wave vector . This splitting approximates the complex cosine-modulated function as the sum of multiple simple Gaussian functions:
[0103] ;
[0104] In this way, the entire scene is represented as a collection of all sub-Gaussians during rendering, and a standard Gaussian sphererendering pipeline can be directly used.
[0105] A differentiable rendering pipeline is constructed. For a given camera view, the rendering process includes: splitting all logical Gaussian cells into sub-Gaussian sets; projecting the sub-Gaussians onto the image plane and calculating their 2D Gaussian parameters; sorting according to depth; accumulating color values by mixing to generate the final image. The entire process maintains differentiability, supporting gradient-based parameter optimization.
[0106] Finally, the model parameters are optimized through end-to-end training. The model uses the rendering loss to drive the learning of the parameters of all the cosine-modulated Gaussian kernels, especially the wave vectors The optimization of the wave vectors enables each kernel to adaptively adjust its frequency characteristics, expressing different scales of details in different regions of the scene.
[0107] With this construction, the sputtering model based on cosine-modulated Gaussian kernels not only maintains the real-time rendering capability of 3DGS, but also significantly improves the expression capability of high-frequency details. The core innovation of the model is to seamlessly integrate the cosine modulation mechanism into the Gaussian sputtering framework, avoiding the overhead of directly calculating the cosine function through a dynamic splitting strategy, achieving a balance between theoretical innovation and engineering efficiency.
[0108] Further, the splitting criterion of the logical Gaussian primitive is that the density of the logical Gaussian primitive is controlled based on an adaptive splitting and merging mechanism, the density of the logical Gaussian primitive is increased in a high-frequency region, the density of the logical Gaussian primitive is reduced in a low-frequency region, and when the logical Gaussian primitive cannot represent local high-frequency details, the logical Gaussian primitive is split into multiple standard Gaussian sub-primitives.
[0109] Specifically, the model controls the density of the logical Gaussian primitive through an adaptive splitting and merging mechanism. The number of cosine-modulated Gaussian primitives is dynamically adjusted according to the complexity of the scene. The primitive density is increased in a high-frequency region, and the number of primitives is reduced in a low-frequency region. When a cosine-modulated Gaussian primitive cannot accurately represent local high-frequency details, it is split into multiple sub-primitives. The primitives are split in the spatial domain, and when initializing the parameters of the sub-primitives, the center position of the parent primitive is maintained, and the covariance and modulation frequency are adjusted. When the frequency spectra of multiple cosine-modulated Gaussian primitives overlap, they are merged into one primitive. Based on the similarity of the primitives, the center position, the covariance, and the frequency are used for merging.
[0110] Further, training the sputtering model using the training set includes: initializing the logistic Gaussian primitives using sparse point clouds; wherein, the center of the logistic Gaussian primitive is initialized to the three-dimensional coordinates of the center point, the basic covariance is initialized to the distance between the center point and its nearest neighbor, the basic opacity is initialized to a first target value, the basic spherical harmonic color coefficient is initialized to the color of the center point, and the learnable parameters are initialized to a second target value; iteratively training the initialized logistic Gaussian primitives using the training set, and in each iteration, generating an original rendered image from the camera perspective of the original image using all current logistic Gaussian primitives; during the generation of the original rendered image, decomposing each logistic Gaussian primitive into multiple standard Gaussian sub-primaries based on the high-frequency modulation vector decoded from the learnable parameters of each logistic Gaussian primitive; calculating the L1 loss and the structural similarity D-SSIM loss between the original rendered image and the original image, weighting and summing the L1 loss and the structural similarity D-SSIM loss, and updating the parameters of all logistic Gaussian primitives until the model converges.
[0111] Specifically, the sputtering model based on cosine-modulated Gaussian kernels aims to optimize the parameters of all logistic Gaussian units by minimizing the loss between the rendered image output by the model and the input real image. The specific rendering optimization process involves: 1) Initializing the cosine-modulated Gaussian kernel and using sparse point clouds... To initialize the logical Gaussian set. For each point in the sparse point cloud. Initialize a logical high-order element. Element The center is initialized as a point Three-dimensional coordinates; fundamental covariance Initialize as a point The distance to its nearest neighbor represents an isotropic Gaussian sphere; the fundamental opacity. Initialize to a small value; basic spherical harmonic color coefficient Initialize as a point The color; the wave vector with learnable parameters. Initializing to zero or a small random value indicates that the high-frequency modulation effect is weak at the beginning of training. 2) In each iteration, a random image is selected from the input images as the target image for the current iteration. And obtain the camera parameters corresponding to the target image. , Given the rotation matrix, translation vector, and camera intrinsics for the k-th image. 3) Using all current logical Gaussian units, generate a rendered image for the camera viewpoint of the target image. During this process, each logical Gaussian unit... Wave vector with learnable parameters decoded high frequency modulation vector . By modulating vector , the basis element can be decomposed into standard Gaussian sub-basis elements. This sub-basis elements have their mean shifted with respect to the mean of the original logical Gaussian basis element . 4) Calculate the L1 loss and structural similarity D-SSIM loss between the rendered image and the target image, the loss function is the weighted sum of the L1 loss and the SSIM loss, . 5) Update the parameters of all logical Gaussian basis elements, especially the wave vector of the learnable parameters , so that they can generate a rendered image closer to the target image in the next iteration. 6) Repeat the above process until the model converges. After the completion of the cosine modulation Gaussian splatting model training, a new perspective image with rich high frequency details can be generated for any given new camera perspective rendered image.
[0112] The cosine modulation Gaussian-based splatting model can be applied to three-dimensional scene reconstruction. The model outputs a three-dimensional model that can be used to render high-quality images of any new perspective by receiving a set of multi-view images of the scene as input. Figure 3 Figures (a)-(c) in the above show the effect of the cosine modulation Gaussian splatting model on scene reconstruction, and the scene rendered image can well reflect the real scene image, verifying the effectiveness of the cosine modulation Gaussian kernel splatting model. The present application applies the cosine modulation Gaussian splatting model to three-dimensional scene reconstruction, obtaining a three-dimensional scene model with fine texture, and rendering a scene rendered image with high frequency details from a new perspective.
[0113] The embodiment provides an image rendering method based on a cosine modulation Gaussian kernel representation, comprising:
[0114] The embodiment describes how to apply the present technology to reconstruct a three-dimensional model of an object from a set of real multi-view images of the object, and to render a rendered image with high frequency details from a new perspective.
[0115] Object multi-view data acquisition and preprocessing: select a target object with rich surface texture, such as carved vases, fabrics, fine mechanical parts, etc.; use a high-resolution camera, at least 12 million pixels is recommended; set appropriate lighting conditions to avoid overexposure or shadows; take 50-100 images around the object, ensuring that the adjacent view overlap reaches , and at least 3 different height levels in the vertical direction and 10-15 degree intervals in the horizontal direction.
[0116] Next, the data is preprocessed, the main steps include: 1) image screening, that is, to eliminate blurred, overexposed or underexposed images. 2) SfM processing using COLMAP to obtain object reconstruction sparse point cloud and camera parameter estimation. The output data of SfM includes the camera intrinsic matrix of each image , the camera extrinsic rotation matrix of each image , and the translation vector , the sparse three-dimensional point cloud ; wherein, is the three-dimensional coordinates of the jth point in the sparse point cloud, is the RGB color value of the jth point in the sparse point cloud, is the total number of points in the sparse point cloud.
[0117] Initialization of the cosine-modulated Gaussian kernel: The first step of model training is to initialize the cosine-modulated Gaussian kernel. For each point in the sparse point cloud, a logical Gaussian cell is initialized. This logical Gaussian cell is the representation of the cosine-modulated Gaussian kernel in the computer system. The parameters of the logical Gaussian cell are the position center , the basis covariance , wherein , the basis spherical harmonic color coefficient , the point is initialized to 0 order coefficient basis opacity , the wave vector parameter vector or a small random value N(0.0.0113); wherein, is the average distance of the jth point from its nearest neighbors, is the number of nearest neighbors, is the nth nearest neighbor, is the three-dimensional coordinates of the nth nearest neighbor, is the center position of the ith logical Gaussian cell.
[0118] The initial value of the wave vector is set close to zero, which means that at the beginning of training, the behavior of the logical Gaussian cell is close to a standard Gaussian function, and the high-frequency modulation effect is weak. This initialization strategy enables the model to start optimization from a stable starting point, and as the training progresses, the wave vector will adaptively adjust according to the local features of the scene, increasing in areas that require high-frequency details and remaining small in smooth areas.
[0119] Splatting model training and optimization: Set the hyperparameters for model training and optimization. The number of iterations is set to 30,000, the Adam optimizer is used, the learning rate , the batch size batch_size=1, and the loss function weight . The process of each epoch includes: 1) image sampling: randomly select an image Lgt and camera parameters (K gt , R gt , t gt ), (K gt , R gt , t gt ) are the camera intrinsic, rotation matrix and translation vector of the k-th image; 2) Gaussian splitting and rendering: for each logical Gaussian cell i, compute the number of sub-Gaussians . Each logical Gaussian cell generates sub-Gaussians. For the k-th sub-Gaussian, its parameters are computed by the following formula. The center position offset of the sub-Gaussian is determined by the wave vector and the covariance, ; where, is the base covariance of the i-th cell, is the inverse of the base covariance of the i-th cell. The covariance matrix of the sub-Gaussian is compressed in the direction of the wave vector, and the Sherman-Morrison matrix inversion formula is used to obtain . The opacity of the sub-Gaussian is weighted according to its distance from the center, ; where, is the base opacity of the i-th cell. The color of the sub-Gaussian is calculated by considering the change in viewing angle caused by the position offset. For the simple case, the color coefficient of the parent Gaussian can be directly inherited.
[0120] Rendering pipeline: project all sub-Gaussians to the image plane, perform depth sorting, and generate the final image L rendered by alpha blending.
[0121] Compute the loss , use the backpropagation algorithm to update the model parameters, and use the Adam optimizer to update all parameters, especially the wave vector .
[0122] System output and application: after training, the system outputs the optimized cosine-modulated Gaussian kernel parameter set. , is the mean of the i-th cell, is the base covariance matrix of the i-th cell, is the base opacity of the i-th cell, is the base color function of the i-th cell, The wave vector of the i-th basis element is denoted as w, and N is the number of Gaussian points in the scene, which can be used for real-time rendering of a three-dimensional model. The new view rendering process includes inputting new view camera parameters, performing a splitting operation on each cosine-modulated Gaussian kernel, and rendering to generate a high-quality image. Based on the sputtering model of the cosine-modulated Gaussian, the present application can be applied to digital content creation in the fields of virtual reality and augmented reality, such as e-commerce product display, cultural heritage digital protection, and film special effect production, etc.
[0123] Cosine-modulated Gaussian decomposition method: by multiplying a Gaussian function with a cosine function and decomposing it into multiple sub-Gaussians with different means using Euler's formula to achieve spectral shifting.
[0124] Dynamic number of Gaussian basis elements splitting mechanism: dynamically determine the number of splits according to the wave vector parameter w and the covariance matrix to ensure that the split sub-Gaussians can fully represent high-frequency details within the range of the Gaussian function .
[0125] Complete derivation of sub-Gaussian parameters: mean shift , covariance adjustment , opacity modulation .
[0126] End-to-end optimization framework of three-dimensional sputtering model based on cosine-modulated Gaussian kernel: integrate the cosine modulation mechanism into the differentiable rendering pipeline of 3DGS, and optimize the basic parameters and wave vector parameters of the cosine-modulated Gaussian function through gradient descent.
[0127] Three-dimensional reconstruction application system: including multi-view data acquisition specification, preprocessing process, training strategy of three-dimensional sputtering model based on cosine-modulated Gaussian kernel, and complete technical solution of rendering output.
[0128] The embodiment also provides an image rendering system based on cosine-modulated Gaussian kernel representation, including: an image acquisition module configured to acquire a target image; an image rendering module configured to input the target image into a sputtering model to obtain a rendered image under a target camera view; the sputtering model is obtained by training a training set; the training set includes: original images and corresponding camera parameters thereof; obtaining the rendered image under the target camera view includes: establishing a corresponding cosine-modulated Gaussian kernel, i.e., a logical Gaussian basis element, for each sparse point in the target image, splitting each logical Gaussian basis element into multiple standard Gaussian sub-basis elements according to a wave vector, and rendering the target image under the target camera view.
[0129] The above described embodiments are only to illustrate the preferred modes of the present application, and are not intended to limit the scope of the present application. Any modification and improvement made by those skilled in the art to the technical solutions of the present application without departing from the design spirit of the present application shall fall within the protection scope of the present application as defined by the claims.
Claims
1. An image rendering method based on cosine-modulated Gaussian kernel representation, characterized in that, include: Acquire a target image, input the target image into the sputtering model, and obtain a rendered image from the target camera's perspective; The sputtering model was obtained by training a training set. The training set includes: original images and their corresponding camera parameters; The process of obtaining a rendered image from the perspective of the target camera includes: establishing a corresponding cosine-modulated Gaussian kernel, i.e., a logical Gaussian primitive, for each sparse point in the target image; splitting each logical Gaussian primitive into multiple standard Gaussian sub-primaries according to the wave vector; and rendering the target image from the perspective of the target camera. Splitting each of the logical Gaussian primitives into multiple standard Gaussian sub-primaries based on the wave vector includes: Based on the behavior of each cosine-modulated Gaussian kernel in the cosine maxima region, when the inner product of the wave vector and the position deviation... Within the first threshold interval, the cosine function is used to expand the cosine-modulated Gaussian kernel, and the expanded cosine-modulated Gaussian kernel is substituted into the cosine-modulated Gaussian kernel; wherein, This is the transpose of the wave vector. The original Gaussian center; When the phase deviates from the square of the k-th cosine peak In the second threshold interval, the cosine-modulated Gaussian is adjusted after substitution, and the exponent is recombined based on the completing method to obtain the logical Gaussian element in the standard Gaussian form. Each logical Gaussian primitive in the standard Gaussian form is split into multiple standard Gaussian sub-primitives based on the wave vector: ; in, This is an upper bound on the number of splits. The weights of the standard Gaussian sub-basic units, For the i-th Gaussian element, Let i be the cosine modulation function of the i-th standard Gaussian. The transpose of the wave vector of the i-th standard Gaussian. The position of a point in three-dimensional space. Let be the mean of the i-th standard Gaussian. For sub-Gaussian index, Let k be the sub-Gaussian center of the i-th primitive. The inverse of the covariance is multiplied by the point position minus the Gaussian mean; The splitting criterion for the logical Gaussian unit is: The density of logical Gaussian primitives is controlled by an adaptive splitting and merging mechanism. The density of logical Gaussian primitives is increased in the high-frequency region and decreased in the low-frequency region. When the logical Gaussian primitives cannot represent local high-frequency details, they are split into multiple standard Gaussian sub-primaries.
2. The image rendering method based on cosine-modulated Gaussian kernel representation according to claim 1, characterized in that, Obtaining logical Gaussian elements in standard Gaussian form includes: An offset vector is set, and the quadratic and second terms of the exponential term in the adjusted cosine-modulated Gaussian kernel are expanded using this offset vector to determine the new center and covariance matrices: ; in, Let x be the position vector and the k-th sub-Gaussian center. Transpose of the difference Let position vector x be the original Gaussian center. Transpose of the difference It is the product of the inverse of the covariance matrix and the positional deviation. The dot product of the transpose of the wave vector and the position deviation represents the phase of the cosine modulation. For the index of sub-Gauss, For constant terms; Based on the new center and covariance matrix, the logical Gaussian elements in the standard Gaussian form are obtained.
3. The image rendering method based on cosine-modulated Gaussian kernel representation according to claim 1, characterized in that, Determining the weights of the standard Gaussian sub-units includes: ; in, The weights of the central sub-Gaussian, For the index of sub-Gauss, The transpose of the wave vector of the i-th standard Gaussian. Let be the wave vector of the i-th standard Gaussian wave.
4. The image rendering method based on cosine-modulated Gaussian kernel representation according to claim 1, characterized in that, The method also includes: Determining the number of the standard Gaussian sub-units includes: ; in, The number of standard Gaussian sub-units. This is the transpose of the wave vector. It is the wave vector.
5. The image rendering method based on cosine-modulated Gaussian kernel representation according to claim 1, characterized in that, Rendering the target image from the perspective of the target camera includes: The standard Gaussian sub-primitive is projected onto the image plane, and the 2D Gaussian parameters of the standard Gaussian sub-primitive are calculated. Sort by standard Gaussian sub-primitives containing 2D Gaussian parameters in depth, and sort by opacity. The accumulated color values are mixed to render the target image from the perspective of the target camera.
6. The image rendering method based on cosine-modulated Gaussian kernel representation according to claim 1, characterized in that, Training the sputtering model using the training set includes: The logical Gaussian unit is initialized using sparse point cloud; wherein, the center of the logical Gaussian unit is initialized to the three-dimensional coordinates of the center point, the basic covariance is initialized to the distance between the center point and its nearest neighbor, the basic opacity is initialized to a first target value, the basic spherical harmonic color coefficient is initialized to the color of the center point, and the learnable parameter is initialized to a second target value. The initialized logical Gaussian primitives are trained iteratively using the training set. In each iteration, the original rendered image is generated from the camera viewpoint of the original image using all the current logical Gaussian primitives. During the generation of the original rendered image, the high-frequency modulation vector decoded based on the learnable parameters of each logical Gaussian primitive is used to decompose each logical Gaussian primitive into multiple standard Gaussian sub-primaries. Calculate the L1 loss and structural similarity D-SSIM loss between the original rendered image and the original image, sum the L1 loss and the structural similarity D-SSIM loss in a weighted manner, and update the parameters of all logistic Gaussian units until the model converges.
7. An image rendering system based on cosine-modulated Gaussian kernel representation implemented according to any one of claims 1-6, characterized in that, include: The image acquisition module is used to acquire target images; The image rendering module is used to input the target image into the sputtering model and obtain a rendered image from the target camera's perspective. The sputtering model is obtained by training a training set, which includes: the original image and its corresponding camera parameters. Acquiring a rendered image from the perspective of the target camera includes: establishing a corresponding cosine-modulated Gaussian kernel, i.e., a logical Gaussian primitive, for each sparse point in the target image; splitting each logical Gaussian primitive into multiple standard Gaussian sub-primaries according to the wave vector; and rendering the target image from the perspective of the target camera.
Citation Information
Patent Citations
Reflecting object inverse rendering method, system and equipment based on two-dimensional Gaussian sputtering and multi-mode diffusion prior and medium
CN120472068A
Human subject gaussian splatting using machine learning
US20250148678A1