4D Gaussian reconstruction method and system based on self-heterogeneity
By introducing a learnable Alpha-self-heterogeneity parameter and an efficient rasterization clipping strategy, the Gaussian kernel transparency is adaptively adjusted to solve the motion blur problem in dynamic 3D scene reconstruction, and efficient and high-quality dynamic multi-view scene reconstruction is achieved.
Patent Information
- Application Number
- CN202511134969.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Existing dynamic 3D scene reconstruction methods face challenges in handling dynamic changes in the temporal dimension and maintaining high-quality rendering. In particular, blurred shadows appear at the edges of objects in dynamic scenes, resulting in a decrease in visual quality. Existing methods fail to effectively solve the motion blur problem.
A learnable Alpha-self-heterogeneity parameter and an efficient rasterization clipping strategy are introduced. By designing a region recognition function and a piecewise activation function for each Gaussian kernel, the transparency of the Gaussian kernel is adaptively adjusted, the alpha transparency is decoupled from the Gaussian shape, and the rendering process of the Gaussian kernel is optimized.
Significantly reduce motion blur, improve spatiotemporal consistency and reconstruction accuracy, enhance rendering efficiency, and achieve high-quality dynamic multi-view scene reconstruction.
Smart Images

Figure CN120747375A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision processing, and in particular relates to a 4D Gaussian reconstruction method and system based on self-heterogeneity. Background Art
[0002] Existing methods for reconstructing dynamic 3D scenes, particularly those based on Gaussian splatting, face challenges in handling dynamic changes in the temporal dimension while maintaining high rendering quality. Traditional Gaussian kernel parameters are typically fixed or simply interpolated, making it difficult to adaptively represent the complex and non-uniform dynamic deformations and transparency changes in the scene. Furthermore, redundant rasterization computations can also impact rendering efficiency.
[0003] Although deformation-field-driven 3D Gaussian Splatting (3DGS) methods have made significant progress in 4D reconstruction tasks in recent years, practical applications still face a core problem: severe motion blur in dynamic content. In particular, when rendering from new perspectives, the edges of objects in the scene often appear blurred and smeared, disrupting spatiotemporal consistency. This blur not only degrades visual quality but also limits the practical application value of related methods for high-quality 4D scene modeling. This problem is widely present in multiple public datasets and real-world data, demonstrating that it is not an isolated phenomenon but a systemic flaw shared by current mainstream methods.
[0004] Most current work attributes motion blur to the quality of the deformation field itself, focusing instead on improving its expressive power, such as designing more complex motion models and introducing temporal attention mechanisms. While these approaches do improve inter-frame alignment to some extent, research has found that they generally overlook the inherent limitations of the 3DGS's representational capabilities. Specifically, existing methods overly rely on external motion models when dealing with temporally varying scenes, failing to delve into the expressive conflicts inherent in the 3DGS's joint appearance-motion modeling in dynamic scenes. This "treating the symptoms rather than the root cause" strategy fails to address the root cause of motion blur.
[0005] A thorough analysis reveals that the key to the problem lies in the alpha blending mechanism employed by 3DGS. In traditional 3DGS, each Gaussian kernel is rendered on the image plane as a "template-like" structure, with high density at the center and gradually transparent at the edges. This design was originally intended to achieve smooth color transitions and high-fidelity appearance reconstruction in static scenes through the overlapping of Gaussians. However, in dynamic scenes, each Gaussian kernel must simultaneously fulfill two conflicting tasks: its central region is responsible for fitting local, high-frequency motion or texture boundaries, while its edges are responsible for smoothing low-frequency transitions across time frames. This dual role creates severe gradient conflicts during the optimization process, causing the Gaussian kernel's predicted motion direction to deviate from the true trajectory, ultimately resulting in noticeable motion blur in new viewing angles. It can be argued that the spatial-frequency coupling in the alpha blending mechanism is a structural bottleneck in the current application of 3DGS to dynamic scenes. Summary of the Invention
[0006] The present invention aims to address the quality, efficiency, and adaptability issues of dynamic Gaussian scene reconstruction in existing technologies, and to provide a 4D Gaussian reconstruction method and system based on self-heterogeneity. This method achieves high-quality reconstruction of dynamic multi-view scenes by introducing a learnable alpha-self-heterogeneity parameter and combining it with an efficient rasterization and cropping strategy.
[0007] In order to achieve the above-mentioned object of the invention, the present invention specifically adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a 4D Gaussian reconstruction method based on self-heterogeneity, which comprises the following steps:
[0009] S1: Accurately parameterize the input original timestamp and original observation angle, and combine the processed timestamp and the processed observation angle with the original image of the corresponding timestamp and observation angle to form training data;
[0010] S2: Design a corresponding region recognition function for each Gaussian kernel in the 4D Gaussian model. Each region recognition function uses its own rendering discrimination condition to determine the renderable region of the Gaussian kernel. For any point on the 2D image plane, determine in turn whether the 2D coordinates of the point meet each rendering discrimination condition: if a rendering discrimination condition is met, the value of the region recognition function corresponding to the rendering discrimination condition at the point is 1 and the Gaussian kernel needs to render the point; otherwise, the value of the region recognition function at the point is 0 and the point is not rendered. Until all points on the 2D image plane are judged, the set of points to be rendered corresponding to each Gaussian kernel is output.
[0011] S3: The 4D Gaussian model is trained using the training data, and a learnable Alpha parameter is set for each Gaussian kernel during the training process for self-heterogeneity learning. Each Gaussian kernel is rendered on its corresponding set of points to be rendered. When the rendering transparency is calculated, the opacity of each Gaussian kernel is adjusted using a piecewise activation function controlled by the Alpha parameter. All adjusted rendering pixel opacities are passed through the renderer in the 4D Gaussian model to form a rendered image. The total loss is constructed by the mean square error loss and the SSIM loss between the rendered image and the original image. The 4D Gaussian model parameters are updated based on minimizing the total loss. After multiple rounds of iterations, the 4D Gaussian model for reconstruction is obtained.
[0012] S4: Input the specified time parameters and viewing angle parameters into the trained 4D Gaussian model to generate the corresponding dynamic scene rendering and complete the 4D Gaussian reconstruction.
[0013] Based on the above solution, each step can be implemented in the following preferred specific manner.
[0014] As a preferred embodiment of the first aspect, in step S1, the specific process of accurately parameterizing the original timestamp and the original viewing angle is as follows:
[0015] S11: The original observation perspective is processed by the projection matrix to transform it from the world coordinate system to the image plane coordinate system to form a processed observation perspective;
[0016] S12: The discrete original timestamp is mapped to the interval [0, 1] using a normalization method to form a processed timestamp.
[0017] As a preferred embodiment of the first aspect, in step S2, the 2D coordinates of any point on the 2D image plane are determined. The specific process of whether a rendering judgment condition is met is: obtaining the major axis size of the ellipse formed by the projection of a Gaussian kernel on the 2D image plane and the minor axis size of the ellipse , and take the eigenvector corresponding to the major axis of the ellipse as the first eigenvector , take the eigenvector corresponding to the minor axis of the ellipse as the second eigenvector , the dot product result of the 2D coordinate and the first eigenvector is taken as the first dot product result, the ratio of the first dot product result to the size of the major axis of the ellipse is taken as the first ratio, the dot product result of the 2D coordinate and the second eigenvector is taken as the second dot product result, the ratio of the second dot product result to the size of the minor axis of the ellipse is taken as the second ratio, the size of the major axis of the ellipse and its opposite are used as the two endpoints of the first discrimination closed interval, and the size of the minor axis of the ellipse and its opposite are used as the two endpoints of the second discrimination closed interval. When the first ratio is within the first discrimination closed interval and the second ratio is within the second discrimination closed interval, the 2D coordinate of the point meets the rendering discrimination condition, the value of the region recognition function corresponding to the rendering discrimination condition at the point is 1, and the Gaussian kernel corresponding to the region recognition function needs to render the point; otherwise, the 2D coordinate of the point does not meet the rendering discrimination condition, the value of the region recognition function corresponding to the rendering discrimination condition at the point is 0, and the Gaussian kernel corresponding to the region recognition function does not need to render the point.
[0018] Furthermore, in step S2, the region recognition function is specifically expressed as follows:
[0019]
[0020] in, Represents the dot product operation of two vectors; Indicates that the region recognition function is in 2D coordinates The value at the point.
[0021] As a preferred embodiment of the first aspect, the process of calculating the major axis size and the minor axis size of the ellipse formed by the Gaussian kernel projection is as follows:
[0022] S21: Calculate the 3D covariance matrix of each Gaussian kernel separately to obtain the 2D covariance matrix of each Gaussian kernel projected onto the image plane;
[0023] S22: Solve the two eigenvalues corresponding to the 2D covariance matrix of each Gaussian kernel, and perform calculations based on the obtained two eigenvalues to obtain the major axis size and the minor axis size of the ellipse projected onto the 2D image plane by the Gaussian kernel.
[0024] As a preferred embodiment of the first aspect, in step S21, a 2D covariance matrix of a Gaussian kernel is obtained as follows: the Jacob matrix , view matrix , the 3D covariance matrix of the Gaussian kernel The 2D covariance matrix obtained after the Gaussian kernel projection is obtained by multiplying the five parts of the perspective matrix transpose and the Jacob matrix transpose. .
[0025] Furthermore, in step S21, the 2D covariance matrix of a Gaussian kernel projected onto the image plane can be calculated as follows:
[0026]
[0027] in, Represents matrix transpose.
[0028] As a preferred embodiment of the first aspect, in step S22, the major axis of the ellipse is three times the arithmetic square root of the larger of the two eigenvalues, and the minor axis of the ellipse is three times the arithmetic square root of the smaller of the two eigenvalues.
[0029] Furthermore, in step S22, the two eigenvalues corresponding to the 2D covariance matrix Calculate as follows:
[0030]
[0031]
[0032] in, They correspond to the values in the 2D covariance matrix respectively.
[0033] Furthermore, the size of the major axis of the ellipse and the minor axis size of the ellipse Respectively expressed as:
[0034]
[0035]
[0036] in, represents the larger of the two eigenvalues of the 2D covariance matrix; Represents the smaller of the two eigenvalues of the 2D covariance matrix.
[0037] As a preferred embodiment of the first aspect, in step S3, Gaussian kernel adjusted opacity of rendered pixels By The self-opacity of the Gaussian kernel And the function value of the piecewise activation function The function value is obtained by adding two parts, the first part is the coefficient of the piecewise activation function and the power term calculated from the 2D covariance matrix The second part is the calculation result of the residual term in the piecewise activation function .
[0038] Further, the Gaussian kernel adjusted opacity The specific expressions are as follows:
[0039]
[0040] As a preferred embodiment of the above-mentioned first aspect, the coefficients are all generated based on the learnable Alpha parameter; when the power term is greater than or equal to 0 and less than or equal to the preset first threshold, the residual term is 0; when the power term is greater than the first threshold and less than or equal to the preset second threshold, the residual term is generated based on the learnable Alpha parameter; when the power term is greater than the second threshold and less than or equal to 1, the residual term is generated based on the learnable Alpha parameter.
[0041] As a preferred embodiment of the first aspect, the learnable Alpha parameters are optimized using a reverse gradient propagation method.
[0042] In a second aspect, the present invention provides a 4D Gaussian reconstruction system based on self-heterogeneity, comprising:
[0043] A data acquisition module is used to accurately parameterize the input original timestamp and original observation angle, and combine the processed timestamp and the processed observation angle with the original image of the corresponding timestamp and observation angle to form training data;
[0044] The region recognition module is used to design a corresponding region recognition function for each Gaussian kernel in the 4D Gaussian model. Each region recognition function uses its own rendering discrimination condition to determine the renderable region of the Gaussian kernel. For any point on the 2D image plane, it is sequentially determined whether the 2D coordinates of the point meet each rendering discrimination condition: if a rendering discrimination condition is met, the value of the region recognition function corresponding to the rendering discrimination condition at the point is 1 and the Gaussian kernel needs to render the point; otherwise, the value of the region recognition function at the point is 0 and the point is not rendered. After all points on the 2D image plane have been identified, the set of points to be rendered corresponding to each Gaussian kernel is output.
[0045] A model training module is used to train a 4D Gaussian model using training data, and to set a learnable Alpha parameter for each Gaussian kernel during the training process for self-heterogeneity learning. Each Gaussian kernel is rendered on its corresponding set of points to be rendered. When the rendering transparency is calculated, the self-opaqueness of each Gaussian kernel is adjusted using a piecewise activation function controlled by the Alpha parameter. All adjusted rendering pixel opacities are passed through a renderer in the 4D Gaussian model to form a rendered image. A total loss is constructed by the mean square error loss and the SSIM loss between the rendered image and the original image. The 4D Gaussian model parameters are updated based on minimizing the total loss. After multiple rounds of iterations, a 4D Gaussian model for reconstruction is obtained.
[0046] The reconstruction module is used to input the specified time parameters and viewing angle parameters into the trained 4D Gaussian model, generate the corresponding dynamic scene rendering, and complete the 4D Gaussian reconstruction.
[0047] Compared with the prior art, the present invention has the following beneficial effects:
[0048] This paper proposes a self-heterogeneous 4D Gaussian reconstruction method for training with limited, specific style images, which can quickly achieve high-fidelity 4D reconstruction results. In this method, a learnable alpha parameter is assigned to each Gaussian kernel for self-heterogeneous learning. The core concept is to decouple alpha transparency from Gaussian shape, breaking the fixed structure of traditional template-based alpha distributions. Furthermore, based on this alpha parameter, the present invention also designs a piecewise activation function controlled by the alpha parameter. This function allows each Gaussian to adaptively adjust its alpha profile based on the frequency characteristics of its region. In high-frequency regions, sharp alpha boundaries are retained to capture clear motion or texture changes; in low-frequency regions, a flatter alpha distribution is adopted to enhance spatial consistency and reduce unnecessary overlap. This structural decoupling allows each Gaussian to perform functional division of labor between different tasks, fundamentally alleviating the optimization conflict between appearance modeling and motion modeling. To further improve representation efficiency and optimization results, the present invention introduces a region recognition function. Traditional 3D GS requires alpha calculation for all Gaussians during the rendering process, which involves a large number of redundant operations that do not contribute substantially to the final image. The method of the present invention significantly reduces unnecessary computation and memory overhead by quickly removing these "invalid Gaussians" before alpha calculation. This strategy not only speeds up convergence during training but also provides a "zero-cost improvement" during inference, further enhancing the practical applicability of the method.
[0049] The present invention and the existing 4D Gaussian reconstruction method have been extensively experimented on multiple data sets. The results show that the method of the present invention significantly reduces the motion blur phenomenon without increasing the inference cost. Compared with the existing methods, the method of the present invention has achieved excellent performance in multiple indicators such as spatiotemporal consistency, boundary clarity, and reconstruction accuracy. In addition, since the method of the present invention is an improvement to the representation structure itself, it can be seamlessly integrated with existing motion modeling or rendering acceleration technology, and has good scalability and engineering practical value. In summary, the present invention starts from the root of the problem and systematically solves the problem of motion blur in current 4D reconstruction, providing a more efficient, accurate and universal solution for dynamic image modeling. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flow chart of the steps of the method of the present invention;
[0051] Figure 2 This is a comparison chart of rendering results when using learnable Alpha parameters and when not using learnable Alpha parameters in this embodiment;
[0052] Figure 3 This is a comparison chart of rendering results of different methods on the DNeRF dataset in this embodiment;
[0053] Figure 4 This is a comparison chart of rendering results of different methods on the Neu3D dataset in this embodiment;
[0054] Figure 5 This is a system block diagram of the present invention. DETAILED DESCRIPTION
[0055] In order to make the above-mentioned objects, features and advantages of the present invention more clearly understood, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined accordingly without conflicting with each other.
[0056] In the description of the present invention, it should be understood that the terms "first" and "second" are used solely for descriptive purposes and are not to be construed as indicating or implying relative importance or implicitly specifying the number of technical features being described. Therefore, features defined as "first" or "second" may explicitly or implicitly include at least one of such features.
[0057] like Figure 1 As shown, in a preferred implementation of the present invention, a 4D Gaussian reconstruction method based on self-heterogeneity is provided. This method is based on invariant representation learning and can achieve fast, high-fidelity 4D reconstruction. This 4D Gaussian reconstruction method based on self-heterogeneity includes the following steps S1 to S4. The specific implementation process is described below.
[0058] S1: Accurately parameterize the input original timestamp and original observation angle respectively, and combine the processed timestamp and the processed observation angle with the original image of the corresponding timestamp and observation angle to form training data.
[0059] It should be noted that step S1 provides temporal and spatially aligned supervision information for the subsequent training of the 4D Gaussian model.
[0060] In step S1 of this embodiment, the specific process of accurately parameterizing the original timestamp and the original viewing angle is as follows:
[0061] S11: The original observation perspective is processed by the projection matrix to transform it from the world coordinate system to the image plane coordinate system to form a processed observation perspective;
[0062] S12: The discrete original timestamp is mapped to the interval [0, 1] using a normalization method to form a processed timestamp.
[0063] S2: Design a corresponding region recognition function for each Gaussian kernel in the 4D Gaussian model. Each region recognition function uses its own rendering discrimination condition to determine the renderable region of the Gaussian kernel. For any point on the 2D image plane, determine in turn whether the 2D coordinates of the point meet each rendering discrimination condition: if a rendering discrimination condition is met, the value of the region recognition function corresponding to the rendering discrimination condition at the point is 1 and the Gaussian kernel needs to render the point; otherwise, the value of the region recognition function at the point is 0 and the point is not rendered. Until all points on the 2D image plane are judged, the set of points to be rendered corresponding to each Gaussian kernel is output.
[0064] It should be noted that in step S2, the renderable area of the Gaussian kernel is clipped by the rendering judgment condition, aiming to identify and remove redundant Gaussian kernel parts that contribute little or no to the final rendering result, thereby avoiding unnecessary rasterization calculations and significantly improving rendering efficiency.
[0065] In step S2 of the present invention, the 2D coordinates of any point on the 2D image plane are determined. The specific process of whether a rendering judgment condition is met is: obtaining the major axis size of the ellipse formed by the projection of a Gaussian kernel on the 2D image plane and the minor axis size of the ellipse , and take the eigenvector corresponding to the major axis of the ellipse as the first eigenvector , take the eigenvector corresponding to the minor axis of the ellipse as the second eigenvector , the dot product result of the 2D coordinate and the first eigenvector is taken as the first dot product result, the ratio of the first dot product result to the size of the major axis of the ellipse is taken as the first ratio, the dot product result of the 2D coordinate and the second eigenvector is taken as the second dot product result, the ratio of the second dot product result to the size of the minor axis of the ellipse is taken as the second ratio, the size of the major axis of the ellipse and its opposite are used as the two endpoints of the first discrimination closed interval, and the size of the minor axis of the ellipse and its opposite are used as the two endpoints of the second discrimination closed interval. When the first ratio is within the first discrimination closed interval and the second ratio is within the second discrimination closed interval, the 2D coordinate of the point meets the rendering discrimination condition, the value of the region recognition function corresponding to the rendering discrimination condition at the point is 1, and the Gaussian kernel corresponding to the region recognition function needs to render the point; otherwise, the 2D coordinate of the point does not meet the rendering discrimination condition, the value of the region recognition function corresponding to the rendering discrimination condition at the point is 0, and the Gaussian kernel corresponding to the region recognition function does not need to render the point.
[0066] In step S2 of this embodiment, before the Gaussian rasterization calculation, a corresponding region identification function is matched for each Gaussian kernel. The region identification function is specifically expressed as follows:
[0067]
[0068] in, Represents the dot product operation of two vectors; Indicates that the region recognition function is in 2D coordinates The value at the point.
[0069] It should be noted that in the present invention, the ellipse projected by the Gaussian kernel on the 2D image plane and the range covered by the ellipse can be calculated based on the 2D covariance matrix of the Gaussian kernel. Since the ellipse is obtained by the 2D projection of the Gaussian kernel, the range within the three eigenvalues of the Gaussian distribution is visible. Based on this, the major and minor axes of the obtained ellipse can be calculated based on the 2D covariance and three eigenvalues of the Gaussian kernel.
[0070] Specifically, the process of calculating the major axis size and minor axis size of the ellipse formed by the Gaussian kernel projection is:
[0071] S21: Calculate the 3D covariance matrix of each Gaussian kernel separately to obtain the 2D covariance matrix of each Gaussian kernel projected onto the image plane.
[0072] In S21 of the present invention, the specific process of obtaining a 2D covariance matrix of a Gaussian kernel is as follows: the Jacob matrix , view matrix , the 3D covariance matrix of the Gaussian kernel The 2D covariance matrix obtained after the Gaussian kernel projection is obtained by multiplying the five parts of the perspective matrix transpose and the Jacob matrix transpose. .
[0073] In this embodiment S21, the 2D covariance matrix of a Gaussian kernel projected onto the image plane can be calculated as follows:
[0074]
[0075] in, Represents matrix transpose; Represents the viewing angle matrix, which is a parameter used by the camera to record the viewing angle during shooting.
[0076] In addition, the center coordinates of the Gaussian kernel can be With the obtained view matrix and projection matrix Multiply them to get the 2D coordinates of the image plane after the Gaussian kernel projection :
[0077]
[0078] S22: Solve the two eigenvalues corresponding to the 2D covariance matrix of each Gaussian kernel, and perform calculations based on the obtained two eigenvalues to obtain the major axis size and the minor axis size of the ellipse projected onto the 2D image plane by the Gaussian kernel.
[0079] In S22 of the present invention, the size of the major axis of the ellipse is three times the arithmetic square root of the larger eigenvalue of the two eigenvalues, and the size of the minor axis of the ellipse is three times the arithmetic square root of the smaller eigenvalue of the two eigenvalues.
[0080] In this embodiment S22, the two eigenvalues corresponding to the 2D covariance matrix Calculate as follows:
[0081]
[0082]
[0083] in, Corresponding to the values in the 2D covariance matrix. Therefore, the size of the major axis of the ellipse and the minor axis size of the ellipse Respectively expressed as:
[0084]
[0085]
[0086] in, represents the larger of the two eigenvalues; represents the smaller of the two eigenvalues.
[0087] In the present invention, the first eigenvector and the second eigenvector are obtained by solving the following formula:
[0088]
[0089] in, Represented by the 2D covariance matrix The eigenvector formed by the two eigenvalues of represents the identity matrix; represents a characteristic matrix composed of the first eigenvector and the second eigenvector, the first row of the characteristic matrix is the first eigenvector, and the second row is the second eigenvector.
[0090] S3: The 4D Gaussian model is trained using the training data, and a learnable Alpha parameter is set for each Gaussian kernel during the training process for self-heterogeneity learning. Each Gaussian kernel is rendered on its corresponding set of points to be rendered. When the rendering transparency is calculated, the opacity of each Gaussian kernel is adjusted using a piecewise activation function controlled by the Alpha parameter. All adjusted rendering pixel opacities are passed through the renderer in the 4D Gaussian model to form a rendered image. The total loss is constructed by the mean square error loss and the SSIM loss between the rendered image and the original image. The 4D Gaussian model parameters are updated based on minimizing the total loss. After multiple rounds of iterations, the 4D Gaussian model for reconstruction is obtained.
[0091] It should be noted that in step S3 of the present invention, the Alpha parameter is adaptively learned during the training process of the 4D Gaussian model to achieve self-heterogeneity changes of the Gaussian kernel, so that it can accurately adjust the transparency or opacity characteristics according to the scene dynamics and its own contribution.
[0092] In the present invention, Gaussian kernel adjusted opacity of rendered pixels By The self-opacity of the Gaussian kernel and piecewise activation function The function value is obtained by multiplying the function value of the segmented activation function. The function value is obtained by adding two parts. The first part is the coefficient of the segmented activation function. and the power term calculated from the 2D covariance matrix The second part is the calculation result of the residual term in the piecewise activation function. express.
[0093] Furthermore, the coefficients are all generated based on a learnable Alpha parameter; when the power term is greater than or equal to 0 and less than or equal to a preset first threshold, the residual term is 0; when the power term is greater than the first threshold and less than or equal to a preset second threshold, the residual term is generated based on a learnable Alpha parameter; when the power term is greater than the second threshold and less than or equal to 1, the residual term is generated based on a learnable Alpha parameter.
[0094] In this embodiment, considering that there is significant blur when reconstructing dynamic 3D scenes using existing methods, a classic scene in dynamic 3D reconstruction is used as an example. Figure 2 As shown, the finger area of the method without alpha self-heterogeneity exhibits significant blurring, and the problem is further exacerbated when the finger model moves from the first moment to the second. Previous methods attributed this problem to the inaccurate driving of the Gaussian kernel by the deformable net. The present invention's analysis reveals that the instability of the Gaussian kernel's alpha value during rendering interferes with the deformable net's control of the Gaussian kernel's motion. Specifically, because the mean square error (MSE) is used to calculate and optimize the error between the rendered and ground truth images, the deformable net must drive the Gaussian kernel to different positions at each moment to render the image. When there is a gap between the two Gaussian kernels at the first moment (i.e., the pixel's alpha value is too small, resulting in a transparent appearance), the optimization method based on minimizing the mean square error loss forces the deformable net to not only learn the true motion direction but also manipulate the Gaussian kernels to stack together to fill the gap (i.e., increasing the pixel's alpha, resulting in a more opaque appearance). However, the vector sum direction (i.e., the deformable net's driving direction) does not correspond to the true motion direction, which is the cause of the significant blurring in dynamic 3D reconstruction using the method without alpha self-heterogeneity.
[0095] To this end, this paper proposes an alpha-autoheterogeneity method, which manipulates the alpha value of Gaussian kernel rendering to enable the deformable net to learn correct motion. The core idea is to calculate the alpha value of each Gaussian kernel at each pixel during rendering and control it into two types: one type of Gaussian kernel consistently represents pixels with small alpha values, and the other type of Gaussian kernel consistently represents pixels with large alpha values. This allows the deformable net to focus on the positional mapping between Gaussian kernels representing pixels with similar alpha values, i.e., the correct motion relationship.
[0096] Based on the above analysis, this embodiment generalizes the existing Gaussian rasterization formula to obtain the corresponding general rendering equation:
[0097]
[0098] in, Represents the opacity of rendered pixels for rasterization (i.e., the adjusted opacity of rendered pixels); Represents the piecewise activation function controlled by the Alpha parameter; Indicates the The self-opacity of the Gaussian kernel; represents the power term calculated based on the 2D covariance matrix; Represents the coefficients in the piecewise activation function; Represents the residual term in the piecewise activation function.
[0099] In this embodiment, the above-mentioned piecewise activation function is used to remap the rasterization formula. After setting the first threshold to 0.25 and the second threshold to 0.75, the functional form of the above-mentioned piecewise activation function is:
[0100]
[0101] in, This is the learnable Alpha parameter introduced in the present invention. When the power term is greater than or equal to 0 and less than or equal to 0.25, the coefficient is , the residual term is 0; when the power term is greater than 0.25 and less than or equal to 0.75, the coefficient is , the residual term is ; When the power term is 0.75 and less than or equal to 1, the coefficient is , the residual term is .
[0102] It should be noted that in step S3 of the present invention, the learnable Alpha parameter is optimized using the reverse gradient propagation method. In this embodiment, according to the functional form of the above-mentioned segmented activation function, the learnable Alpha parameter The corresponding partial derivatives are:
[0103]
[0104] In this embodiment, when training a 4D Gaussian model using training data, a set of training data (including a processed timestamp, a processed viewing angle, and an original image) is given and input into the 4D Gaussian model. A rendered image corresponding to the processed viewing angle and timestamp is obtained. The rendered image is then compared with the original image and the mean square error loss and the mean square error (SSIM) loss are calculated. After multiple rounds of iterations, the 4D Gaussian model used for reconstruction is obtained.
[0105] S4: Input the specified time parameters and viewing angle parameters into the trained 4D Gaussian model to generate the corresponding dynamic scene rendering and complete the 4D Gaussian reconstruction.
[0106] It should be noted that the 4D Gaussian model trained through the above steps supports efficient and high-quality rendering. Specifically, by directly inputting the specified time parameters and viewing angle parameters, the corresponding dynamic scene rendering can be quickly generated.
[0107] The 4D Gaussian reconstruction method based on self-heterogeneity in the above embodiment is applied to a specific data set for classification testing. The specific steps are as described in S1-S4 and will not be repeated here. The specific parameters and technical effects are mainly shown.
[0108] Example
[0109] This embodiment follows the implementation process of the aforementioned S1 to S4 steps. To quantify the indicators, this embodiment conducts qualitative and quantitative experiments on the DNeRF and Neu3D datasets, and compares the method of the present invention with the existing technologies 4D-GS, DeformGS, SC-GS, Grid4D, Tensor-4D, HexPlane, and TiNeuVox-B. The comparison results on the DNeRF dataset are shown in Figure 2. Figure 3 As shown in the figure, the comparison results on the Neu3D dataset are as follows Figure 4 The quantitative evaluation results on the DNeRF dataset are shown in Tables 1, 2, and 3. Figure 3 In the figure, the first comparison method refers to the prior art 4D-GS, the second comparison method refers to the prior art DeformGS, the third comparison method refers to the prior art SC-GS, and the fourth comparison method refers to the prior art Grid4D.
[0110] Table 1. PSNR quantification results of test data
[0111] Table 2. SSIM quantization results of test data
[0112] Table 3. LPIPS quantization results of test data
[0113] It can be clearly observed from Tables 1 to 3 that the present invention has achieved the best quantitative indicators in terms of fidelity, consistency and visual perception, which shows that the present invention can accurately reconstruct dynamic 4D scenes. Figure 3 and Figure 4 It can be clearly observed that the present invention has a clear advantage in reconstructing details, which is due to the Gaussian kernel proposed in the present invention, which can maintain stability during motion, thereby accurately reconstructing set and texture information during training.
[0114] It should also be noted that the 4D Gaussian reconstruction method based on self-heterogeneity in the above embodiment can essentially be executed by a computer program or module. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a 4D Gaussian reconstruction system based on self-heterogeneity corresponding to the 4D Gaussian reconstruction method based on self-heterogeneity provided in the above embodiment, such as Figure 5 As shown, it includes:
[0115] A data acquisition module is used to accurately parameterize the input original timestamp and original observation angle, and combine the processed timestamp and the processed observation angle with the original image of the corresponding timestamp and observation angle to form training data;
[0116] The region recognition module is used to design a corresponding region recognition function for each Gaussian kernel in the 4D Gaussian model. Each region recognition function uses its own rendering discrimination condition to determine the renderable region of the Gaussian kernel. For any point on the 2D image plane, it is sequentially determined whether the 2D coordinates of the point meet each rendering discrimination condition: if a rendering discrimination condition is met, the value of the region recognition function corresponding to the rendering discrimination condition at the point is 1 and the Gaussian kernel needs to render the point; otherwise, the value of the region recognition function at the point is 0 and the point is not rendered. After all points on the 2D image plane have been identified, the set of points to be rendered corresponding to each Gaussian kernel is output.
[0117] A model training module is used to train a 4D Gaussian model using training data, and to set a learnable Alpha parameter for each Gaussian kernel during the training process for self-heterogeneity learning. Each Gaussian kernel is rendered on its corresponding set of points to be rendered. When the rendering transparency is calculated, the self-opaqueness of each Gaussian kernel is adjusted using a piecewise activation function controlled by the Alpha parameter. All adjusted rendering pixel opacities are passed through a renderer in the 4D Gaussian model to form a rendered image. A total loss is constructed by the mean square error loss and the SSIM loss between the rendered image and the original image. The 4D Gaussian model parameters are updated based on minimizing the total loss. After multiple rounds of iterations, a 4D Gaussian model for reconstruction is obtained.
[0118] The reconstruction module is used to input the specified time parameters and viewing angle parameters into the trained 4D Gaussian model, generate the corresponding dynamic scene rendering, and complete the 4D Gaussian reconstruction.
[0119] It should also be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the various embodiments provided in this application, the division of steps or modules in the system and method is only a logical function division. In actual implementation, there may be other division methods, for example, multiple modules or steps can be combined or integrated together, and a module or step can also be split.
[0120] The embodiment described above is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Persons skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent substitution or equivalent transformation falls within the scope of protection of the present invention.
Claims
1. A 4D Gaussian reconstruction method based on self-heterogeneity, characterized in that: The following steps are involved: S1: Accurately parameterize the input original timestamp and original observation angle, and combine the processed timestamp and the processed observation angle with the original image of the corresponding timestamp and observation angle to form training data; S2: Design a corresponding region recognition function for each Gaussian kernel in the 4D Gaussian model. Each region recognition function uses its own rendering discrimination condition to determine the renderable region of the Gaussian kernel. For any point on the 2D image plane, determine in turn whether the 2D coordinates of the point meet each rendering discrimination condition: if a rendering discrimination condition is met, the value of the region recognition function corresponding to the rendering discrimination condition at the point is 1 and the Gaussian kernel needs to render the point; otherwise, the value of the region recognition function at the point is 0 and the point is not rendered. Until all points on the 2D image plane are judged, the set of points to be rendered corresponding to each Gaussian kernel is output. S3: The 4D Gaussian model is trained using the training data, and a learnable Alpha parameter is set for each Gaussian kernel during the training process for self-heterogeneity learning. Each Gaussian kernel is rendered on its corresponding set of points to be rendered. When the rendering transparency is calculated, the opacity of each Gaussian kernel is adjusted using a piecewise activation function controlled by the Alpha parameter. All adjusted rendering pixel opacities are passed through the renderer in the 4D Gaussian model to form a rendered image. The total loss is constructed by the mean square error loss and the SSIM loss between the rendered image and the original image. The 4D Gaussian model parameters are updated based on minimizing the total loss. After multiple rounds of iterations, the 4D Gaussian model for reconstruction is obtained. S4: Input the specified time parameters and viewing angle parameters into the trained 4D Gaussian model to generate the corresponding dynamic scene rendering and complete the 4D Gaussian reconstruction.
2. The 4D Gaussian reconstruction method based on self-heterogeneity according to claim 1, characterized in that: In step S1, the specific process of accurately parameterizing the original timestamp and the original observation angle is as follows: S11: The original observation perspective is processed by the projection matrix to transform it from the world coordinate system to the image plane coordinate system to form a processed observation perspective; S12: The discrete original timestamp is mapped to the interval [0, 1] using a normalization method to form a processed timestamp.
3. The 4D Gaussian reconstruction method based on self-heterogeneity according to claim 1, characterized in that: In step S2, the specific process of determining whether the 2D coordinates of any point on the 2D image plane meet a rendering judgment condition is as follows: obtaining the size of the major axis and the minor axis of the ellipse formed by the projection of a Gaussian kernel on the 2D image plane, and taking the eigenvector corresponding to the major axis of the ellipse as the first eigenvector, and taking the eigenvector corresponding to the minor axis of the ellipse as the second eigenvector, taking the dot product result of the 2D coordinate and the first eigenvector as the first dot product result, taking the ratio of the first dot product result to the size of the major axis of the ellipse as the first ratio, and taking the 2D coordinate and the second eigenvector as the first ratio. The dot product result of is used as the second dot product result, the ratio of the second dot product result to the size of the minor axis of the ellipse is used as the second ratio, the size of the major axis of the ellipse and its opposite are used as the two endpoints of the first discrimination closed interval, and the size of the minor axis of the ellipse and its opposite are used as the two endpoints of the second discrimination closed interval. When the first ratio is within the first discrimination closed interval and the second ratio is within the second discrimination closed interval, the 2D coordinates of the point meet the rendering discrimination condition, the value of the region recognition function corresponding to the rendering discrimination condition at the point is 1, and the Gaussian kernel corresponding to the region recognition function needs to render the point; Otherwise, the 2D coordinate of the point does not meet the rendering discrimination condition, the value of the region recognition function corresponding to the rendering discrimination condition at the point is 0, and the Gaussian kernel corresponding to the region recognition function does not need to render the point.
4. The 4D Gaussian reconstruction method based on self-heterogeneity according to claim 3, wherein: The process of calculating the major axis and minor axis of the ellipse formed by the Gaussian kernel projection is: S21: Calculate the 3D covariance matrix of each Gaussian kernel separately to obtain the 2D covariance matrix of each Gaussian kernel projected onto the image plane; S22: Solve the two eigenvalues corresponding to the 2D covariance matrix of each Gaussian kernel, and perform calculations based on the obtained two eigenvalues to obtain the major axis size and the minor axis size of the ellipse projected onto the 2D image plane by the Gaussian kernel.
5. The 4D Gaussian reconstruction method based on self-heterogeneity according to claim 4, characterized in that: In step S21, the specific process of obtaining a 2D covariance matrix of a Gaussian kernel is as follows: the Jacob matrix, the view matrix, the 3D covariance matrix of the Gaussian kernel, the view matrix transpose, and the Jacob matrix transpose are multiplied to obtain the 2D covariance matrix obtained after the Gaussian kernel is projected.
6. The 4D Gaussian reconstruction method based on self-heterogeneity according to claim 4, characterized in that: In step S22 , the size of the major axis of the ellipse is three times the arithmetic square root of the larger eigenvalue of the two eigenvalues, and the size of the minor axis of the ellipse is three times the arithmetic square root of the smaller eigenvalue of the two eigenvalues.
7. The 4D Gaussian reconstruction method based on self-heterogeneity according to claim 1, characterized in that: In step S3, the opacity of the rendered pixel after adjustment by the i-th Gaussian kernel is obtained by multiplying the opacity of the Gaussian kernel itself and the function value of the piecewise activation function, and the function value is obtained by adding two parts, the first part is the multiplication result between the coefficient of the piecewise activation function and the power term calculated according to the 2D covariance matrix, and the second part is the calculation result of the residual term in the piecewise activation function.
8. The 4D Gaussian reconstruction method based on self-heterogeneity according to claim 7, characterized in that: The coefficients are all generated based on the learnable Alpha parameter; when the power term is greater than or equal to 0 and less than or equal to the preset first threshold, the residual term is 0; when the power term is greater than the first threshold and less than or equal to the preset second threshold, the residual term is generated based on the learnable Alpha parameter; when the power term is greater than the second threshold and less than or equal to 1, the residual term is generated based on the learnable Alpha parameter.
9. The 4D Gaussian reconstruction method based on self-heterogeneity according to claim 8, characterized in that: Use the back gradient propagation method to optimize the learnable Alpha parameters.
10. A 4D Gaussian reconstruction system based on self-heterogeneity, characterized in that: include: A data acquisition module is used to accurately parameterize the input original timestamp and original observation angle, and combine the processed timestamp and the processed observation angle with the original image of the corresponding timestamp and observation angle to form training data; The region recognition module is used to design a corresponding region recognition function for each Gaussian kernel in the 4D Gaussian model. Each region recognition function uses its own rendering discrimination condition to determine the renderable region of the Gaussian kernel. For any point on the 2D image plane, it is sequentially determined whether the 2D coordinates of the point meet each rendering discrimination condition: if a rendering discrimination condition is met, the value of the region recognition function corresponding to the rendering discrimination condition at the point is 1 and the Gaussian kernel needs to render the point; otherwise, the value of the region recognition function at the point is 0 and the point is not rendered. After all points on the 2D image plane have been identified, the set of points to be rendered corresponding to each Gaussian kernel is output. A model training module is used to train a 4D Gaussian model using training data, and to set a learnable Alpha parameter for each Gaussian kernel during the training process for self-heterogeneity learning. Each Gaussian kernel is rendered on its corresponding set of points to be rendered. When the rendering transparency is calculated, the self-opaqueness of each Gaussian kernel is adjusted using a piecewise activation function controlled by the Alpha parameter. All adjusted rendering pixel opacities are passed through a renderer in the 4D Gaussian model to form a rendered image. A total loss is constructed by the mean square error loss and the SSIM loss between the rendered image and the original image. The 4D Gaussian model parameters are updated based on minimizing the total loss. After multiple rounds of iterations, a 4D Gaussian model for reconstruction is obtained. The reconstruction module is used to input the specified time parameters and viewing angle parameters into the trained 4D Gaussian model, generate the corresponding dynamic scene rendering, and complete the 4D Gaussian reconstruction.
Citation Information
Patent Citations
Novel view angle synthesis method based on Gaussian splash and fusing learnable basis function
CN118505541A
3D Gaussian sputtering scene reconstruction method based on view dependence difference decoupling
CN118710792A
Dynamic scene reconstruction method and device, equipment, medium and product
CN119169183A
FCN-based multivariate time series data classification method and device
US20220180129A1
Human subject gaussian splatting using machine learning
US20250148678A1