A three-dimensional gaussian animatable human avatar modeling method based on graph structure optimization
By constructing a static dressed Gaussian avatar template based on graph structure optimization and introducing multi-level kinematic relationships and material-sensitive motion propagation, the problem of insufficient pose generalization ability and local appearance dynamic confusion in monocular video 3D Gaussian human avatar modeling is solved. The appearance consistency and geometric continuity under complex pose changes are achieved, and the local deformation stability and visual realism in the animation process are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF TECH
- Filing Date
- 2026-03-04
- Publication Date
- 2026-06-05
AI Technical Summary
Existing monocular video 3D Gaussian human avatar modeling methods suffer from insufficient pose generalization ability, dynamic confusion of local appearance, and instability of non-rigid deformation in clothing areas under pose change conditions.
A graph-based optimization method is adopted to construct a static clothing Gaussian avatar template, and a structured modeling mechanism for multi-level kinematic relationships and material-sensitive motion propagation is introduced. By combining the graph structure with human kinematic topology and spatial neighborhood information, the structured propagation of motion and deformation signals is realized. The skin weight is progressively adaptively optimized by combining local neighborhood graph learning.
Maintaining consistent appearance and geometric continuity during complex posture changes enhances the stability of local deformations and visual realism in animation, thereby improving the dynamic stability and visual realism of the dressed human avatar.
Smart Images

Figure CN122156411A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of computer vision, 3D reconstruction and animable human avatar modeling, specifically involving a 3D Gaussian Splatting method for driving human avatar reconstruction based on monocular video, which can be used in scenarios such as virtual humans, digital twins, and XR / interactive rendering. Background Technology
[0002] Human avatar reconstruction is a key technology in applications such as virtual reality, digital human generation, remote interaction, and intelligent content creation. Its goal is to obtain photorealistic 3D human models capable of posture-driven operation under limited observation conditions. Due to the low cost and low barrier to entry of monocular video acquisition, constructing animable 3D human avatars based on monocular video has significant engineering value. However, monocular video suffers from challenges such as missing viewpoint information and coupling between occlusion and appearance changes. Especially in scenarios involving clothed individuals, clothing wrinkles, tightness, and non-rigid movements further amplify modeling errors, leading to problems such as unstable appearance, detail distortion, and unreasonable motion propagation during posture-driven operation. This negatively impacts the immersive experience and credibility of virtual performances and real-time interactions.
[0003] In recent years, 3D Gaussian Splatting (3DGS) has been widely used in 3D reconstruction and avatar modeling due to its differentiable rendering and efficient real-time compositing capabilities. Existing techniques for using 3DGS for animable human avatars typically combine parametric human models (such as SMPL) to provide structural priors and employ linear hybrid skinning (LBS) for skeletal drive. Simultaneously, to characterize clothing and dynamic surface changes, non-rigid deformation networks or point-level residual modules are often introduced to modulate Gaussian properties. While these methods can improve rendering detail and reconstruction efficiency to some extent, their dynamic modeling process largely relies on unstructured point-level learning: on the one hand, common practices use a single global pose code to uniformly modulate all Gaussian units, making it difficult to express the differences in motion patterns, semantic constraints, and interdependencies among different body parts; on the other hand, skinning weights are often directly inherited from the static weights of the human template or modified independently point-by-point, ignoring the consistency between local neighborhoods and the nonlinear deformation behavior caused by differences in clothing materials. The aforementioned limitations can easily lead to overfitting of the model within the training pose range, and cause unreasonable appearance changes, local dynamic confusion, and stretching and shaking of clothing surfaces when no pose is seen, thus reducing the robustness of pose transfer and the fidelity of local deformation.
[0004] To improve the dynamic consistency and realism under monocular conditions, existing research has begun to focus on introducing structured modeling mechanisms, such as graph learning methods based on skeleton, mesh, or point set topological relationships. Graph neural networks, through message passing between nodes and their neighborhoods, can model structural associations and consistency constraints while maintaining computational efficiency, and have been applied in human body deformation and clothing dynamic modeling. However, existing graph learning methods are often designed for multi-view or explicit mesh dynamic scenes, relying on stronger observation conditions or additional geometric constraints, making it difficult to directly adapt to the reconstruction of dressed human avatars under monocular video. At the same time, in 3D Gaussian representation, dynamic modulation involves not only pose-driven appearance changes but also the stability of skin weights and motion propagation. If only single-path graph modeling is used or there is a lack of collaborative characterization of kinematic hierarchy relationships and material-sensitive deformation, it is still difficult to ensure both component-level appearance stability and local geometric continuity under complex movements, thus limiting the generalization ability and visual quality of monocular animated Gaussian avatars.
[0005] To address the aforementioned issues, this invention proposes a monocular video 3D Gaussian animable human avatar modeling method based on graph structure optimization. This method constructs a static, dressed Gaussian avatar template and introduces a structured modeling mechanism that incorporates multi-level kinematic relationships and material-sensitive motion propagation. This improves the consistency of appearance and the fidelity of clothing details in unseen poses, and enhances the stability and physical plausibility of local deformations during animation. Summary of the Invention
[0006] This invention addresses the shortcomings of existing monocular video 3D Gaussian avatar modeling methods, such as insufficient pose generalization ability, dynamic confusion of local appearance, and instability due to non-rigid deformation of clothing areas. It proposes a graph-structure-optimized 3D Gaussian animated avatar modeling method. This invention aims to achieve animated avatar reconstruction that maintains appearance consistency and geometric continuity during complex pose changes by structurally modeling the human pose dependencies and local geometric propagation process under monocular acquisition conditions. This improves the dynamic stability and visual realism of the clothed avatar. Unlike methods that treat Gaussian primitives as independent or rely solely on a single global pose feature for modulation, the graph structure combines human kinematic topology and spatial neighborhood information to achieve structured propagation of motion and deformation signals during pose-driven processes, thus ensuring local continuity while suppressing unreasonable global coupling. Furthermore, by introducing hierarchical node representation and feature propagation mechanisms into the graph structure, decoupled modeling of pose semantics for different human body parts can be achieved. This allows each part to obtain a relatively independent and controllable dynamic response under pose changes, avoiding interference from global features on local actions. Meanwhile, the graph learning method based on local neighborhoods provides stable structural constraints for the progressive and adaptive optimization of skin weights, enabling non-rigid areas such as clothing to achieve more reasonable motion propagation under the guidance of material and geometric priors, thereby improving the geometric stability and realism in the animation process.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] Step 1: Construction of a static, dressed 3D Gaussian human avatar template. In the normalized pose space, a geometric prior for the human body is constructed based on a parametric human model, and 3D Gaussian primitives are generated on its surface. The geometric and appearance properties of the Gaussian primitives are jointly optimized to obtain a static, dressed 3D Gaussian human avatar template. This static template is used to establish unified and stable geometric topological relationships, pose alignment basis, and local material feature priors, thereby providing reliable initial constraints for subsequent pose-driven and dynamic modulation processes and avoiding the introduction of structural inconsistencies during the dynamic modeling stage.
[0009] Step 2: Dynamic Appearance Modulation Based on Multi-Level Posture Graphs. A multi-level posture graph is constructed based on human kinematics. The posture information of different parts of the human body is hierarchically modeled using a graph structure, and feature propagation and aggregation are performed within the graph to obtain part-level and global-level posture representations. These posture representations are mapped to corresponding 3D Gaussian elements. The geometric position, scale, and appearance attributes of the Gaussian elements are dynamically modulated according to posture, enabling different body parts to obtain relatively independent and controllable appearance responses under posture changes. This reduces local dynamic confusion caused by global posture features and improves appearance stability under unseen postures.
[0010] Step 3: Motion propagation optimization and pose generation based on material-aware skin map. A local neighborhood graph structure of 3D Gaussian primitives is constructed in the normalized space. Geometric deviation information, spatial location features, and material-related features extracted from the static template are integrated. A graph learning mechanism is used to perform structured propagation and adaptive correction of the skin weights. By introducing material-aware gating modulation, the rigid human body region and the flexible clothing region exhibit differentiated motion responses during pose-driven processes. This ensures local continuity while improving the stability and rationality of non-rigid deformation. Based on the optimized skin weights, pose transformation and rendering of the 3D Gaussian primitives are completed, generating an animated human avatar.
[0011] Through the above steps, this invention constructs a 3D Gaussian animable human avatar modeling method for monocular video. The method first establishes a static, clothed Gaussian avatar template to ensure the accuracy of posture alignment and local material priors. Then, based on the human kinematic structure and local topological relationships, it introduces a collaborative modeling mechanism of multi-level posture maps and material-aware skin maps to achieve structured constraints on posture-driven appearance modulation and motion propagation processes. Compared with existing 3D Gaussian avatar modeling methods that rely solely on global posture features or point-level independent optimization, this invention effectively suppresses local dynamic confusion and geometric propagation distortion during posture changes, enabling different body parts and clothed areas to obtain more stable and reasonable responses under posture-driven conditions. This improves the posture generalization ability and visual realism of animable human avatars under monocular conditions. Attached Figure Description
[0012] Figure 1 A monocular video human avatar reconstruction framework based on graph structure optimization of 3D Gaussian;
[0013] Figure 2 A rendering illustration of different 3D human avatar reconstruction results on the DynVideo dataset;
[0014] Figure 3 This is a rendering diagram of the reconstructed 3D human avatar generated on the DynVideo-male sequence under the condition of introducing different functional components. Detailed Implementation
[0015] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0016] Network framework introduction: static dressed Gaussian avatar template modeling, multi-level pose graph, material-aware skin map, loss function.
[0017] Static Dressed Gaussian Avatar Template Modeling: This stage aims to establish a static dressed Gaussian avatar template, accurately fit the pose, and construct Gaussian meta-features with material-aware capabilities. Unlike two typical strategies in previous methods: one relies on multi-view data (such as Animatable Gaussians) to explicitly construct the dressing template, achieving good geometric consistency but difficult to apply to monocular scenes; the other (such as GaussianAvatar, 3DGS-Avatar), while achieving Gaussian avatar reconstruction on monocular videos, ignores the rich material and relaxation information contained in the static template, resulting in a lack of stability of clothing surface details under dynamic driving. Therefore, we propose a static dressed Gaussian template modeling strategy starting from monocular videos, jointly optimizing the pose and geometric and appearance parameters of the dressing template, thus providing accurate 3D pose parameters and high-quality static priors for subsequent pose generalization and skinning optimization, such as... Figure 1 As shown in the blue module.
[0018] Using the SMPL model as the prior for human geometry, based on its canonical A-pose mesh... Gaussian elements are generated by uniformly sampling the surface based on UV mapping. The sampled points are projected onto the mesh surface via UV-to-3D mapping and used as the Gaussian center locations. Initial LBS weights The feature maps are obtained through SMPL vertex weight interpolation. This UV-based initialization ensures the geometric consistency of the Gaussian distribution with the surface manifold, providing a stable topological foundation for subsequent property learning. Learnable feature maps. and UV barycentric coordinate maps as surface manifold information Input to UNet encoding network Then, by the Gaussian attribute decoder Decoding yields the Gaussian parameters in the normalized space:
[0019]
[0020]
[0021] in , and These represent the position offset, scale, and color attributes in the canonical space, respectively. To ensure the avatar model's pose generalization ability, we use isotropic scaling for all Gaussian primitives and fix rotation and opacity. Gaussian points in the canonical space are transformed into pose space by skinning for rendering.
[0022]
[0023]
[0024] After convergence in the static modeling phase, the following eigenvector is defined for each Gaussian element: Standard deviation vector. , It explicitly quantifies the tightness of clothing in the canonical space, providing key input for the Material-Aware Skin Map (MASG); material surrogate vectors : Based on color characteristics The encoded data contains rich local surface material properties; it shares a three-plane position encoding ( ): Introducing a three-plane encoder, from the standard position Extracting low-dimensional location features As a unified spatial prior for MLPG and MASG, it is more effective in aggregating 3D spatial neighborhood information compared to traditional cosine position or UV encoding.
[0025] Multi-Level Pose Graph (MLPG): This paper proposes a multi-level pose graph (MLPG) for dynamic adjustment of pose-related parameters in component decoupling. Using the human kinematic chain as a priori, it achieves a hierarchical graph structure to represent human pose from local to global perspectives. Furthermore, it combines a graph neural network encoder and a lightweight decoder to achieve pose-driven Gaussian meta-attribute manipulation (dynamic adjustment of position, scale, and color brightness). Figure 1 As shown in the green module, existing 3D Gaussian-based human avatar methods typically use uniform global pose features to modulate all Gaussian points, ignoring the differences in motion patterns and semantic constraints among different body parts. This leads to two key problems: First, regarding local dynamic obfuscation, when different body regions move inconsistently (e.g., the arm swings while the torso remains still), global features struggle to reflect local changes. Second, regarding excessive coupling between pose and appearance changes, all Gaussian points share the same pose vector, causing appearance changes of different parts to be modulated in the same way, weakening the fine-grained modeling capability of Gaussian representations. To address these issues, we propose a hierarchical pose modeling framework based on the kinematic topology of SMPL. Through the collaborative design of hierarchical pose representation and contrast constraints with part-level pose embedding, we structurally achieve separation and spatial decoupling of pose semantics, enabling the model to achieve independent and controllable dynamic adjustments to different parts under complex movements.
[0026] For pose graph construction and multi-layer contrastive learning, MLPG employs a two-layer graph structure: a fine-grained layer (15 nodes, corresponding to body parts with low-level semantics) and a coarse-grained layer (7 nodes, corresponding to body parts with high-level semantics). The graph edges are constructed based on the kinematic hierarchy of SMPL, using parent-child dependencies between joints or parts as a priori, such as the bottom-up hierarchical connections of the spine, the branching topology of the limbs from root to tip, and the kinematic coupling between the trunk and limbs. The pose features input to each node are obtained by encoding SMPL rotation parameters using an MLP, and then processed through two lightweight GraphSAGE modules and residual aggregation.
[0027]
[0028]
[0029] in These are learnable parameters. They are obtained through hierarchical pose representation:
[0030]
[0031] To further enhance semantic decoupling between different levels, MLPG introduces a hierarchical contrastive learning mechanism in the 15 component and 7 region layers. Let the... layer( ) Node characteristics are ,definition:
[0032]
[0033] in The similarity between nodes is defined as:
[0034]
[0035] The loss function is:
[0036]
[0037]
[0038] This hierarchical contrast constraint promotes a more consistent feature distribution among neighboring components, while maintaining differences among components further away, forming a natural hierarchical separation structure in pose space, effectively mitigating the dynamic confusion problem caused by global embedding.
[0039] Based on component feature mapping and attribute decoding, after obtaining hierarchical pose features, we map them to the Gaussian metaspace, so that each Gaussian point only receives pose information related to its respective component. Let the Gaussian point be... The fine-grained and mid-layer component labels are respectively and ,but:
[0040]
[0041] and with three-plane position encoding By piecing them together, we get:
[0042]
[0043] This component-aware mapping approach offers several advantages: In terms of multi-layered driving consistency, local, mid-level, and global features work together on each Gaussian point, enabling it to respond simultaneously to local pose changes and overall appearance change trends; in terms of cross-frame pose stability, when local actions are consistent across different frames (e.g., the same arm pose), the corresponding Gaussian point obtains a consistent component embedding, thereby enhancing cross-frame consistency; and in terms of explicit decoupling of appearance change influence, pose features at different levels only act on the Gaussian subset of the corresponding semantic range, avoiding interference from global features with local appearance changes.
[0044] Lightweight 1D convolutional decoder As input, simultaneously predict the increments of three attributes (position, scale, and color brightness):
[0045]
[0046] in This represents the single-channel brightness change, which is replicated to the RGB three channels during application. Ultimately, the three types of increments are applied as residuals to the normalized Gaussian property:
[0047]
[0048] in For epoch-based linear warm-up factors, The scaling factor is used to ensure training stability. MLPG combines hierarchical pose representation, contrastive decoupling constraints, and part semantic mapping to enable Gaussian avatars to have fine-grained responsiveness at the local level while maintaining structural and appearance stability at the global level.
[0049] Material-Aware Skinning Graph (MASG): For refined skinning weight optimization, our proposed Material-Aware Skinning Graph (MASG) aims to adaptively optimize the initial linear blended skin (LBS) weights within the canonical space of a 3D Gaussian Avatar. MASG learns the correlation between geometric and material features within the local topological structure based on a Graph Neural Network (GNN), thereby modeling material-sensitive deformations in non-rigid regions such as clothing while maintaining geometric continuity. Figure 1 The purple module is shown in the middle. Traditional LBS skinning methods assume fixed weights for each point, which is insufficient when modeling clothing wrinkles, changes in tightness, and non-rigid motion. Fixed weights cannot adapt to nonlinear deformations under different materials, while directly optimizing high-dimensional weight fields can easily lead to training instability and cross-bone drift. Especially in Gaussian point representation, weight perturbations at clothing boundaries can cause rendering discontinuities and jitter. To address this, MASG achieves a balance between geometric consistency and material adaptability through structured graph learning and residual gating mechanisms, significantly improving the robustness and expressiveness of skinning weights. This module consists of two parts: Local Graph Encoding and Residual-Gated Refinement. Through dynamic modulation based on geometric deviations, it achieves differentiated weight optimization for rigid body surfaces and flexible clothing regions.
[0050] Local Graph Encoding: MASG constructs a nearest-neighbor-based local graph structure in the canonical space. Each node represents a Gaussian point, with edges connecting its geometric neighborhood. The input features of each node integrate geometric, appearance, material, and skinning attribute information:
[0051]
[0052] in, For the initial SMPL skinning weights, It is a three-plane position coding feature. It represents the deviation vector in the normal space (reflecting the residual displacement of the Gaussian point relative to the human body template, which can characterize the tightness of clothing and local thickness changes). The color projection features are used as material proxies. On this graph, MASG propagates features through an attention-enhanced Local Graph Neural Network (LocalGNNLayer). Given node features... and its neighborhood The update rules are as follows:
[0053]
[0054] in Activate GELU A lightweight Sigmoid attention mechanism based on feature concatenation is used to adaptively adjust the influence of neighboring nodes. To simultaneously capture local geometry and cross-part dependencies, MASG models geometric relationships in parallel across multi-scale neighborhoods. Features from different scales are integrated into a unified representation through a fusion function.
[0055]
[0056] in It consists of two layers of MLP. This structure effectively captures global consistency across clothing areas and body parts while maintaining local smoothness.
[0057] Residual-Gated Refinement: In the fusion of features Based on this, MASG predicts the skin weight correction for each node using a linear mapping:
[0058]
[0059] in To prevent training instability, we control the upper limit of the correction magnitude. To achieve region-adaptive weight adjustment, we introduce a residual gating function based on the norm deviation magnitude:
[0060]
[0061] in Control the gating response slope, This is the deviation threshold. When... (When fitting a rigid area) Suppress ineffective disturbances; when (In areas with flexible clothing) Sufficient adjustment is allowed. Weight adjustments are performed in logarithmic space to maintain numerical stability and a warm-up factor is introduced. Asymptotic convergence of control over the magnitude of the correction:
[0062]
[0063] This ensures the corrected weights Normalization constraints are satisfied. This gating fusion mechanism maintains physical consistency while avoiding drift risks caused by over-correction. Finally, the posetized position of the Gaussian point is calculated through a linear blending skinning process:
[0064]
[0065] in For the first The pose transformation matrix for each skeleton. MASG employs pre-computed nearest neighbor indices to reduce neighborhood search overhead, all linear layers use ReLU or GELU activation, and color features are encoded through lightweight linear projection layers. The module output includes corrected weights. Correction amount and gate value , used for joint optimization.
[0066] Loss Function: To ensure geometric stability while maintaining consistent modulation of dynamic appearance and structure, we employ a staged optimization strategy, enabling the model to gradually transition from a static clothing template to an animable Gaussian avatar. The entire training process comprises Stage 1 (static clothing avatar template modeling) and Stage 2 (graph learning avatar dynamic modulation), responsible for static geometric reconstruction and dynamic graph-driven learning, respectively. In the second stage, all components except the graph network parameters are frozen, and the graph learning module is gradually activated through a linear warm-up mechanism to ensure the stability and physical consistency of the optimization process.
[0067] The Static Clothed Avatar Template Modeling Stage aims to construct a high-fidelity canonical space clothing template and accurately fit the pose to provide a stable geometric and appearance basis. The optimization objective simultaneously considers photometric consistency, perceptual consistency, and canonical residual constraints to prevent excessive geometric divergence. The overall loss function is defined as:
[0068]
[0069] in, L1 norm is the difference in perceptual features between the rendered image and the ground truth image; The L2 norm is used to normalize the offset of the Gaussian center in space, for stabilizing the geometric topology; The mean value penalty for each point scale parameter is applied to suppress abnormal expansion / contraction; L2 regularization is applied to the learnable geometric feature maps to prevent unbounded growth of feature amplitudes. After this optimization stage, the model maintains a high-fidelity static template for clothing boundaries and body surface details.
[0070] The Graph-based Dynamic Modulation stage focuses on achieving dynamic adaptive adjustment of the animated avatar through graph learning mechanisms. In this stage, the network and pose parameters from stage one are frozen, and only the Multi-Level Pose Map (MLPG) and Material-Aware Skin Map (MASG) are optimized to model pose-driven appearance changes and material-sensitive skin weight adjustments, respectively. These two work together to achieve two-layer dynamic consistency between appearance and geometry. The loss function is defined as:
[0071]
[0072] The first two items continue photometric and perceptual supervision to maintain rendering consistency, while the latter three optimize pose representation, structural constraints, and skin correction, respectively: Overflow regularization. Suppressing attitude deviations exceeding the upper limit of the specification residuals and avoiding unreasonable deformation; skin weight regularization. Applying to MASG, it standardizes the weighting process for material perception:
[0073]
[0074] in, Local smoothing is constrained by the squared mean of the weight differences between adjacent edges in the adjacency graph. The L2 regularization of the logit is applied to the weights to constrain the magnitude of the correction and improve convergence stability; together, these two measures ensure that the correction of MASG satisfies geometric continuity and physical rationality.
[0075] Experimental Section
[0076] Experimental Datasets: To verify the effectiveness and stability of the proposed 3D Gaussian animable human avatar modeling method based on graph structure optimization under monocular video conditions, this invention selected two publicly available human video datasets for experimental evaluation. The People-Snapshot dataset contains video sequences of human bodies performing relatively restricted pose movements from a fixed camera perspective, typically showing the human body slowly rotating in front of the camera. This dataset is widely used for evaluating monocular video-driven animable human avatar modeling methods, and can verify the appearance consistency and pose-driven stability of the method under limited pose change conditions. In specific implementation, to maintain consistency with the evaluation settings of existing technical solutions, this invention selected four video sequences as experimental objects and used a unified data partitioning method for training and testing. The DynVideo dataset consists of human videos collected by mobile terminals, with a video duration of approximately one minute, covering diverse pose changes of the human body during natural movement, especially including relatively rich rotational movements. This dataset also provides human parameter sequences corresponding to video frames, which can be used to evaluate the dynamic performance of human geometry and appearance modeling. Compared to the People-Snapshot dataset, the DynVideo dataset contains richer human pose variations and places higher demands on the modeling of high-frequency appearance details such as clothing wrinkles and local shadows. It is suitable for verifying the pose generalization ability and detail fidelity of the present invention under complex dynamic conditions.
[0077] Evaluation indicators:
[0078] To quantitatively evaluate the reconstruction quality of animable 3D Gaussian human avatars, this invention employs three commonly used evaluation metrics: Peak Signal-to-Noise Ratio (PSNR), used to measure the pixel-level accuracy of the reconstructed image; Structural Similarity Index (SSIM), used to assess the structural similarity between the reconstructed result and the real image; and Perceptual Similarity Index (LPIPS), used to measure the visual perceptual similarity of the reconstructed result, better reflecting the consistency of high-frequency details and local appearance. These metrics can comprehensively evaluate the reconstruction effect of the method described in this invention from multiple dimensions, including pixel accuracy, structural consistency, and perceptual quality.
[0079] Experimental Setup: The specific implementation of this invention is based on the PyTorch deep learning framework, and all experiments are conducted on a single NVIDIA GeForce RTX 4090 graphics processing unit. During model training, the Adam optimizer is used to update the network parameters. Depending on the dataset size and video duration, the training time for a single 3D human avatar model ranges from approximately 0.5 hours to 6 hours, with a GPU memory usage of approximately 8GB during training. In terms of inference performance, the forward inference speed of this invention is no less than 67 FPS, which can meet the requirements of real-time animation generation and rendering.
[0080] Comparative Experiment: To verify the effectiveness of the proposed 3D Gaussian animable human avatar modeling method based on graph structure optimization, this invention was compared and evaluated with various existing monocular human avatar reconstruction methods on two publicly available human video datasets, People-Snapshot and DynVideo. Table 1 presents the quantitative evaluation results of different methods on the above datasets. Figure 2 The diagram illustrates the reconstruction results of the method described in this invention on the DynVideo dataset. Experimental results show that the method of this invention can maintain a stable appearance reconstruction effect under conditions of human posture change, verifying the effectiveness of the technical solution.
[0081] Table 1: Quantitative comparison of different methods on the People-Snapshot and DynVideo datasets. Higher PSNR or SSIM and lower LPIPS indicate better performance.
[0082]
[0083] Ablation Experiment: To verify the function of each technical module in this invention, an ablation experiment was conducted on the DynVideo dataset. Table 2 presents the quantitative evaluation results under different component introduction conditions. Figure 3 The corresponding reconstruction effect diagram is shown. Experimental results demonstrate that each technical module plays a positive role in human posture driving and dynamic appearance modeling.
[0084] Table 2: Ablation studies of different components on the DynVideo dataset
[0085] .
Claims
1. A method for modeling a 3D Gaussian-animable human avatar based on graph structure optimization, characterized in that, Includes the following steps: Step 1: Construction of static dressed 3D Gaussian human avatar template. In the normalized pose space, construct the human geometric prior based on the human parametric model, and generate 3D Gaussian primitives on its surface. Jointly optimize the geometric and appearance properties of the Gaussian primitives to obtain the static dressed 3D Gaussian human avatar template. Step 2: Dynamic modulation of appearance based on multi-level pose graph. A multi-level pose graph is constructed based on the human kinematic structure. The pose information of different parts of the human body is hierarchically modeled through the graph structure, and feature propagation and aggregation are performed in the graph to obtain part-level and global-level pose representations. The pose representations are mapped to the corresponding three-dimensional Gaussian elements. The geometric position, scale and appearance attributes of the Gaussian elements are dynamically modulated according to the pose, so that different body parts can obtain relatively independent and controllable appearance responses under pose change conditions. Step 3: Motion propagation optimization and pose generation based on material-aware skin map. A local neighborhood graph structure of 3D Gaussian primitives is constructed in the normed space. Geometric deviation information, spatial position features and material-related features extracted from the static template are integrated. The skin weights are structured and adaptively corrected through graph learning mechanism. By introducing material-aware gating modulation, the rigid human body region and the flexible clothing region exhibit differentiated motion responses during pose driving. Based on the optimized skin weights, the pose transformation and rendering of 3D Gaussian primitives are completed to generate an animated human avatar.
2. The method for modeling a 3D Gaussian animable human avatar based on graph structure optimization according to claim 1, characterized in that, In step 1, Gaussian primitives are generated by uniformly sampling from the surface of the SMPL model in the canonical pose based on UV mapping, and the Gaussian position offset, scale and color attributes in the canonical space are obtained by decoding through learnable feature maps and UNet encoding network.
3. The method for modeling a 3D Gaussian animated human avatar based on graph structure optimization according to claim 1, characterized in that, The multi-level pose graph constructed in step 2 includes fine-grained and coarse-grained layers. Node features are propagated and aggregated through graph neural networks, and a hierarchical contrastive learning mechanism is combined to enhance the semantic decoupling ability of pose features at different levels.
4. The method for modeling a 3D Gaussian animated human avatar based on graph structure optimization according to claim 1, characterized in that, In step 2, each Gaussian unit receives hierarchical pose features corresponding to its component, including fine-grained component features, mid-level component features, and global features, and predicts the dynamic increments of position, scale, and color brightness through a lightweight 1D convolutional decoder.
5. The method for modeling a 3D Gaussian animated human avatar based on graph structure optimization according to claim 1, characterized in that, The material-aware skin map constructed in step 3 uses Gaussian elements as nodes and geometric neighborhoods as edges. The input features include initial skin weights, three-plane position codes, normed deviation vectors, and material proxy features. Feature propagation and fusion are performed through attention-enhanced graph neural networks.
6. The method for modeling a 3D Gaussian animated human avatar based on graph structure optimization according to claim 5, characterized in that, In step 3, a residual gating function based on the norm deviation modulus is introduced to dynamically adjust the correction magnitude of the skin weight according to the degree of fit between the Gaussian point and the human body surface, thereby achieving hierarchical optimization of rigid and flexible regions.
7. The method for modeling a 3D Gaussian animated human avatar based on graph structure optimization according to claim 1, characterized in that, In step 3, the skin weight correction is performed in logarithmic space, and the effectiveness of the weight is ensured by Softmax normalization. Finally, Gaussian points in normal space are transformed to pose space by linear blending skin to complete the rendering.
8. The method for modeling a 3D Gaussian animable human avatar based on graph structure optimization according to claim 1, characterized in that, The method employs a phased optimization strategy: the first phase freezes the posture parameters and optimizes the geometry and appearance of the static clothing template; The second stage freezes the static template network, optimizes only the network parameters of the multi-level pose map and the material-aware skin map, and gradually activates the graph learning module through a linear warm-up mechanism.
9. The method for modeling a 3D Gaussian animated human avatar based on graph structure optimization according to claim 1, characterized in that, The loss function of the method includes photometric consistency loss, perceptual loss, contrastive learning loss, skin weight smoothing constraint and correction amplitude regularization term to ensure geometric continuity and appearance consistency under attitude-driven conditions.