A dual-stream collaborative reconstruction method and system based on template frame initialization
By using template frame initialization and patch binding to form a 3D Gaussian representation, combined with dual-stream collaborative deformation prediction and UV domain appearance refinement technology, the modeling problem of global contour and local wrinkles in dynamic human clothing reconstruction is solved. This achieves high fidelity and efficient optimization of clothing reconstruction, and improves the visual realism and physical rationality of the reconstruction results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAQIAO UNIVERSITY
- Filing Date
- 2026-05-27
- Publication Date
- 2026-06-26
AI Technical Summary
Existing dynamic human body clothing reconstruction technology lacks a layered modeling mechanism that matches the deformation frequency and structure of clothing, making it difficult to take into account both the global outline and local wrinkles. It also lacks 3D Gaussian cue features consistent with the clothing mesh topology, resulting in a disconnect between geometric reconstruction and rendering cueing. Furthermore, it lacks a mechanism that links clothing collision feedback to human body shape correction, making it difficult to suppress clothing deformities caused by human body estimation errors.
A dual-stream collaborative reconstruction method based on template frame initialization is adopted. By using template frame initialization and 3D Gaussian representation bound to surface patches, combined with dual-stream collaborative deformation prediction and UV domain appearance refinement technology, high-fidelity and efficient joint optimization of geometric deformation and appearance texture in clothing dynamic reconstruction is achieved.
It effectively preserves the fine texture and light and shadow details of the clothing, enhances the visual realism of the reconstruction results, ensures the stability of the overall movement of the clothing and the detailed expression of local folds, avoids the phenomenon of clothing penetrating the human body mesh, and ensures the physical rationality and wearability of the reconstruction process.
Smart Images

Figure CN122289507A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and specifically to a dual-stream collaborative reconstruction method and system based on template frame initialization. Background Technology
[0002] Dynamic human clothing reconstruction is a key foundational technology in digital human content generation, virtual try-on, film and television visual effects, and immersive interaction. Unlike rigid objects, real clothing, driven by human movement, typically exhibits two types of deformations that differ significantly in scale but are coupled with each other: one is low-frequency macroscopic deformation determined by human posture, gravity, and inertia, manifested as overall swaying, draping, folding, and large-scale contour changes; the other is high-frequency local deformation caused by material elasticity, friction, contact, and local stress concentration, manifested as fine-grained wrinkles, local stretching, and edge undulations. How to stably model both global coarse deformation and local detailed deformation within a unified framework is a core problem that urgently needs to be solved in the field of dynamic clothing reconstruction.
[0003] Among existing methods, some focus on temporal simulation of clothing based on explicit physical models. However, such methods often rely on material parameters, collision parameters, and complex solution processes, resulting in high computational costs and difficulty in directly utilizing the real appearance constraints in multi-view image observations. Other methods focus on directly regressing clothing deformation from images or point clouds, but they usually lack the ability to distinguish and model the deformation characteristics of clothing at different frequency bands, which can easily lead to problems such as unstable global contours, blurred local wrinkles, lining with the human body surface, and geometric inconsistencies across frames.
[0004] With the development of 3D Gaussian representation and differentiable rendering technology, researchers have attempted to use 3D Gaussian primitives to explicitly model dynamic objects in order to improve reconstruction quality and rendering efficiency. However, if 3D Gaussian is directly used as an independent representation in free space, its topological constraints on the clothing surface are not tight enough, making it prone to drift or unstable aggregation in local areas of the clothing, thus affecting the frame-by-frame geometric reconstruction accuracy. In addition, when relying solely on a unified spatial network to regress clothing deformation, it is easy to mix low-frequency overall contour changes with high-frequency local details in the model, leading to mutual interference between deformations at different scales.
[0005] Furthermore, human shape estimation in real-world videos often contains errors. If the estimated human body is too large, the clothing reconstruction result is prone to continuous compression due to collision constraints; if the estimated human body is too thin, unreasonable concavity or body-fitting artifacts are likely to appear on the clothing surface. Most existing methods only treat the human body as a static or fixed prior, and cannot dynamically correct the human body shape based on collision feedback during the clothing reconstruction process. Therefore, it is difficult to fundamentally solve the cascading effects of human body errors on clothing geometry.
[0006] In summary, existing dynamic human body clothing reconstruction technologies have at least the following problems: First, they lack a layered modeling mechanism that matches the deformation frequency structure of clothing, making it difficult to balance global contours and local wrinkles; second, they lack 3D Gaussian cueing features consistent with clothing mesh topology, leading to a disconnect between geometric reconstruction and rendering cueing; and third, they lack a mechanism that links clothing collision feedback to human body shape correction, making it difficult to suppress clothing deformities caused by human body estimation errors. Summary of the Invention
[0007] To address the aforementioned issues, this invention proposes a dual-stream collaborative reconstruction method and system based on template frame initialization. By initializing the template frame and binding the three-dimensional Gaussian representation with the surface patch, combined with dual-stream collaborative deformation prediction and UV domain appearance refinement technology, high-fidelity and efficient joint optimization of geometric deformation and appearance texture in clothing dynamic reconstruction is achieved.
[0008] On the one hand, a dual-stream collaborative reconstruction method based on template frame initialization includes:
[0009] S1. Obtain a set of template frame images from multiple perspectives. Use the camera intrinsic and extrinsic parameters and the multi-view matching results to recover the sparse 3D point set and segment the clothing area to obtain the template clothing mesh. Generate the template clothing point cloud and the corresponding UV unfolding result based on the template clothing mesh.
[0010] S2, based on the template clothing point cloud, the Gaussian parameters are initialized using the nearest neighbor inheritance strategy to obtain Gaussian primitives, and the Gaussian primitives are bound to the facets of the template clothing mesh to construct the facet-bound 3D Gaussian representation. Based on the facet-bound 3D Gaussian representation, the vertex-wise physical features and 3D Gaussian cue features are extracted to form the vertex-wise input feature matrix.
[0011] S3, input the feature matrix input vertex by vertex into the low-frequency branch of the spectral domain and the high-frequency branch of the spatial domain, predict the global deformation offset and local fine wrinkle displacement of the clothing in the current frame, obtain the total deformation displacement of the clothing in the current frame, and update the vertex coordinates of the reconstructed mesh of the clothing in the current frame based on the total deformation displacement of the clothing in the current frame.
[0012] S4, based on the updated vertex coordinates of the clothing reconstruction mesh in the current frame, maintain the learnable offsets on the original vertex coordinates of the human body mesh; based on rendering loss and collision loss, jointly optimize the learnable offsets to obtain the optimized human body and clothing reconstruction mesh;
[0013] S5: Based on the UV unwrapping results and the optimized human body and clothing reconstruction mesh, construct the UV domain Gaussian appearance representation; based on the UV domain Gaussian appearance representation, predict the Gaussian parameter correction amount through the appearance refinement network, and obtain the final clothing appearance reconstruction result of the current frame through differentiable rendering optimization.
[0014] Furthermore, in S2, 3D Gaussian cue features are extracted based on the 3D Gaussian representation of patch binding, and the calculation formula is as follows:
[0015] ;
[0016] ;
[0017] ;
[0018] ;
[0019] ;
[0020] ;
[0021] in, Indicates the first A Gaussian element-bound face index; Index representing a three-dimensional Gaussian; Indicates the number of Gaussian elements; Indices representing the coordinate components of the centroid; This indicates the offset along the direction normal to the bound patch; Indicates the nearest neighbor; This represents the opacity parameter; Indicates color and appearance characteristic parameters; Indicates rotation parameters; Indicates the scale parameter; , and These represent mapping functions for appearance, rotation, and scale parameters generated based on the local geometric properties of nearest neighbors, respectively. Indicates the first The centroid coordinates of a 3D Gaussian within the bound surface; , and This represents the centroid coordinates of the Gaussian center relative to the first, second, and third vertices of the bound face; Indicates the first The vertex correlation coefficients of a 3D Gaussian at the three vertices of the bound face; , and These represent the scale control parameters of the Gaussian element along the first tangential direction, the second tangential direction, and the normal direction, respectively. Represents vertices No. The correlation weights between each Gaussian element; This represents the feature mapping function used to encode face plates by binding Gaussian parameters. It is a stable term; This represents the distance attenuation parameter; Represents an exponential function; Represents the L2 norm; The position on the bound patch is represented by a combination of centroid coordinates and normal offset; Indicates the first A three-dimensional Gaussian hint feature.
[0022] Furthermore, in S3, the feature matrix input vertex-by-vertex is input into the low-frequency branch of the spectral domain to predict the global deformation offset of the clothing in the current frame, as follows:
[0023] Laplacian operator for template clothing mesh Eigenvalue decomposition is performed to obtain the Laplacian matrix, which is then divided into low-frequency and high-frequency Laplacian eigenvector matrices according to frequency. These are used as the feature matrices for vertex-by-vertex input, and the calculation formula is as follows:
[0024]
[0025]
[0026] in, This represents the Laplacian eigenvector matrix of the template clothing mesh; This represents the diagonal matrix of eigenvalues corresponding to the Laplacian eigenvector matrix; The Laplacian operator represents the template clothing mesh. Represents the real number field; This represents the low-frequency Laplacian eigenvector matrix. Indicates the low-frequency component; Represents the high-frequency Laplacian eigenvector matrix. Indicates the high-frequency portion; Indicates the number of vertices in the clothing; Indicates the number of low-frequency substrates;
[0027] The feature matrix, input vertex-by-vertex, is fed into the low-frequency branch of the spectral domain to predict the low-frequency coefficients of the current frame. The global deformation offset is then reconstructed using the low-frequency basis, as shown in the following formula:
[0028] ;
[0029] ;
[0030] in, This represents the mapping function corresponding to the low-frequency branch of the spectral domain. Indicates the current frame The low-spectral coefficient matrix; This indicates the global deformation offset of the clothing in the current frame; This represents the vertex-by-vertex input feature matrix of the clothing mesh in the template frame.
[0031] Furthermore, in S3, the vertex coordinates of the clothing reconstruction mesh in the current frame are updated based on the total deformation displacement of the clothing in the current frame. The calculation formula is as follows:
[0032] ;
[0033] ;
[0034] ;
[0035] in, This represents the mapping function corresponding to the high-frequency branch in the spatial domain; This indicates the local fine fold displacement of the clothing in the current frame; The confidence weights represent the local displacements; This indicates element-wise multiplication; This indicates the total deformation and displacement of the clothing in the current frame; This indicates the updated vertex coordinates of the clothing reconstruction mesh in the current frame; This represents the vertex coordinate matrix of the clothing mesh in the template frame.
[0036] Furthermore, in S4, the formula for calculating rendering loss is as follows:
[0037] ;
[0038] in, This represents the rendered image from the front viewpoint. Represents actual observed images; This represents the L1 reconstruction error between the rendered image and the real image; and These represent the weight coefficients of the L1 loss term and the SSIM structural similarity loss term, respectively; This represents the SSIM loss function.
[0039] Furthermore, in S4, collision loss The calculation formula is as follows:
[0040] ;
[0041] in, Indicates the first The projection distance of each clothing vertex onto the reference point on the human body surface in the direction of the normal; This represents the preset security boundary threshold; This indicates the number of vertices in the clothing mesh.
[0042] Furthermore, in S5, the final clothing appearance reconstruction result of the current frame is obtained through differentiable rendering optimization, and the calculation formula is as follows:
[0043] ;
[0044] ;
[0045] ;
[0046] in, Indicates the first The amount of local position correction for each Gaussian element in the current frame; Indicates the first A Gaussian unit in the current frame The initial anchoring position in the middle; This indicates the correction amount for the spherical harmonic appearance coefficient; Indicates the first The initial spherical harmonic appearance coefficients of each Gaussian element; This represents a differentiable rendering process; Indicates the first The final result of clothing appearance reconstruction in the frame; Indicates the number of Gaussian elements; Indicates the first The camera projection parameters corresponding to the current viewpoint of the frame; and These represent the updated Gaussian center position and Gaussian spherical harmonic appearance coefficients for the current frame, respectively.
[0047] On the other hand, a dual-stream collaborative reconstruction system based on template frame initialization includes:
[0048] The template initialization module is used to acquire a set of template frame images from multiple perspectives, recover a sparse 3D point set using camera intrinsic and extrinsic parameters and multi-view matching results, and segment the clothing region to obtain a template clothing mesh; based on the template clothing mesh, a template clothing point cloud and the corresponding UV unfolding results are generated.
[0049] The 3D Gaussian representation acquisition module is used to initialize Gaussian parameters based on the obtained template clothing point cloud using the nearest neighbor inheritance strategy to obtain Gaussian primitives, and establish binding relationships between the Gaussian primitives and the facets of the template clothing mesh to construct the facet-bound 3D Gaussian representation. Based on the facet-bound 3D Gaussian representation, the module extracts vertex-wise physical features and 3D Gaussian cue features to form the vertex-wise input feature matrix.
[0050] The grid vertex coordinate reconstruction module is used to input the feature matrix input vertex by vertex into the low-frequency branch of the spectral domain and the high-frequency branch of the spatial domain, predict the global deformation offset and local fine wrinkle displacement of the clothing in the current frame, obtain the total deformation displacement of the clothing in the current frame, and update the reconstructed grid vertex coordinates of the clothing in the current frame based on the total deformation displacement of the clothing in the current frame.
[0051] The joint optimization module is used to maintain the learnable offsets on the original vertex coordinates of the human body mesh based on the updated vertex coordinates of the clothing reconstruction mesh in the current frame; based on the rendering loss and collision loss, the learnable offsets are jointly optimized to obtain the optimized human body and clothing reconstruction mesh.
[0052] The rendering module is used to construct a UV-domain Gaussian appearance representation based on the UV unwrapping results and the optimized human body and clothing reconstruction mesh. Based on the UV-domain Gaussian appearance representation, the Gaussian parameter correction amount is predicted through the appearance refinement network, and the final clothing appearance reconstruction result of the current frame is obtained through differentiable rendering optimization.
[0053] The present invention adopts the above technical solution and has the following beneficial effects:
[0054] (1) By combining UV domain Gaussian appearance representation with differentiable rendering optimization, this invention effectively preserves the fine texture and light and shadow details of clothing, and enhances the visual realism of the reconstruction results.
[0055] (2) The present invention adopts dual-stream branch collaborative prediction in the spectral domain low frequency and spatial domain high frequency, which not only ensures the stability of the overall movement of the clothing, but also enhances the detail expression of local folds.
[0056] (3) By introducing human-clothing collision loss and learnable offset optimization, this invention effectively avoids the phenomenon of clothing and human body mesh penetration, ensuring the physical rationality and wearability of the reconstruction process. Attached Figure Description
[0057] Figure 1 This is a flowchart of the dual-stream collaborative reconstruction method based on template frame initialization according to an embodiment of the present invention;
[0058] Figure 2 This is a schematic diagram of the overall operation process of an embodiment of the present invention;
[0059] Figure 3 This is a flowchart illustrating the low-frequency branch of the spectral domain in an embodiment of the present invention.
[0060] Figure 4 This is a schematic diagram of the spatial high-frequency branching process according to an embodiment of the present invention;
[0061] Figure 5 This is a schematic diagram of the linkage optimization process according to an embodiment of the present invention;
[0062] Figure 6 This is a flowchart illustrating the appearance rendering stage of an embodiment of the present invention;
[0063] Figure 7 These are subjective comparison images of the rendering effects of embodiments of the present invention;
[0064] Figure 8 This is a diagram of a dual-stream collaborative reconstruction system based on template frame initialization, according to an embodiment of the present invention. Detailed Implementation
[0065] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0066] like Figure 1 As shown, the present invention provides a dual-stream collaborative reconstruction method based on template frame initialization, comprising:
[0067] S1. Obtain a set of template frame images from multiple perspectives. Use the camera's intrinsic and extrinsic parameters and the results of multi-view matching to recover the sparse 3D point set and segment the clothing area to obtain a template clothing mesh. Generate a template clothing point cloud and the corresponding UV unfolding results based on the template clothing mesh.
[0068] S2. Based on the obtained template clothing point cloud, the Gaussian parameters are initialized using the nearest neighbor inheritance strategy to obtain Gaussian primitives. The Gaussian primitives are then bound to the facets of the template clothing mesh to construct a facet-bound 3D Gaussian representation. Based on the facet-bound 3D Gaussian representation, vertex-wise physical features and 3D Gaussian cue features are extracted to form a vertex-wise input feature matrix.
[0069] like Figure 2 As shown, the overall workflow of this invention is illustrated. Starting with a template frame, the process begins by obtaining a template clothing mesh and its UV unwrapping results through multi-view reconstruction and clothing region extraction. Subsequently, a 3D Gaussian representation bound to facets is constructed around the template clothing mesh, and vertex-by-vertex input features are extracted. During the frame-by-frame reconstruction stage, these input features are fed into the low-frequency branch in the spectral domain and the high-frequency branch in the spatial domain to model the global low-frequency deformation and local high-frequency wrinkles of the clothing, respectively. After obtaining the geometric reconstruction results, geometric consistency and physical rationality are further constrained through human-clothing linkage optimization. Finally, based on the optimized geometric results, a Gaussian appearance representation is constructed in the UV domain, and the target frame clothing appearance reconstruction result is output by combining an appearance refinement network and differentiable rendering. This overall workflow embodies a collaborative reconstruction process that starts from template geometric priors and gradually completes geometric and appearance restoration.
[0070] Specifically, the 3D Gaussian cue features are extracted based on the 3D Gaussian representation of patch binding, and the calculation formula is as follows:
[0071] ;
[0072] ;
[0073] ;
[0074] ;
[0075] ;
[0076] ;
[0077] in, Indicates the first A Gaussian element-bound face index; Index representing a three-dimensional Gaussian; Indicates the number of Gaussian elements; Indices representing the coordinate components of the centroid; This indicates the offset along the direction normal to the bound patch; Indicates the nearest neighbor; This represents the opacity parameter; Indicates color and appearance characteristic parameters; Indicates rotation parameters; Indicates the scale parameter; , and These represent mapping functions for appearance, rotation, and scale parameters generated based on the local geometric properties of nearest neighbors, respectively. Indicates the first The centroid coordinates of a 3D Gaussian within the bound surface; , and This represents the centroid coordinates of the Gaussian center relative to the first, second, and third vertices of the bound face; Indicates the first The vertex correlation coefficients of a 3D Gaussian at the three vertices of the bound face; , and These represent the scale control parameters of the Gaussian element along the first tangential direction, the second tangential direction, and the normal direction, respectively. Represents vertices No. The correlation weights between each Gaussian element; This represents the feature mapping function used to encode face plates by binding Gaussian parameters. It is a stable term; This represents the distance attenuation parameter; Represents an exponential function; Represents the L2 norm; The position on the bound patch is represented by a combination of centroid coordinates and normal offset; Indicates the first A three-dimensional Gaussian hint feature.
[0078] Specifically, in this embodiment, a template clothing dot cloud is used. Initialize the 3D Gaussian metaset as a geometric reference. , among which, the Each Gaussian element is represented as:
[0079] ;
[0080] in, Indicates the first The center coordinates of each Gaussian element; Represent the covariance matrix; This represents the opacity parameter; Indicates color or appearance characteristic parameters; Indicates rotation parameters; Indicates the scale parameter; The template frame represents the first layer on the clothing surface. The location of each sampling point or anchor point; This indicates the total number of sampling points or anchor points on the surface of the clothing in the template frame.
[0081] For each Gaussian center to be initialized In template clothing dot cloud Search for its nearest neighbor in the middle to obtain the nearest neighbor index:
[0082] ;
[0083] Based on the nearest neighbor This completes the parameter inheritance initialization of the Gaussian meta-element, specifically as follows:
[0084] ;
[0085] ;
[0086] ;
[0087] in, , , , and These represent mapping functions that generate Gaussian covariance, opacity, appearance, rotation, and scale parameters based on the local geometric properties of nearest neighbors.
[0088] After initializing the Gaussian elements, each Gaussian element is bound to a triangular face in the template clothing mesh. For the first Gaussian element... Each high-level element has a bound patch index. The minimum distance from the center of Gauss to each triangular facet is determined as follows:
[0089] ;
[0090] in, Indicates the middle of Each facet corresponds to a triangular facet region formed by the three vertices; This represents the shortest Euclidean distance from a point to a triangular facet.
[0091] Specifically, in this embodiment, let the first... The triangular facet bound to each Gaussian element The three vertices are respectively , and The position of the Gaussian center on the bound surface is then expressed using the combined barycentric coordinates and normal offset as follows:
[0092] ;
[0093] And satisfy:
[0094] ;
[0095] in, , , This represents the centroid coordinates of the Gaussian center relative to the three vertices of the bound face; This indicates the offset along the direction normal to the bound patch; Indicates binding facet The unit normal vector.
[0096] Specifically, in this embodiment, to describe the anisotropic distribution of Gaussian elements in the local space of the bound surface, its covariance matrix is parameterized using a local coordinate system. Let the bound surface... Corresponding local orthogonal basis for:
[0097] ;
[0098] Then the first The covariance matrix of Gaussian elements is expressed as:
[0099] ;
[0100] in, , , These represent the scale control parameters of the Gaussian element along the two tangential and normal directions, respectively; Represents a local orthogonal basis The transpose of .
[0101] Therefore, the 3D Gaussian representation of patch binding is constructed as follows:
[0102] ;
[0103] in, ; .
[0104] The patch-bound 3D Gaussian representation encodes the attachment relationship of Gaussian elements to the template clothing mesh by binding patch indices, centroid coordinates, normal offsets, and local anisotropic scale parameters, thus enabling Gaussian elements to maintain stable binding with local geometric changes in the template clothing mesh. After constructing the patch-bound 3D Gaussian representation, per-vertex 3D Gaussian cue features are further extracted based on this representation. For the th vertex in the template clothing mesh... vertices .
[0105] S3 inputs the feature matrix per vertex into the low-frequency branch of the spectral domain and the high-frequency branch of the spatial domain to predict the global deformation offset and local fine wrinkle displacement of the clothing in the current frame, and obtains the total deformation displacement of the clothing in the current frame. Based on the total deformation displacement of the clothing in the current frame, the vertex coordinates of the reconstructed mesh of the clothing in the current frame are updated.
[0106] Specifically, the feature matrix input vertex-by-vertex is input into the low-frequency branch of the spectral domain to predict the global deformation offset of the clothing in the current frame, as follows:
[0107] like Figure 3 The diagram illustrates the process of using the low-frequency branch of the spectral domain to model the low-frequency overall deformation of clothing. First, the Laplacian operator is applied to the template clothing mesh. Eigenvalue decomposition is performed to obtain the Laplacian matrix, which is then divided into low-frequency and high-frequency Laplacian eigenvector matrices based on frequency. These are used as the feature matrices for vertex-by-vertex input, and the calculation formula is as follows:
[0108] ;
[0109] ;
[0110] in, This represents the Laplacian eigenvector matrix of the template clothing mesh; This represents the diagonal matrix of eigenvalues corresponding to the Laplacian eigenvector matrix; The Laplacian operator represents the template clothing mesh. Represents the real number field; This represents the low-frequency Laplacian eigenvector matrix. Indicates the low-frequency component; Represents the high-frequency Laplacian eigenvector matrix. Indicates the high-frequency portion; Indicates the number of vertices in the clothing; Indicates the number of low-frequency substrates;
[0111] The feature matrix, input vertex-by-vertex, is fed into the low-frequency branch of the spectral domain to predict the low-frequency coefficients of the current frame. The global deformation offset is then reconstructed using the low-frequency basis, as shown in the following formula:
[0112] ;
[0113] ;
[0114] in, This represents the mapping function corresponding to the low-frequency branch of the spectral domain. Indicates the current frame The low-spectral coefficient matrix; This indicates the global deformation offset of the clothing in the current frame; This represents the vertex-by-vertex input feature matrix of the clothing mesh in the template frame.
[0115] like Figure 4 As shown, the spatial high-frequency branch is used to model the local fine wrinkles and high-frequency detail displacements on the garment surface. This branch directly predicts the local fine wrinkle displacements of the garment in the current frame in the spatial domain based on vertex-by-vertex input features, and further outputs the corresponding local displacement confidence weights. Subsequently, the local fine wrinkle displacements and confidence weights are fused element-wise to highlight reliable local high-frequency details and suppress unstable prediction regions. Finally, the output of the spatial high-frequency branch is combined with the output of the spectral low-frequency branch to obtain the total deformation displacement of the garment in the current frame, thus balancing global contour reconstruction capabilities with local detail representation capabilities.
[0116] Specifically, the vertex coordinates of the clothing reconstruction mesh in the current frame are updated based on the total deformation displacement of the clothing in the current frame. The calculation formula is as follows:
[0117] ;
[0118] ;
[0119] ;
[0120] in, This represents the mapping function corresponding to the high-frequency branch in the spatial domain; This indicates the local fine fold displacement of the clothing in the current frame; The confidence weights represent the local displacements; This indicates element-wise multiplication; This indicates the total deformation and displacement of the clothing in the current frame; This indicates the updated vertex coordinates of the clothing reconstruction mesh in the current frame; This represents the vertex coordinate matrix of the clothing mesh in the template frame.
[0121] S4: Based on the updated vertex coordinates of the clothing reconstruction mesh in the current frame, maintain the learnable offsets on the original vertex coordinates of the human body mesh; based on rendering loss and collision loss, jointly optimize the learnable offsets to obtain the optimized human body and clothing reconstruction mesh.
[0122] like Figure 5 As shown, this invention introduces a human-clothing linkage optimization mechanism in the geometric reconstruction stage. After obtaining the clothing reconstruction mesh of the current frame, the learnable offsets of the original vertex coordinates of the human body mesh and the offsets of the clothing vertices are maintained simultaneously, and a joint optimization objective is constructed based on rendering loss and collision loss. The rendering loss is used to constrain the consistency between the optimized geometric results and the actual observed image, while the collision loss is used to constrain the physical contact relationship between the clothing mesh and the human body mesh, avoiding unreasonable penetration. The gradients of the two types of losses are jointly backpropagated to the geometric parameters of the human body and clothing, enabling the human body geometric adjustment and clothing deformation compensation to be updated in a linked manner under a unified objective. This alleviates the problems of clothing compression, body-fitting artifacts, and local clipping caused by human body estimation errors, improving the geometric rationality and dynamic stability of the reconstruction results.
[0123] Specifically, the formula for calculating rendering loss is as follows:
[0124] ;
[0125] in, This represents the rendered image from the front viewpoint. Represents actual observed images; This represents the L1 reconstruction error between the rendered image and the real image; and These represent the weight coefficients of the L1 loss term and the SSIM structural similarity loss term, respectively; This represents the SSIM loss function.
[0126] Specifically, collision damage The calculation formula is as follows:
[0127] ;
[0128] in, Indicates the first The projection distance of each clothing vertex onto the reference point on the human body surface in the direction of the normal; This represents the preset security boundary threshold; This indicates the number of vertices in the clothing mesh.
[0129] Specifically, in this embodiment, during the process of jointly optimizing the learnable offset of human body vertices and the offset of clothing vertices to obtain the optimized human body and clothing reconstruction mesh for each frame:
[0130] For the Frame, let the human body template mesh vertex coordinate matrix Clothing template mesh vertex coordinate matrix The learnable offset of the human vertex is Clothing vertex offset The optimized human body reconstruction mesh and clothing reconstruction mesh are represented as follows:
[0131] ;
[0132] ;
[0133] Furthermore, a joint optimization objective is constructed based on rendering loss and collision loss, and the learnable offsets of human body vertices and clothing vertices are updated synchronously, as follows:
[0134] ;
[0135] in, Indicates the first Frame rendering loss Indicates the first Frame collision loss, This represents the collision loss weighting coefficient. and These represent the learnable offsets of the human body vertices and the clothing vertices after joint optimization, respectively.
[0136] Therefore, the first The vertex coordinates of the reconstructed human body and clothing meshes after frame optimization are as follows:
[0137]
[0138]
[0139] in, Indicates the first Frame-optimized vertex coordinate matrix of the human body reconstruction mesh. Indicates the first The frame-optimized vertex coordinate matrix of the clothing reconstruction mesh. By jointly optimizing the learnable offsets of human body vertices and clothing vertex offsets, human body geometric adjustment and clothing deformation compensation are collaboratively updated under the same objective function constraint, thereby improving the geometric consistency between the human body and clothing reconstruction results. Furthermore, ambient light characteristics... Local normal features and perspective features As input, the Gaussian local position correction and spherical harmonic appearance coefficient correction of the current frame are predicted by the appearance refinement network, and are expressed as follows:
[0140] ;
[0141] in, This indicates a detailed network pattern. Indicates the first The amount of local position correction for each Gaussian element in the current frame. This represents the correction amount for the spherical harmonic appearance coefficient. By superimposing the above two correction amounts onto the initial local position and initial spherical harmonic appearance coefficient of each Gaussian element, the updated Gaussian center position and spherical harmonic appearance coefficient under the current frame and current viewpoint can be obtained, thus completing the frame-by-frame viewpoint adaptive refinement of the Gaussian appearance parameters in the UV domain.
[0142] S5, based on the UV unwrapping results and the optimized human and clothing reconstruction mesh, constructs a UV-domain Gaussian appearance representation; based on the UV-domain Gaussian appearance representation, predicts Gaussian parameter corrections through an appearance refinement network, and obtains the final clothing appearance reconstruction result for the current frame through differentiable rendering optimization. Specifically, the entire process first superimposes the local position correction and spherical harmonic appearance coefficient correction onto the initial parameters of each Gaussian element to obtain the updated Gaussian center position and spherical harmonic appearance coefficients; then, the updated complete Gaussian parameters are fed into a differentiable renderer, and combined with the current frame's camera projection parameters to render the final clothing appearance reconstruction result for the current frame.
[0143] like Figure 6 As shown, the appearance rendering stage constructs a Gaussian appearance representation in the UV domain based on the UV unwrapping results of the template clothing mesh and the optimized human and clothing reconstruction mesh of the current frame. For effective texture sampling points in the UV domain, the binding relationship between them and clothing mesh patches is first established, and their anchor point positions on the 3D clothing surface of the current frame are obtained through barycentric coordinate mapping. Subsequently, using ambient light features, local normal features, and viewpoint features as inputs, the appearance refinement network predicts the Gaussian local position correction and spherical harmonic appearance coefficient correction under the current viewpoint. Finally, the updated Gaussian parameters are fed into the differentiable renderer to obtain the clothing rendering result under the current viewpoint. This stage can further restore the texture, lighting, and local appearance details of the clothing based on geometric reconstruction, thereby obtaining a higher fidelity visual reconstruction effect.
[0144] Specifically, firstly, the local position correction and spherical harmonic appearance coefficient correction are superimposed on the initial parameters of each Gaussian element to obtain the updated Gaussian center position and spherical harmonic appearance coefficients; then, the updated complete Gaussian parameters are fed into the differentiable renderer, and combined with the current frame's camera projection parameters to render the final clothing appearance reconstruction result for the current frame. The calculation formula is as follows:
[0145] ;
[0146] ;
[0147] ;
[0148] in, Indicates the first The amount of local position correction for each Gaussian element in the current frame; Indicates the first A Gaussian unit in the current frame The initial anchoring position in the middle; This indicates the correction amount for the spherical harmonic appearance coefficient; Indicates the first A Gaussian unit in the current frame The initial anchoring position in the middle; Indicates the first The initial spherical harmonic appearance coefficients of each Gaussian element; This represents a differentiable rendering process; Indicates the first The final result of clothing appearance reconstruction in the frame; Indicates the number of Gaussian elements; Indicates the first The camera projection parameters corresponding to the current viewpoint of the frame; and These represent the updated Gaussian center position and Gaussian spherical harmonic appearance coefficients for the current frame, respectively. By superimposing these two correction values onto the initial local position and initial spherical harmonic appearance coefficients of each Gaussian element, the updated Gaussian center position and spherical harmonic appearance coefficients for the current frame and current viewpoint can be obtained, thus completing the frame-by-frame viewpoint adaptive refinement of the Gaussian appearance parameters in the UV domain.
[0149] Specifically, in this embodiment, the Gaussian parameter correction amount is predicted through an appearance refinement network, and the final clothing appearance reconstruction result of the current frame is obtained through differentiable rendering optimization. The calculation formula is as follows:
[0150] First, the UV unwrapping results based on the template clothing mesh. With UV sheet Within the UV domain, a patch binding relationship is established for each valid texture sampling point. Let the first... The binding patch corresponding to each valid UV sampling point is Its centroid coefficient at the three vertices of the triangular facet The coordinates of the anchor point mapped from the effective UV sampling point to the three-dimensional clothing surface are expressed as:
[0151]
[0152]
[0153] in, , and They represent the first Binding facets in frame clothing reconstruction mesh The coordinates of the three vertices, Indicates the first The coordinates of the three-dimensional anchor point corresponding to each valid UV sampling point in the current frame.
[0154] Specifically, in this embodiment, during the linkage optimization process, rendering loss and collision loss The gradient will be simultaneously propagated back to the learnable offset at the human vertex. and clothing vertex offset This updates the parameters of the clothing deformation prediction network and the human body vertex offset, so that the human body vertex offset and clothing vertex offset are no longer optimized independently, but are updated in a coordinated manner under the combined effect of rendering consistency constraints and collision constraints. This effectively alleviates clothing deformity, local compression, and clipping problems caused by human body shape estimation errors, thereby improving the accuracy and stability of dynamic clothing reconstruction results. In the appearance rendering stage, based on the UV unfolding results of the current frame's clothing reconstruction mesh and the template clothing mesh, a UV domain Gaussian appearance representation is constructed. Combining ambient light features, normal features, and viewpoint features, the Gaussian local position correction amount and the Gaussian spherical harmonic appearance coefficient correction amount are predicted. The clothing appearance reconstruction result of the current frame is obtained through differentiable rendering optimization, including the following steps:
[0155] ;
[0156] in, This represents the first value obtained based on barycentric coordinate interpolation. The three-dimensional coordinates of each sampling point or anchor point; , , This represents the weight of the centroid coordinates of the sampling point relative to the three vertices of the triangle it belongs to; , , Indicates the first The coordinates of the three vertices of a triangular facet.
[0157] A further appearance refinement network is constructed to jointly correct the surface appearance and local geometric details of the garment. In this embodiment, the appearance refinement network adopts the StyleUNet architecture, which is an image domain neural network that introduces a style modulation mechanism on the basis of the standard U-Net encoding and decoding structure. It is used to perform high-quality joint correction of the surface appearance and local geometric details of the garment in the UV texture domain.
[0158] Specifically, the network uses ambient light feature maps, normal feature maps, and viewpoint feature maps unfolded in the UV domain as multi-channel inputs. The encoder extracts multi-scale spatial features by downsampling layer by layer. Simultaneously, the global style vector of the current frame is encoded by the viewpoint direction and illumination conditions. This vector is then injected into each scale feature layer through adaptive instance normalization (AdaIN) to achieve global style modulation of viewpoint-related appearance changes. The decoder fuses the encoder features of the corresponding scale through skip connections and upsamples layer by layer to restore the UV image resolution, thus preserving local details while maintaining global illumination consistency. Finally, StyleUNet outputs the local position correction and spherical harmonic appearance coefficient correction for each Gaussian element in the UV domain through position correction output heads and appearance coefficient correction output heads, respectively. These are then added to the initial position and initial spherical harmonic coefficients of the corresponding Gaussian element to complete the frame-by-frame viewpoint adaptive refinement of the UV domain Gaussian appearance parameters.
[0159] Specifically, using ambient light features, normal features, and viewpoint features as inputs, the appearance refinement network outputs the Gaussian local position correction amount under the current viewpoint. Correction amount for Gaussian spherical harmonic appearance coefficient Therefore, the updated local position of Gauss from the current perspective. Gaussian spherical harmonic appearance coefficient They are represented as follows:
[0160] ;
[0161] ;
[0162] Updated Gaussian parameters The image is fed into a differentiable renderer to obtain the rendered image from the current viewpoint. :
[0163] ;
[0164] in, This represents the camera projection parameters corresponding to the current viewpoint. Therefore, by minimizing the reconstruction error between the rendered image and the real multi-view image, and applying Gaussian position, scale, and opacity regularization terms, the UV domain Gaussian appearance parameters are optimized end-to-end, ultimately yielding a high-fidelity clothing appearance reconstruction result for the current frame.
[0165] Finally, an appearance optimization objective is constructed based on the reconstruction error, structural similarity error, and Gaussian position, scale, and opacity regularization terms estimated from the original multi-view images.
[0166] Specifically, this embodiment was validated on an NVIDIA 4090 GPU server, and a dynamic clothing reconstruction and differentiable rendering training process was built based on a deep learning framework. In the template frame generation stage, 3D reconstruction, clothing region extraction, and template mesh restoration were performed on multi-view template frame images to obtain the template clothing mesh, point cloud, and corresponding UV unwrapping results. In the frame-by-frame geometric reconstruction stage, the template clothing mesh was used as the initial geometric prior. A 3D Gaussian representation of patch binding was constructed for each frame of the target video sequence, and vertex-by-vertex physical features and 3D Gaussian cue features were extracted. The spectral domain low-frequency branch was used to predict the global deformation offset of the clothing in the current frame, and the spatial domain high-frequency branch was used to predict the local fine wrinkles of the clothing in the current frame. Based on rendering loss and collision loss, the human vertex offset and clothing vertex offset were optimized in a linked manner to obtain the clothing reconstruction mesh for each frame. In the appearance rendering stage, based on the UV unfolding results of the clothing reconstruction mesh and the template clothing mesh in the current frame, a Gaussian appearance representation in the UV domain is constructed. Combining ambient light features, normal features, and viewpoint features, the Gaussian local position correction and Gaussian spherical harmonic appearance coefficient correction are predicted. The clothing appearance reconstruction result for the current frame is then obtained through differentiable rendering optimization. This invention can improve the stability, realism, and practicality of dynamic clothing reconstruction results under complex pose changes and multi-view observation conditions while maintaining high reconstruction accuracy.
[0167] The proposed method was comprehensively compared with the representative state-of-the-art algorithm (Gaussian-Garments) in the field of dynamic human costume reconstruction on the Actor01 sequence of the ActorsHQ dataset. Through multiple comparative experiments, the comprehensive advantages of the proposed method were verified in terms of both geometric reconstruction accuracy and multi-view rendering quality. Specifically, the geometric reconstruction metric was obtained by averaging the results from eight camera views, and the rendering quality metric was obtained by averaging the results from three camera views. Comparative experiments between the method of this invention and Gaussian-Garments on the Actor01 sequence of the ActorsHQ dataset show that, in terms of geometric reconstruction accuracy, the average IoU (0.7839) and average F1 (0.8747) of the method of this invention are superior to those of Gaussian-Garments (0.7697, 0.8657), while the average Chamfer Distance (11.4636 px) is lower than that of Gaussian-Garments (12.5906 px). In terms of multi-view rendering quality, the average L1 (0.001302) and average LPIPS (0.016004) of the method of this invention are lower than those of Gaussian-Garments (0.001519, 0.017141), while the average PSNR (41.7815) and average SSIM (0.9909) are superior to those of Gaussian-Garments (40.9853, 0.9904). In summary, the method of this invention outperforms Gaussian-Garments in both geometric reconstruction accuracy and multi-view rendering quality.
[0168] like Figure 7 As shown in the figure, the rendering effects of the method of this invention and the comparative method in a dynamic clothing reconstruction task are subjectively compared. It can be seen that, compared to the comparative method, the method of this invention performs better in preserving the overall outline of the clothing, restoring local wrinkle details, and maintaining rendering consistency across multiple perspectives. Especially in complex pose changes, edge areas, and fine-grained wrinkle areas, it presents a clearer, more stable, and realistically observed appearance. This subjective comparison further verifies the comprehensive advantages of this invention in both geometric reconstruction and appearance restoration.
[0169] like Figure 8 As shown, this embodiment also discloses a dual-stream collaborative reconstruction system based on template frame initialization, including:
[0170] The template initialization module 81 is used to acquire a set of template frame images from multiple perspectives, recover a sparse 3D point set using camera intrinsic and extrinsic parameters and multi-view matching results, and segment the clothing region to obtain a template clothing mesh; and generate a template clothing point cloud and corresponding UV unfolding results based on the template clothing mesh.
[0171] The 3D Gaussian representation acquisition module 82 is used to initialize Gaussian parameters based on the obtained template clothing point cloud using the nearest neighbor inheritance strategy to obtain Gaussian primitives, and to establish a binding relationship between the Gaussian primitives and the facets of the template clothing mesh to construct a facet-bound 3D Gaussian representation. Based on the facet-bound 3D Gaussian representation, the physical features and 3D Gaussian cue features per vertex are extracted to form a feature matrix for per vertex input.
[0172] The mesh vertex coordinate reconstruction module 83 is used to input the feature matrix input vertex by vertex into the low-frequency branch of the spectral domain and the high-frequency branch of the spatial domain, predict the global deformation offset and local fine wrinkle displacement of the clothing in the current frame, obtain the total deformation displacement of the clothing in the current frame, and update the reconstructed mesh vertex coordinates of the clothing in the current frame based on the total deformation displacement of the clothing in the current frame.
[0173] The joint optimization module 84 is used to maintain the learnable offset on the original vertex coordinates of the human body mesh based on the updated vertex coordinates of the clothing reconstruction mesh in the current frame; based on the rendering loss and collision loss, the learnable offset is jointly optimized to obtain the optimized human body and clothing reconstruction mesh.
[0174] Rendering module 85 is used to construct a UV domain Gaussian appearance representation based on the UV unwrapping results and the optimized human body and clothing reconstruction mesh; based on the UV domain Gaussian appearance representation, the Gaussian parameter correction amount is predicted through the appearance refinement network, and the final clothing appearance reconstruction result of the current frame is obtained through differentiable rendering optimization.
[0175] A specific implementation of a dual-stream collaborative reconstruction system based on template frame initialization is described in this embodiment, which is the same dual-stream collaborative reconstruction method based on template frame initialization.
[0176] Although the invention has been specifically shown and described in conjunction with preferred embodiments, those skilled in the art should understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined in the appended claims, all of which shall be within the scope of protection of the invention.
Claims
1. A dual-stream collaborative reconstruction method based on template frame initialization, characterized in that, Includes the following steps: S1. Obtain a set of template frame images from multiple perspectives. Use the camera intrinsic and extrinsic parameters and the multi-view matching results to recover the sparse 3D point set and segment the clothing area to obtain the template clothing mesh. Generate the template clothing point cloud and the corresponding UV unfolding result based on the template clothing mesh. S2, based on the template clothing point cloud, the Gaussian parameters are initialized using the nearest neighbor inheritance strategy to obtain Gaussian primitives, and the Gaussian primitives are bound to the facets of the template clothing mesh to construct the facet-bound 3D Gaussian representation. Based on the facet-bound 3D Gaussian representation, the vertex-wise physical features and 3D Gaussian cue features are extracted to form the vertex-wise input feature matrix. S3, input the feature matrix input vertex by vertex into the low-frequency branch of the spectral domain and the high-frequency branch of the spatial domain, predict the global deformation offset and local fine wrinkle displacement of the clothing in the current frame, obtain the total deformation displacement of the clothing in the current frame, and update the vertex coordinates of the reconstructed mesh of the clothing in the current frame based on the total deformation displacement of the clothing in the current frame. S4, based on the updated vertex coordinates of the clothing reconstruction mesh in the current frame, maintain the learnable offsets on the original vertex coordinates of the human body mesh; based on rendering loss and collision loss, jointly optimize the learnable offsets to obtain the optimized human body and clothing reconstruction mesh; S5: Based on the UV unwrapping results and the optimized human body and clothing reconstruction mesh, construct the UV domain Gaussian appearance representation; based on the UV domain Gaussian appearance representation, predict the Gaussian parameter correction amount through the appearance refinement network, and obtain the final clothing appearance reconstruction result of the current frame through differentiable rendering optimization.
2. The dual-stream collaborative reconstruction method based on template frame initialization according to claim 1, characterized in that, In S2, 3D Gaussian cue features are extracted based on the 3D Gaussian representation of patch binding, and the calculation formula is as follows: ; ; ; ; ; ; in, Indicates the first A Gaussian element-bound face index; Index representing a three-dimensional Gaussian; Indicates the number of Gaussian elements; Indices representing the coordinate components of the centroid; This indicates the offset along the direction normal to the bound patch; Indicates the nearest neighbor; This represents the opacity parameter; Indicates color and appearance characteristic parameters; Indicates rotation parameters; Indicates the scale parameter; , and These represent mapping functions for appearance, rotation, and scale parameters generated based on the local geometric properties of nearest neighbors, respectively. Indicates the first The centroid coordinates of a 3D Gaussian within the bound surface; , and This represents the centroid coordinates of the Gaussian center relative to the first, second, and third vertices of the bound face; Indicates the first The vertex correlation coefficients of a 3D Gaussian at the three vertices of the bound face; , and These represent the scale control parameters of the Gaussian element along the first tangential direction, the second tangential direction, and the normal direction, respectively. Represents vertices No. The correlation weights between each Gaussian element; This represents the feature mapping function used to encode face plates by binding Gaussian parameters. It is a stable term; This represents the distance attenuation parameter; Represents an exponential function; Represents the L2 norm; The position on the bound patch is represented by a combination of centroid coordinates and normal offset; Indicates the first A three-dimensional Gaussian hint feature.
3. The dual-stream collaborative reconstruction method based on template frame initialization according to claim 1, characterized in that, In S3, the feature matrix input vertex by vertex is input into the low-frequency branch of the spectral domain to predict the global deformation offset of the clothing in the current frame, as follows: Laplacian operator for template clothing mesh Eigenvalue decomposition is performed to obtain the Laplacian matrix, which is then divided into low-frequency and high-frequency Laplacian eigenvector matrices according to frequency. These are used as the feature matrices for vertex-by-vertex input, and the calculation formula is as follows: ; ; in, This represents the Laplacian eigenvector matrix of the template clothing mesh; This represents the diagonal matrix of eigenvalues corresponding to the Laplacian eigenvector matrix; The Laplacian operator represents the template clothing mesh. Represents the real number field; This represents the low-frequency Laplacian eigenvector matrix. Indicates the low-frequency component; Represents the high-frequency Laplacian eigenvector matrix. Indicates the high-frequency portion; Indicates the number of vertices in the clothing; Indicates the number of low-frequency substrates; The feature matrix, input vertex-by-vertex, is fed into the low-frequency branch of the spectral domain to predict the low-frequency coefficients of the current frame. The global deformation offset is then reconstructed using the low-frequency basis, as shown in the following formula: ; ; in, This represents the mapping function corresponding to the low-frequency branch of the spectral domain. Indicates the current frame The low-spectral coefficient matrix; This indicates the global deformation offset of the clothing in the current frame; This represents the vertex-by-vertex input feature matrix of the clothing mesh in the template frame.
4. The dual-stream collaborative reconstruction method based on template frame initialization according to claim 3, characterized in that, In S3, the vertex coordinates of the clothing reconstruction mesh in the current frame are updated based on the total deformation displacement of the clothing in the current frame. The calculation formula is as follows: ; ; ; in, This represents the mapping function corresponding to the high-frequency branch in the spatial domain; This indicates the local fine fold displacement of the clothing in the current frame; The confidence weights represent the local displacements; This indicates element-wise multiplication; This indicates the total deformation and displacement of the clothing in the current frame; This indicates the updated vertex coordinates of the clothing reconstruction mesh in the current frame; This represents the vertex coordinate matrix of the clothing mesh in the template frame.
5. The dual-stream collaborative reconstruction method based on template frame initialization according to claim 1, characterized in that, In S4, the formula for calculating rendering loss is as follows: ; in, This represents the rendered image from the front viewpoint. Represents actual observed images; This represents the L1 reconstruction error between the rendered image and the real image; and These represent the weight coefficients of the L1 loss term and the SSIM structural similarity loss term, respectively; This represents the SSIM loss function.
6. The dual-stream collaborative reconstruction method based on template frame initialization according to claim 1, characterized in that, In S4, collision loss The calculation formula is as follows: ; in, Indicates the first The projection distance of each clothing vertex onto the reference point on the human body surface in the direction of the normal; This represents the preset security boundary threshold; This indicates the number of vertices in the clothing mesh.
7. The dual-stream collaborative reconstruction method based on template frame initialization according to claim 1, characterized in that, In S5, the final clothing appearance reconstruction result of the current frame is obtained through differentiable rendering optimization. The calculation formula is as follows: ; ; ; in, Indicates the first The amount of local position correction for each Gaussian element in the current frame; Indicates the first A Gaussian unit in the current frame The initial anchoring position in the middle; This indicates the correction amount for the spherical harmonic appearance coefficient; Indicates the first The initial spherical harmonic appearance coefficients of each Gaussian element; This represents a differentiable rendering process; Indicates the first The final result of clothing appearance reconstruction in the frame; Indicates the number of Gaussian elements; Indicates the first The camera projection parameters corresponding to the current viewpoint of the frame; and These represent the updated Gaussian center position and Gaussian spherical harmonic appearance coefficients for the current frame, respectively.
8. A dual-stream collaborative reconstruction system based on template frame initialization, characterized in that, include: The template initialization module is used to acquire a set of template frame images from multiple perspectives, recover a sparse 3D point set using camera intrinsic and extrinsic parameters and multi-view matching results, and segment the clothing region to obtain a template clothing mesh; based on the template clothing mesh, a template clothing point cloud and the corresponding UV unfolding results are generated. The 3D Gaussian representation acquisition module is used to initialize Gaussian parameters based on the obtained template clothing point cloud using the nearest neighbor inheritance strategy to obtain Gaussian primitives, and establish binding relationships between the Gaussian primitives and the facets of the template clothing mesh to construct the facet-bound 3D Gaussian representation. Based on the facet-bound 3D Gaussian representation, the module extracts vertex-wise physical features and 3D Gaussian cue features to form the vertex-wise input feature matrix. The grid vertex coordinate reconstruction module is used to input the feature matrix input vertex by vertex into the low-frequency branch of the spectral domain and the high-frequency branch of the spatial domain, predict the global deformation offset and local fine wrinkle displacement of the clothing in the current frame, obtain the total deformation displacement of the clothing in the current frame, and update the reconstructed grid vertex coordinates of the clothing in the current frame based on the total deformation displacement of the clothing in the current frame. The joint optimization module is used to maintain learnable offsets on the original vertex coordinates of the human body mesh based on the updated current frame clothing reconstruction mesh vertex coordinates; Based on rendering loss and collision loss, the learnable offset is jointly optimized to obtain the optimized human body and clothing reconstruction mesh. The rendering module is used to construct a UV-domain Gaussian appearance representation based on the UV unwrapping results and the optimized human body and clothing reconstruction mesh. Based on the UV-domain Gaussian appearance representation, the Gaussian parameter correction amount is predicted through the appearance refinement network, and the final clothing appearance reconstruction result of the current frame is obtained through differentiable rendering optimization.