Virtual fitting video generation method and system based on multi-view face fixation
Through the combination of SAM, CLIP, diffusion repair and WAN model, the inconsistency problem of multi-view angle and dynamic texture processing in virtual fitting video generation is solved, and a high-quality and personalized virtual fitting experience is achieved.
Patent Information
- Application Number
- CN202511097921.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-06
AI Technical Summary
The existing virtual fitting video generation methods have inconsistencies, identity features blurred, failed to reconstruct occlusion area and texture mapping jagging effects in multi-view and dynamic texture processing, and cannot be deployed in real time and lack multi-view geometric consistency constraints.
The SAM model is used to segment the face and clothing areas, and the multi-view image is generated by combining rigid geometric transformation. The CLIP model is used to analyze the action prompt words to generate the target pose sequence. The texture is optimized through the diffusion repair model, the WAN model is used for smooth interframe transitions, and texture repair and pose alignment are combined with ACE-LoRA and Vista2Vogue technology.
It achieves natural interaction between clothing and characters in multiple perspectives and complex postures, maintains realistic fit, consistency of texture details, and smooth transition between frames, improving the fluidity of user experience and personalized virtual fitting effects.
Smart Images

Figure CN120602604A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of generative artificial intelligence and computer vision, and more specifically, to a method and system for generating virtual fitting videos based on multi-view facial fixation. Background Art
[0002] With the rapid development of virtual fitting technology, especially in e-commerce, fashion, virtual reality and other fields, the generation of virtual fitting videos has become an important part of user experience. Traditional virtual fitting technology is mostly based on the processing of static images, by taking photos of products or people, and then using computer graphics technology to generate fitting effects.
[0003] Explicit geometric constraints are a commonly used technique in computer graphics, computer vision, and machine learning. They are used to enforce certain geometric conditions so that the generated models or images are more consistent with the geometric rules of the real world. Unlike implicit constraints, explicit geometric constraints are directly specified in the algorithm design through clear mathematical relationships, thereby enhancing the accuracy and interpretability of the model. This makes explicit geometric constraints particularly important in tasks such as virtual fitting, image generation, and 3D modeling, especially when ensuring the consistency and naturalness between the generated images and the actual objects or scenes.
[0004] However, existing virtual fitting video generation methods and systems based on explicit geometric constraints are incoherent during use. In addition, virtual fitting methods rely on 3D modeling or preset posture templates and cannot flexibly generate facial views from different angles. The traditional Fill-Redux module has not been extended to the video field and cannot process dynamic textures. At the same time, existing virtual fitting video generation methods and systems rely on single-view image generation and cannot directly expand multi-view content through text prompts. Methods that rely on 3D modeling or physical simulation require a large amount of resources and are difficult to deploy in real time. In addition, identity consistency, geometric coherence, and high-frequency detail fidelity across perspectives remain challenges. At the same time, current mainstream solutions are all based on single-view image generation and lack a constraint framework for multi-view geometric consistency. When attempting to expand the perspective through text prompts, problems such as blurred identity features and failed reconstruction of occluded areas often occur. In addition, there is a significant jagged effect in the texture mapping of dynamically deformed areas. As a result, existing methods still face significant challenges in key indicators such as cross-view geometric coherence and high-frequency detail fidelity.
[0005] Currently, no effective solutions have been proposed for the problems in related technologies. Summary of the Invention
[0006] In response to the problems in the related art, the present invention proposes a virtual fitting video generation method and system based on multi-view facial fixation to overcome the above-mentioned technical problems existing in the existing related art.
[0007] In order to achieve the above object, the specific technical solutions adopted by the present invention are as follows: According to one aspect of the present invention, a method and system for generating a virtual fitting video based on multi-view facial fixation is provided, comprising the following steps: S1. Obtain original fitting images and action prompt words, perform image normalization and text vectorization processing respectively, and construct a multi-source input dataset; S2. Use the SAM model to segment the original fitting images in the multi-source input dataset, obtain the facial area and clothing area in the image, and extract the facial mask of the facial area. Then, based on the rigid geometric transformation algorithm, perform multi-view modeling on the facial mask to generate multi-angle facial images, and extract the clothing texture feature parameters of the clothing area. As a preferred solution, the method uses the SAM model to segment the original fitting image in the multi-source input data set, obtains the facial area and clothing area in the image, extracts the facial mask of the facial area, and then performs multi-view modeling on the facial mask based on the rigid geometric transformation algorithm to generate multi-angle facial images, and extracts clothing texture feature parameters of the clothing area, including the following steps: S21. Use the SAM model to perform semantic segmentation on the original fitting images in the multi-source input dataset to obtain the facial region and clothing region in the image, and extract the facial mask of the facial region; S22, performing similarity transformation on the facial mask using a rigid geometric transformation algorithm to generate a head transformation map, and outputting the map as a multi-angle facial image; S23. Extract clothing texture feature parameters of the clothing area, and unify the clothing texture parameters of multi-angle facial images through a material-aware style transfer algorithm.
[0008] S3, based on the CLIP model, parses action prompt words, generates target pose sequences, and constructs a dual-view spatial topology prior framework to fuse multi-view facial images with clothing texture feature parameters to generate the first and last frame virtual fitting images; As a preferred solution, the CLIP model is used to parse action prompt words, generate target posture sequences, construct a dual-view space topology prior framework, fuse multi-view facial images with clothing texture feature parameters, and generate first and last frame virtual fitting images, including the following steps: S31, extracting the semantic feature vector of the action prompt word through the CLIP text encoder, and generating a continuous posture key point coordinate sequence by combining it with the predefined posture dictionary; S32: The original fitting image and the geometrically transformed facial image are spliced into a left and right image, and the spatial attention mechanism is used to achieve feature alignment, and the SRIM mechanism is used to maintain identity consistency; As a preferred solution, the process of splicing the original fitting image and the geometrically transformed facial image into a left and right dual image, using a spatial attention mechanism to achieve feature alignment, and maintaining identity consistency through a SRIM mechanism includes the following steps: S321, using the global features of the original image as identity anchors, aligning the identity key points in the dual images through a feature matching algorithm, and splicing the original fitting image and the geometrically transformed facial image into a left and right dual image; S322, adopting spatial attention mechanism and introducing context-aware mask to keep the texture of clothing area continuous and aligned with the original image; S323. By optimizing the spatial alignment error of the dual-view spatial topology prior framework, the interpolation distortion caused by posture transformation is suppressed, and identity consistency is maintained through the SRIM mechanism.
[0009] S33, generate high-fidelity images of the clothing area through the diffusion restoration model, and use a high-frequency texture preservation strategy to optimize details. Then, perform pixel-level fusion with the transformed facial area to obtain the first and last frames of the virtual fitting image; As a preferred solution, the method of generating a high-fidelity image of the clothing area by using a diffusion restoration model, optimizing details by adopting a high-frequency texture preservation strategy, and then performing pixel-level fusion with the transformed facial area to obtain the first and last frame virtual fitting images includes the following steps: S331, performing posture alignment on the reference image according to the target posture, and performing physical rationality calibration on the generated image after posture alignment to ensure the coupling relationship between the clothing and the human body posture; As a preferred solution, the method of performing posture alignment on the reference image according to the target posture, and performing physical rationality calibration on the generated image after posture alignment to ensure the coupling relationship between the clothing and the human body posture includes the following steps: Assume the target posture is , the pose of the reference image is , the posture alignment process can be expressed as: ; in, Represents a geometric transformation function, which is implemented by transformations such as rotation and scaling; The generated images are calibrated for physical plausibility to ensure the coupling relationship between clothing and human body posture. This process can be expressed by the following formula: ; in, is the physical plausibility loss function, and represent the mechanical constraint and deformation constraint of the i-th pixel respectively.
[0010] The S332 and Flux models guide the generation process by combining textual cues, and the Flux diffusion model generates clear images by gradually guiding noisy images; S333: When repairing details and local textures in clothing areas, Vista2Vogue uses a dedicated high-frequency texture preservation strategy to repair high-frequency textures. It also uses the Fill-Redux module to repair high-frequency details in non-head areas, and maintains the consistency of details when the pose changes. As a preferred solution, when repairing the details and local textures of the clothing area, Vista2Vogue uses a special high-frequency texture preservation strategy to repair the high-frequency texture, and uses the Fill-Redux module to repair the high-frequency details of the non-head area, and maintains the consistency of the details when the posture changes, including the following steps: The S3331 and Fill modules use image restoration algorithms to fill in missing details and combine contextual information to ensure the naturalness of the restoration effect. The high-frequency detail restoration process is optimized using the following loss function: ; in, is the high frequency detail recovery loss, Represents the repaired image of the current pixel p, is the true value of the target image; The S3332 and Redux modules use style transfer technology to uniformly optimize the textures of the head and non-head areas. The goal is to minimize texture distortion during posture transformation. The optimization is performed using the following formula: ; in, Optimize the results for textures, represents the texture after posture transformation, For reference texture.
[0011] S334 and ACE-LoRA use a low-rank adaptive optimization strategy to update some parameters of the model and adjust the image content during the generation process based on the alignment information; S335 and Vista2Vogue input the static generation results into the WAN model and achieve smooth transition between frames through spatiotemporal consistency constraints.
[0012] S4, performing pairing analysis on the first and last frames of the virtual fitting image, calculating the change vector between adjacent postures of the first and last frames of the virtual fitting image, and generating an inter-frame motion vector field, injecting the motion vector field as input into the diffusion model to generate posture transition parameters; S5. Input the first and last frame virtual fitting images and posture transition parameters into the WAN video generation model, perform dynamic texture restoration on the image, and generate a continuous transition frame image sequence based on the first and last frames; As a preferred solution, the process of inputting the first and last frame virtual fitting images and posture transition parameters into the WAN video generation model, performing dynamic texture restoration on the images, and generating a continuous transition frame image sequence based on the first and last frames includes the following steps: S51. Load the first and last frame images according to the FLF2V model, quantize them through FP16 and set the frame number. At the same time, use the T5 text encoder to parse the posture transition parameters to obtain the motion guidance vector. The motion guidance vector is used to assist in controlling the dynamic strategy of generating inter-frame transitions, but is not used as an explicit input to the WAN model. After receiving the RGB first and last frame images output by Vista2Vogue, the S52 and WAN models rely on their built-in DiT spatiotemporal attention mechanism to achieve dynamic alignment and natural transition of clothing textures. At the same time, the WAN model implicitly models the coupling relationship between clothing and human posture changes through its VCU unit, and calculates the continuous change values of cloth wrinkles and motion amplitude; S53, finely generate transition frames through its built-in frame interpolation module, and output a continuous and natural transition frame image sequence. S54. Output the transition sequence verified by physical constraints, trigger regeneration for frames that fail collision detection, and obtain a continuous transition frame image sequence.
[0013] S6. Generate a virtual fitting video based on the transition frame image sequence, and verify the output of the virtual fitting video.
[0014] According to another aspect of the present invention, a multi-view virtual fitting video generation system based on facial rigid transformation is provided, the system comprising: The data acquisition and processing module obtains the original fitting images and action prompt words, performs image normalization and text vectorization processing respectively, and constructs a multi-source input dataset; The image segmentation processing module uses the SAM model to segment the original fitting images in the multi-source input data set, obtain the facial area and clothing area in the image, and extract the facial mask of the facial area. Then, based on the rigid geometric transformation algorithm, the facial mask is modeled from multiple perspectives to generate multi-angle facial images, and the clothing texture feature parameters of the clothing area are extracted. The virtual fitting image module parses action prompt words based on the CLIP model, generates target pose sequences, and constructs a dual-view spatial topology prior framework to fuse multi-view facial images with clothing texture feature parameters to generate the first and last frames of virtual fitting images. The image posture transition module performs pairing analysis on the first and last frames of the virtual fitting image, calculates the change vector between adjacent postures of the first and last frames of the virtual fitting image, and generates an inter-frame motion vector field. The motion vector field is injected into the diffusion model as input to generate posture transition parameters; The transition image sequence module inputs the first and last frame virtual fitting images and posture transition parameters into the WAN video generation model, performs dynamic texture restoration on the image, and generates a continuous transition frame image sequence based on the first and last frames; The video verification output module generates a virtual fitting video based on a transition frame image sequence and verifies the output of the virtual fitting video.
[0015] The beneficial effects of the present invention are: 1. This invention combines the image segmentation of the SAM model with the rigid geometric transformation algorithm to achieve accurate segmentation and multi-angle modeling of facial and clothing areas, making the interaction between clothing and characters more natural. In particular, it can maintain a realistic fit between face and clothing under multiple perspective changes and complex postures. The Flux diffusion model is used to guide the generation process, gradually generating clear images. At the same time, Vista2Vogue is used to repair high-frequency textures to ensure the consistency of texture details of clothing under different posture changes. In addition, the Fill-Redux module is used to successfully repair high-frequency details in non-head areas during posture changes, avoiding common texture distortion problems.
[0016] 2. The present invention ensures the natural coupling between clothing and human body posture changes by utilizing physical rationality calibration, avoids deformation anomalies by introducing deformation energy parameters and cloth dynamics models, and enhances the physical consistency of the generated image, making it more consistent with the actual wearing effect in different postures. At the same time, through the constraints of spatiotemporal consistency, it ensures the smooth transition between frames in the virtual fitting video, so that the character remains consistent in the transition of multiple postures, avoiding the common inter-frame jumps or abruptness, and improving the smoothness and comfort of the user experience. In combination with the clothing texture feature fusion of the CLIP model, it can generate highly customized virtual fitting videos according to the action prompts and clothing requirements of different users. Through the fusion of these technologies, high-quality simulation effects are achieved in visual effects, and a more personalized virtual fitting experience is provided for users.
[0017] 3. The present invention uses the Vista2Vogue method to generate the first and last frames and reconstruct the multi-view generation process using explicit geometric constraints. Its core idea is to use the rigid transformation of the head to construct a dual-view stitching reference image, then use the spatial topology prior to lock the identity anchor and establish cross-view geometric associations. At the same time, texture restoration of non-head areas is achieved through pixel-level loss and style transfer constraints, and a low-rank adaptive strategy is used to efficiently align the semantics of posture text. Vista2Vogue outperforms advanced solutions in identity consistency, clothing structure restoration, and multi-view coherence, while maintaining real-time efficiency, providing a deployable solution for high-fidelity virtual fitting. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 is a flowchart of a method for generating a multi-view virtual fitting video based on facial rigid transformation according to an embodiment of the present invention; Figure 2 4 is a system block diagram of a virtual fitting video generation system based on multi-view face fixation according to an embodiment of the present invention.
[0020] In the picture: 1. Data acquisition and processing module; 2. Image segmentation and processing module; 3. Virtual fitting image module; 4. Image posture transition module; 5. Transition image sequence module; 6. Video verification output module. DETAILED DESCRIPTION
[0021] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0022] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.
[0023] According to an embodiment of the present invention, a method and system for generating a virtual fitting video based on multi-view facial fixation are provided.
[0024] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. According to one embodiment of the present invention, Figure 1 As shown, the method for generating a virtual fitting video based on explicit geometric constraints according to an embodiment of the present invention includes the following steps: S1. Obtain original fitting images and action prompt words, perform image normalization and text vectorization processing respectively, and construct a multi-source input dataset; S2. Use the SAM model to segment the original fitting images in the multi-source input dataset, obtain the facial area and clothing area in the image, and extract the facial mask of the facial area. Then, based on the rigid geometric transformation algorithm, perform multi-view modeling on the facial mask to generate multi-angle facial images, and extract the clothing texture feature parameters of the clothing area. As a preferred solution, the method uses the SAM model to segment the original fitting image in the multi-source input data set, obtains the facial area and clothing area in the image, extracts the facial mask of the facial area, and then performs multi-view modeling on the facial mask based on the rigid geometric transformation algorithm to generate multi-angle facial images, and extracts clothing texture feature parameters of the clothing area, including the following steps: S21. Use the SAM model to perform semantic segmentation on the original fitting images in the multi-source input dataset to obtain the facial region and clothing region in the image, and extract the facial mask of the facial region; S22, performing similarity transformation on the facial mask using a rigid geometric transformation algorithm to generate a head transformation map, and outputting the map as a multi-angle facial image; Specifically, by rotating the array and scaling factor Adjust the head posture and proportion, the formula is: ; where the translation vector t is driven by the keypoint coordinates of the target pose.
[0025] S23. Extract clothing texture feature parameters of the clothing area, and unify the clothing texture parameters of multi-angle facial images through a material-aware style transfer algorithm.
[0026] Specifically, this step is cyclic. First, the facial mask is extracted, then the facial mask is rigidly transformed, scaled and rotated, and then a virtual fitting image with a new perspective is generated. Then the same operation is performed again to generate a virtual fitting image with a new perspective. Finally, the corresponding number is generated according to the needs of use, and multi-perspective modeling is constructed based on the generated virtual fitting images.
[0027] S3, based on the CLIP model, parses action prompt words, generates target pose sequences, and constructs a dual-view spatial topology prior framework to fuse multi-view facial images with clothing texture feature parameters to generate the first and last frame virtual fitting images; As a preferred solution, the CLIP model is used to parse action prompt words, generate target posture sequences, construct a dual-view space topology prior framework, fuse multi-view facial images with clothing texture feature parameters, and generate first and last frame virtual fitting images, including the following steps: S31, extracting the semantic feature vector of the action prompt word through the CLIP text encoder, and generating a continuous posture key point coordinate sequence by combining it with the predefined posture dictionary; S32: The original fitting image and the geometrically transformed facial image are spliced into a left and right image, and the spatial attention mechanism is used to achieve feature alignment, and the SRIM mechanism is used to maintain identity consistency; As a preferred solution, the process of splicing the original fitting image and the geometrically transformed facial image into a left and right dual image, using a spatial attention mechanism to achieve feature alignment, and maintaining identity consistency through a SRIM mechanism includes the following steps: S321, using the global features of the original image as identity anchors, aligning the identity key points in the dual images through a feature matching algorithm, and splicing the original fitting image and the geometrically transformed facial image into a left and right dual image; S322, adopting spatial attention mechanism and introducing context-aware mask to keep the texture of clothing area continuous and aligned with the original image; Specifically, the pixel-level difference of the background area is constrained by calculating the L1 loss function: ; Among them, M non-head is the binary mask outside the head area.
[0028] S323. By optimizing the spatial alignment error of the dual-view spatial topology prior framework, the interpolation distortion caused by posture transformation is suppressed, and identity consistency is maintained through the SRIM mechanism.
[0029] S33, generate high-fidelity images of the clothing area through the diffusion restoration model, and use a high-frequency texture preservation strategy to optimize details. Then, perform pixel-level fusion with the transformed facial area to obtain the first and last frames of the virtual fitting image; As a preferred solution, the method of generating a high-fidelity image of the clothing area by using a diffusion restoration model, optimizing details by adopting a high-frequency texture preservation strategy, and then performing pixel-level fusion with the transformed facial area to obtain the first and last frame virtual fitting images includes the following steps: S331, performing posture alignment on the reference image according to the target posture, and performing physical rationality calibration on the generated image after posture alignment to ensure the coupling relationship between the clothing and the human body posture; As a preferred solution, the method of performing posture alignment on the reference image according to the target posture, and performing physical rationality calibration on the generated image after posture alignment to ensure the coupling relationship between the clothing and the human body posture includes the following steps: Assume the target posture is , the pose of the reference image is , the posture alignment process can be expressed as: ; in, Represents a geometric transformation function, which is implemented by transformations such as rotation and scaling; The generated images are calibrated for physical plausibility to ensure the coupling relationship between clothing and human body posture. This process can be expressed by the following formula: ; in, is the physical plausibility loss function, and They represent the mechanical constraint and deformation constraint of the i-th pixel respectively, and by minimizing this loss function, the physical rationality of the image is corrected.
[0030] The S332 and Flux models guide the generation process by combining textual cues, and the Flux diffusion model generates clear images by gradually guiding noisy images; Specifically, by combining with textual prompts, the Flux model can generate details that are more consistent with expectations. For example, given a description of a model making a victory sign with their left hand, the Flux model will automatically adjust the hand posture and texture of the relevant area based on this information to ensure consistency of details across different viewpoints. The Flux diffusion model generates a clear image by gradually guiding the noisy image. Its generation process can be expressed as the following equation: ; in, represents the current image in the diffusion process, is the noise amplitude, For the image given the text description y The gradient information of the image is obtained, and this process gradually optimizes the image to make it more consistent with the text description and the target pose.
[0031] S333: When repairing details and local textures in clothing areas, Vista2Vogue uses a dedicated high-frequency texture preservation strategy to repair high-frequency textures. It also uses the Fill-Redux module to repair high-frequency details in non-head areas, and maintains the consistency of details when the pose changes. As a preferred solution, when repairing the details and local textures of the clothing area, Vista2Vogue uses a special high-frequency texture preservation strategy to repair the high-frequency texture, and uses the Fill-Redux module to repair the high-frequency details of the non-head area, and maintains the consistency of the details when the posture changes, including the following steps: The S3331 and Fill modules use image restoration algorithms to fill in missing details and combine contextual information to ensure the naturalness of the restoration effect. The high-frequency detail restoration process is optimized using the following loss function: ; in, is the high frequency detail recovery loss, Represents the repaired image of the current pixel p, is the true value of the target image; Specifically, the Fill module focuses on repairing high-frequency details in non-head areas, such as clothing texture, wrinkles, and fabric sheen.
[0032] The S3332 and Redux modules use style transfer technology to uniformly optimize the textures of the head and non-head areas. The goal is to minimize texture distortion during posture transformation. The optimization is performed using the following formula: ; in, Optimize the results for textures, represents the texture after posture transformation, is the reference texture, and the loss function ensures texture coherence and detail consistency during pose transformation.
[0033] S334 and ACE-LoRA use a low-rank adaptive optimization strategy to update some parameters of the model and adjust the image content during the generation process based on the alignment information; Specifically, ACE-LoRA uses a low-rank adaptive optimization strategy to update only some parameters of the model, thereby achieving efficient alignment with fewer computing resources. The optimization goal of ACE-LoRA can be expressed as: ; in, is the optimized weight matrix, is the reference weight, and Ϝ represents the Frobenius norm.
[0034] The ACE-LoRA technology can quickly align the input posture and text description, and adjust the image content in the generation process based on this alignment information. The alignment process can be expressed as the following formula: ; in, and The loss function minimizes the difference between the generated image and the target pose for the generated pose and the reference pose, respectively.
[0035] S335 and Vista2Vogue input the static generation results into the WAN model and achieve smooth transition between frames through spatiotemporal consistency constraints.
[0036] As a preferred solution, Vista2Vogue inputs the static generation results into the WAN model and implements smooth transition between frames through spatiotemporal consistency constraints, including the following steps: In the process of generating the first and last frames of video in Alibaba WAN models (such as Wan2.1-FLF2V-14B), the core relies on the following key technologies to achieve smooth transitions between frames: Conditional control branch: First and last frames as control conditions: The model uses the first and last frames provided by the user as input conditions and compresses them into conditional latent representations through 3DVAE. This process implicitly extracts semantic features of the first and last frames (such as object position, style, and scene), eliminating the need for manual modeling of motion trajectories.
[0037] Binary Mask: By introducing a binary mask (1 for retained frames, 0 for frames to be generated), the model explicitly knows which frames to keep (first and last frames) and which to generate dynamically. This avoids explicit calculation of complex motion fields.
[0038] Semantic feature injection: CLIP Encoder and Global Context: The model extracts semantic features from the first and last frames through the CLIP image encoder and converts them into global context through the MLP. This semantic information is injected into the cross-attention mechanism of the DiT model, ensuring that the generation process always revolves around the semantic content of the first and last frames, thereby achieving consistency in style, content, and structure.
[0039] Dynamic constraints: The injection of semantic features implicitly constrains the dynamic change direction of the generated frames (for example, the object position and action trend in the first and last frames), avoiding the problem of explicitly estimating the motion trajectory in traditional methods.
[0040] Spatiotemporal modeling of latent spaces: Causal 3DVAE: The WAN model uses a causal 3DVAE (spatiotemporal variational autoencoder). The encoder compresses the video into a latent space, while the decoder gradually reconstructs the spatiotemporal information. Because the VAE's latent space implicitly captures spatiotemporal continuity, the model achieves smooth transitions between frames without explicitly computing dense fields.
[0041] Diffusion-based Time-Dependent Interaction (DiT) spatiotemporal modeling: The DiT architecture directly learns the spatiotemporal relationship between the first and last frames through self-attention and cross-attention mechanisms in the Transformer block. The diffusion process gradually removes noise and generates intermediate frames, implicitly addressing the problem of inter-frame motion coherence.
[0042] Data-driven implicit learning: Large-scale training data: The WAN model is trained on a high-quality dataset consisting of 1.5 billion videos and 10 billion images. During training, the model learns a rich set of spatiotemporal dynamic patterns (such as object motion patterns and light and shadow variations). Therefore, inference requires no additional modeling of motion trajectories, relying instead on data-driven implicit learning.
[0043] Four-step data cleaning process: Strict training data screening ensures that the model can learn natural inter-frame transition patterns from the data, rather than relying on manually designed motion modeling.
[0044] Efficiency optimization and hardware adaptation: High-Compression 3DVAE: WAN 2.2's 3DVAE achieves temporal and spatial compression ratios up to 4×16×16, significantly reducing video memory usage (only 22GB of video memory is needed to generate a 5-second 720P video). This efficient compression allows the model to run on consumer-grade hardware without explicitly processing the high-dimensional motion field.
[0045] MoE architecture and parallel strategy: The mixture of experts (MoE) architecture and 2D context parallel strategy (RingAttention+Ulysses) further optimize computational efficiency and avoid the high computational cost of traditional motion modeling.
[0046] S4, performing pairing analysis on the first and last frames of the virtual fitting image, calculating the change vector between adjacent postures of the first and last frames of the virtual fitting image, and generating an inter-frame motion vector field, injecting the motion vector field as input into the diffusion model to generate posture transition parameters; Specifically, the posture changes of adjacent frames are calculated to generate the inter-frame motion vector field. Combined with the ACE-LoRA technology, the motion vector field is injected into the diffusion model to guide the generation of inter-frame posture transitions.
[0047] Then by optimizing the L1 loss function, the formula is: ; And correct the physical rationality of clothing and human body between frames (such as fabric tension and joint movement).
[0048] S5. Input the first and last frame virtual fitting images and posture transition parameters into the WAN video generation model, perform dynamic texture restoration on the image, and generate a continuous transition frame image sequence based on the first and last frames; As a preferred solution, the process of inputting the first and last frame virtual fitting images and posture transition parameters into the WAN video generation model, performing dynamic texture restoration on the images, and generating a continuous transition frame image sequence based on the first and last frames includes the following steps: S51. Load the first and last frame images according to the FLF2V model, quantize them through FP16 and set the frame number. At the same time, use the T5 text encoder to parse the posture transition parameters to obtain the motion guidance vector. The motion guidance vector is used to assist in controlling the dynamic strategy of generating inter-frame transitions, but is not used as an explicit input to the WAN model. Specifically, the process of parsing posture transition parameters can also input prompt words to describe how the first and last frames change, such as a model in a white dress making a "V" gesture, then smiling and moving away from the camera, and sitting on a chair, and the action enhancement words are output with parameter changes.
[0049] After receiving the RGB first and last frame images output by Vista2Vogue, the S52 and WAN models rely on their built-in DiT spatiotemporal attention mechanism to achieve dynamic alignment and natural transition of clothing textures. At the same time, the WAN model implicitly models the coupling relationship between clothing and human posture changes through its VCU unit, and calculates the continuous change values of cloth wrinkles and motion amplitude; Specifically, a dynamic texture restoration algorithm is applied to non-head areas (such as clothing bodies), combined with the inter-frame motion vector field, to repair details such as fabric fluttering and wrinkle changes. The loss formula for high-frequency details restoration through time perception is:
[0050] Optimize repair effects.
[0051] S53 , finely generating transition frames through its built-in frame interpolation module, and outputting a continuous and natural transition frame image sequence.
[0052] Specifically, we use style transfer technology to uniformly optimize the textures of the head and clothing areas to ensure style consistency between frames. The formula is: ; S54. Output the transition sequence verified by physical constraints, trigger regeneration for frames that fail collision detection, and obtain a continuous transition frame image sequence.
[0053] S6. Generate a virtual fitting video based on the transition frame image sequence, and verify the output of the virtual fitting video.
[0054] Specifically, only the low-rank parameters related to posture and motion in the model are updated to reduce the amount of calculation. The formula is: ; Efficient video generation is achieved by processing time series through recurrent neural networks or Transformer architectures.
[0055] The output video contains a complete video sequence of the first frame (front view), the last frame (back / side view), and the intermediate transition frames. The video quality is evaluated using indicators such as KID, SSIM, and LPIPS, and the optical flow consistency loss formula is used:
[0056] Quantify inter-frame coherence.
[0057] The WAN model has a built-in quality assurance mechanism. The VCU unit of Wan2.1-VACE has been implemented during the generation process: Inter-frame consistency check (spatial-temporal attention weight monitoring); Physical plausibility verification (implicit constraints on cloth dynamics); Identity drift suppression (CLIP embedding space similarity calculation); Adding an additional verification module is equivalent to a secondary check of the WAN output.
[0058] According to another aspect of the present invention, Figure 2 A multi-view virtual fitting video generation system based on facial rigid transformation is provided, the system comprising: Data acquisition and processing module 1 obtains the original fitting images and action prompt words, performs image normalization and text vectorization processing respectively, and constructs a multi-source input dataset; Image segmentation processing module 2 uses the SAM model to segment the original fitting images in the multi-source input dataset, obtain the facial area and clothing area in the image, and extract the facial mask of the facial area. Then, based on the rigid geometric transformation algorithm, the facial mask is modeled from multiple perspectives to generate multi-angle facial images, and clothing texture feature parameters of the clothing area are extracted. Virtual fitting image module 3: Based on the CLIP model, it parses action prompt words, generates target pose sequences, and constructs a dual-view spatial topology prior framework to fuse multi-view facial images with clothing texture feature parameters to generate the first and last frames of virtual fitting images; Image posture transition module 4 performs pairing analysis on the first and last frames of the virtual fitting image, calculates the change vector between adjacent postures of the first and last frames of the virtual fitting image, and generates an inter-frame motion vector field. The motion vector field is injected into the diffusion model as input to generate posture transition parameters; The transition image sequence module 5 inputs the first and last frame virtual fitting images and posture transition parameters into the WAN video generation model, performs dynamic texture restoration on the image, and generates a continuous transition frame image sequence based on the first and last frames; The video verification and output module 6 generates a virtual fitting video based on the transition frame image sequence, and verifies and outputs the virtual fitting video.
[0059] In summary, with the help of the above technical solution of the present invention, the present invention realizes accurate segmentation and multi-angle modeling of facial and clothing areas by combining the image segmentation of the SAM model with the rigid geometric transformation algorithm, making the interaction between clothing and characters more natural. In particular, under multi-view changes and complex postures, it can maintain the true fit of the face and clothing, and adopts the Flux diffusion model to guide the generation process to gradually generate clear images. At the same time, Vista2Vogue is used to repair high-frequency textures to ensure the consistency of texture details of clothing under different posture changes. In addition, the Fill-Redux module is used to successfully repair high-frequency details in non-head areas when posture changes, avoiding common texture distortion problems.
[0060] In addition, the present invention ensures the natural coupling between clothing and human body posture changes by utilizing physical rationality calibration, avoids deformation anomalies by introducing deformation energy parameters and cloth dynamics models, and enhances the physical consistency of the generated image, making it more consistent with the actual wearing effect under different postures. At the same time, through time and space consistency constraints, it ensures smooth transition between frames in the virtual fitting video, so that the character remains consistent in the transition of multiple postures, avoiding common frame jumps or abruptness, and improving the smoothness and comfort of the user experience. Combined with the clothing texture feature fusion of the CLIP model, it can generate highly customized virtual fitting videos according to the action prompts and clothing needs of different users. Through the fusion of these technologies, high-quality simulation effects are achieved in visual effects, and a more personalized virtual fitting experience is provided for users.
[0061] In addition, the present invention uses the Vista2Vogue method to generate the first and last frame parts and reconstructs the multi-view generation process using explicit geometric constraints. The core idea is to use the rigid transformation of the head to construct a dual-view stitching reference image, and then use the spatial topology prior to lock the identity anchor and establish cross-view geometric associations. At the same time, the texture restoration of the non-head area is achieved through pixel-level loss and style transfer constraints, and a low-rank adaptive strategy is used to efficiently align the posture text semantics. Vista2Vogue outperforms advanced solutions in identity consistency, clothing structure restoration and multi-view coherence, while maintaining real-time efficiency, providing a deployable solution for high-fidelity virtual fitting.
[0062] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A virtual fitting video generation method based on multi-view face fixation, characterized in that: The following steps are involved: S1. Obtain original fitting images and action prompt words, perform image normalization and text vectorization processing respectively, and construct a multi-source input dataset; S2. Use the SAM model to segment the original fitting images in the multi-source input dataset, obtain the facial area and clothing area in the image, and extract the facial mask of the facial area. Then, based on the rigid geometric transformation algorithm, perform multi-view modeling on the facial mask to generate multi-angle facial images, and extract the clothing texture feature parameters of the clothing area. S3, based on the CLIP model, parses action prompt words, generates target pose sequences, and constructs a dual-view spatial topology prior framework to fuse multi-view facial images with clothing texture feature parameters to generate the first and last frame virtual fitting images; S4, performing pairing analysis on the first and last frames of the virtual fitting image, calculating the change vector between adjacent postures of the first and last frames of the virtual fitting image, and generating an inter-frame motion vector field, injecting the motion vector field as input into the diffusion model to generate posture transition parameters; S5. Input the first and last frame virtual fitting images and posture transition parameters into the WAN video generation model, perform dynamic texture restoration on the image, and generate a continuous transition frame image sequence based on the first and last frames; S6. Generate a virtual fitting video based on the transition frame image sequence, and verify the output of the virtual fitting video.
2. The method for generating virtual fitting videos based on multi-view face fixation according to claim 1, characterized in that: The method uses the SAM model to segment the original fitting image in the multi-source input data set, obtains the facial area and clothing area in the image, extracts the facial mask of the facial area, and then performs multi-view modeling on the facial mask based on the rigid geometric transformation algorithm to generate multi-angle facial images, and extracts clothing texture feature parameters of the clothing area. The method includes the following steps: S21. Use the SAM model to perform semantic segmentation on the original fitting images in the multi-source input dataset to obtain the facial region and clothing region in the image, and extract the facial mask of the facial region; S22, performing similarity transformation on the facial mask using a rigid geometric transformation algorithm to generate a head transformation map, and outputting the map as a multi-angle facial image; S23. Extract clothing texture feature parameters of the clothing area, and unify the clothing texture parameters of multi-angle facial images through a material-aware style transfer algorithm.
3. The method for generating virtual fitting videos based on multi-view face fixation according to claim 1, characterized in that: The CLIP model is used to parse action prompt words, generate target posture sequences, construct a dual-view space topology prior framework, fuse multi-view facial images with clothing texture feature parameters, and generate first and last frame virtual fitting images, which includes the following steps: S31, extracting the semantic feature vector of the action prompt word through the CLIP text encoder, and generating a continuous posture key point coordinate sequence in combination with the predefined posture dictionary; S32: The original fitting image and the geometrically transformed facial image are spliced into a left and right image, and the spatial attention mechanism is used to achieve feature alignment, and the SRIM mechanism is used to maintain identity consistency; S33. Generate high-fidelity images of the clothing area through the diffusion repair model, and use a high-frequency texture retention strategy to optimize details. Then, perform pixel-level fusion with the transformed facial area to obtain the first and last frame virtual fitting images.
4. The method for generating virtual fitting videos based on multi-view face fixation according to claim 1, characterized in that: The step of inputting the first and last frame virtual fitting images and posture transition parameters into the WAN video generation model, performing dynamic texture restoration on the images, and generating a continuous transition frame image sequence based on the first and last frames includes the following steps: S51. Load the first and last frame images according to the FLF2V model, quantize them through FP16 and set the frame number. At the same time, use the T5 text encoder to parse the posture transition parameters to obtain the motion guidance vector. The motion guidance vector is used to assist in controlling the dynamic strategy of generating inter-frame transitions, but is not used as an explicit input to the WAN model. After receiving the RGB first and last frame images output by Vista2Vogue, the S52 and WAN models rely on their built-in DiT spatiotemporal attention mechanism to achieve dynamic alignment and natural transition of clothing textures. At the same time, the WAN model implicitly models the coupling relationship between clothing and human posture changes through its VCU unit, and calculates the continuous change values of cloth wrinkles and motion amplitude; S53 , finely generating transition frames through its built-in frame interpolation module, and outputting a continuous and natural transition frame image sequence.
5. The method for generating virtual fitting videos based on multi-view face fixation according to claim 3, characterized in that: The process of splicing the original fitting image and the geometrically transformed facial image into a left and right image, using a spatial attention mechanism to achieve feature alignment, and maintaining identity consistency through a SRIM mechanism includes the following steps: S321, using the global features of the original image as identity anchors, aligning the identity key points in the dual images through a feature matching algorithm, and splicing the original fitting image and the geometrically transformed facial image into a left and right dual image; S322, adopting spatial attention mechanism and introducing context-aware mask to keep the texture of clothing area continuous and aligned with the original image; S323. By optimizing the spatial alignment error of the dual-view spatial topology prior framework, the interpolation distortion caused by posture transformation is suppressed, and identity consistency is maintained through the SRIM mechanism.
6. The method for generating virtual fitting videos based on multi-view face fixation according to claim 3, characterized in that: The method of generating a high-fidelity image of the clothing area by using a diffusion restoration model, optimizing details by using a high-frequency texture preservation strategy, and then performing pixel-level fusion with the transformed facial area to obtain the first and last frame virtual fitting images includes the following steps: S331, performing posture alignment on the reference image according to the target posture, and performing physical rationality calibration on the generated image after posture alignment to ensure the coupling relationship between the clothing and the human body posture; The S332 and Flux models guide the generation process by combining textual cues, and the Flux diffusion model generates clear images by gradually guiding noisy images; S333: When repairing details and local textures in clothing areas, Vista2Vogue uses a high-frequency texture preservation strategy to repair high-frequency textures and uses the Fill-Redux module to repair high-frequency details in non-head areas, while maintaining the consistency of details when the pose changes. S334 and ACE-LoRA use a low-rank adaptive optimization strategy to update some parameters of the model and adjust the image content during the generation process based on the alignment information; S335 and Vista2Vogue input the static generation results into the WAN model and achieve smooth transition between frames through spatiotemporal consistency constraints.
7. The method for generating virtual fitting videos based on multi-view face fixation according to claim 6, characterized in that: The method of performing posture alignment on the reference image according to the target posture, and performing physical rationality calibration on the generated image after posture alignment to ensure the coupling relationship between the clothing and the human body posture includes the following steps: Assume the target posture is , the pose of the reference image is , the posture alignment process is expressed as: ; in, Represents a geometric transformation function, which is implemented by rotation and scaling transformation; The generated images are calibrated for physical plausibility to ensure the coupling relationship between clothing and human body posture. This process can be expressed by the following formula: ; in, is the physical plausibility loss function, and represent the mechanical constraint and deformation constraint of the i-th pixel respectively.
8. The method for generating virtual fitting videos based on multi-view face fixation according to claim 6, characterized in that: When repairing the details and local textures of the clothing area, Vista2Vogue uses a high-frequency texture preservation strategy to repair high-frequency textures, and uses the Fill-Redux module to repair high-frequency details in non-head areas. Maintaining the consistency of details when the pose changes includes the following steps: The S3331 and Fill modules use image restoration algorithms to fill in missing details and combine contextual information to ensure the naturalness of the restoration effect. The high-frequency detail restoration process is optimized using the following loss function: ; in, is the high frequency detail recovery loss, Represents the repaired image of the current pixel p, is the true value of the target image; The S3332 and Redux modules use style transfer technology to uniformly optimize the textures of the head and non-head areas. The goal is to minimize texture distortion during posture transformation. The optimization is performed using the following formula: ; in, Optimize the results for textures, represents the texture after posture transformation, For reference texture.
9. A virtual fitting video generation system based on multi-view face fixation, used to implement the virtual fitting video generation method based on multi-view face fixation according to any one of claims 1 to 8, characterized in that: The system includes: The data acquisition and processing module obtains the original fitting images and action prompt words, performs image normalization and text vectorization processing respectively, and constructs a multi-source input dataset; The image segmentation processing module uses the SAM model to segment the original fitting images in the multi-source input data set, obtain the facial area and clothing area in the image, and extract the facial mask of the facial area. Then, based on the rigid geometric transformation algorithm, the facial mask is modeled from multiple perspectives to generate multi-angle facial images, and the clothing texture feature parameters of the clothing area are extracted. The virtual fitting image module parses action prompt words based on the CLIP model, generates target pose sequences, and constructs a dual-view spatial topology prior framework to fuse multi-view facial images with clothing texture feature parameters to generate the first and last frames of virtual fitting images. The image posture transition module performs pairing analysis on the first and last frames of the virtual fitting image, calculates the change vector between adjacent postures of the first and last frames of the virtual fitting image, and generates an inter-frame motion vector field. The motion vector field is injected into the diffusion model as input to generate posture transition parameters; The transition image sequence module inputs the first and last frame virtual fitting images and posture transition parameters into the WAN video generation model, performs dynamic texture restoration on the image, and generates a continuous transition frame image sequence based on the first and last frames; The video verification output module generates a virtual fitting video based on a transition frame image sequence and verifies the output of the virtual fitting video.
Citation Information
Patent Citations
Virtual fitting method and device
CN112884638A
Virtual fitting method under any human body posture
CN116342879A
Personalized virtual fitting method based on shape control and texture guidance
CN117670893A
Virtual fitting method of deep learning 2D picture
CN118505835A
Virtual fitting method and device
CN119444349A
Cited By
Digital human-oriented four-dimensional virtual fitting method and system
CN121582420A
A four-dimensional virtual fitting method and system for digital humans
CN121582420B
Layered modeling and double-domain optimization-based clothes edge leveling method
CN121708291A
Clothes edge flattening method based on hierarchical modeling and dual-domain optimization
CN121708291B