Virtual fitting method and device, storage medium and electronic device
By acquiring user videos for human posture detection and 3D model construction, virtual clothing is rendered to generate AR fitting videos, solving the problem of not being able to try on clothes purchased online, reducing the return and exchange rate, and improving user experience.
Patent Information
- Application Number
- CN202511033383.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Traditional online clothing purchases cannot provide a fitting experience, which leads to problems such as unsuitable sizes and unsatisfactory styles after purchase, resulting in high return and exchange rates and poor consumer experience.
By acquiring the user's target video, performing human posture detection and tracking processing, building a 3D human body model, and rendering virtual clothing based on a dynamic fitting algorithm, an AR fitting video is generated to show the user the fitting effect.
It effectively solves the problem of not being able to try on clothes purchased online, reduces the probability of returns and exchanges, and improves the user shopping experience.
Smart Images

Figure CN120525619B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer image processing, and in particular to a virtual fitting method and device, a storage medium and an electronic device. Background Art
[0002] With the widespread adoption of internet technology and the iterative upgrades of mobile devices, the global e-commerce market has experienced explosive growth. As a key category in the retail industry, apparel has become a core competitive area for e-commerce platforms due to its low degree of standardization, strong personalized demand, and high consumption frequency.
[0003] According to statistics, the scale of the global apparel e-commerce market has exceeded one trillion US dollars. Online shopping has broken through geographical and time limitations, and has reshaped consumers' shopping habits through massive product displays, precise marketing recommendations and convenient payment systems.
[0004] However, the traditional way of purchasing clothing online cannot provide consumers with a fitting experience, which results in consumers easily having problems such as inappropriate size and unsatisfactory style when purchasing clothing online, leading to high return and exchange rates and poor consumer experience. Summary of the Invention
[0005] In view of this, the embodiments of the present application provide a virtual fitting method and device, a storage medium and an electronic device. By applying the solution provided by the present application, the user can try on the selected clothing in the form of AR, so that the user can see the fitting effect of the clothing. The user can determine whether to purchase the clothing based on the effect, thereby effectively solving the problem of not being able to try on clothing purchased online, thereby reducing the probability of returns and exchanges of clothing purchased online, and providing users with a good shopping experience.
[0006] To achieve the above objectives, the present invention provides the following technical solutions:
[0007] In a first aspect, the present application discloses a virtual fitting method, comprising:
[0008] Obtain a target video containing a user to try on clothes;
[0009] Performing human posture detection and human posture tracking processing on the target video to obtain human posture data of the user in each video frame of the target video;
[0010] For each of the video frames, determining human body model parameters in multiple dimensions based on the human body posture data of the video frame, and constructing a 3D human body model of the user in the video frame using the respective human body model parameters;
[0011] Obtaining a virtual garment of a target garment from a preset garment model library; the target garment is the garment to be tried on selected by the user;
[0012] For each 3D human body model in the video frame, obtaining clothing rendering information of the 3D human body model, and rendering the virtual clothing on the 3D human body model based on a preset dynamic fitting algorithm and the clothing rendering information, to obtain a virtual rendered video frame corresponding to the video frame;
[0013] An AR fitting video is generated based on each of the virtual rendering video frames, and the AR fitting video is displayed to the user.
[0014] A second aspect of the present application discloses a virtual fitting device, comprising:
[0015] A first acquisition unit is configured to acquire a target video containing a user to be fitted with clothing;
[0016] A second acquisition unit is configured to perform human posture detection and human posture tracking processing on the target video to obtain human posture data of the user in each video frame of the target video;
[0017] a construction unit, configured to determine, for each video frame, human body model parameters in multiple dimensions based on human body posture data of the video frame, and construct a 3D human body model of the user in the video frame using the human body model parameters;
[0018] A third acquiring unit is configured to acquire a virtual garment of a target garment from a preset garment model library; the target garment is the garment to be tried on selected by the user;
[0019] a rendering unit, configured to obtain, for each of the 3D human body models in the video frame, clothing rendering information of the 3D human body model, and render the virtual clothing on the 3D human body model based on a preset dynamic fitting algorithm and the clothing rendering information, to obtain a virtual rendered video frame corresponding to the video frame;
[0020] A display unit is used to generate an AR fitting video based on each of the virtual rendering video frames and display the AR fitting video to the user.
[0021] A third aspect of the present application discloses a storage medium, which includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the virtual fitting method as described above.
[0022] In a fourth aspect, the present application discloses an electronic device comprising a memory and one or more instructions, wherein the one or more instructions are stored in the memory and configured to be executed by one or more processors to implement the virtual fitting method as described above.
[0023] Compared with the prior art, this application has the following advantages:
[0024] The present application discloses a virtual fitting method and device, storage medium and electronic device, including: processing a target video containing a user to be fitted, obtaining human body posture data for each video frame in the target video; using the human body posture data for each video frame to construct a 3D human body model for each video frame; for the 3D human body model of each video frame, obtaining clothing rendering information for the 3D human body model, rendering virtual clothing on the 3D human body model based on a dynamic fitting algorithm and clothing rendering information, and obtaining a virtual rendered video frame for the video frame; generating an AR fitting video based on each virtual rendered video frame, and displaying the AR fitting video to the user. By applying this application, the selected clothing can be tried on in the form of AR, showing the user the fitting effect of the clothing, effectively solving the problems of inappropriate size and mismatched styles caused by the inability to try on clothing purchased online, reducing the probability of consumer returns and exchanges, and providing users with a good shopping experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0026] Figure 1 A flowchart of a virtual fitting method provided in an embodiment of the present application;
[0027] Figure 2 A flow chart of a method for obtaining a target video containing a user to try on clothes provided in an embodiment of the present application;
[0028] Figure 3 A flowchart of rendering a virtual garment on a 3D human body model based on a preset dynamic fitting algorithm and garment rendering information provided in an embodiment of the present application to obtain a virtual rendered video frame corresponding to the video frame;
[0029] Figure 4 A flowchart of another virtual fitting method provided in an embodiment of the present application;
[0030] Figure 5 A schematic structural diagram of a virtual fitting device provided in an embodiment of the present application;
[0031] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0033] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0034] The present application can be used in an intelligent video dynamic fitting system or an AR fitting system composed of a plurality of general or special computing device environments or configurations.
[0035] Reference Figure 1 , is a flow chart of a virtual fitting method provided in an embodiment of the present application, and is specifically described as follows:
[0036] S101: Obtain a target video containing a user to try on clothes.
[0037] In the embodiment provided in the present application, the target video may be a pre-processed video, and the pre-processing process includes scene-adaptive denoising, fitting scene-specific color correction, fitting optimization resolution standardization, human motion perception frame rate stabilization, fitting enhancement pre-processing and other aspects.
[0038] Reference Figure 2 , which is a flow chart of a method for obtaining a target video containing a user to try on clothes provided in an embodiment of the present application, and is specifically described as follows:
[0039] S201: Obtain an initial video containing a user to try on clothes.
[0040] The initial video can be a real-time recorded video or a video file uploaded by the user.
[0041] S202: Determine scene complexity in the initial video, and adjust filter parameters of a preset denoising filter based on the scene complexity.
[0042] The preset denoising filter may be a combination of improved bilateral filtering and Gaussian filtering, and the filtering parameters of the denoising filter may be dynamically adjusted according to the impetuousness of the scene.
[0043] S203 : Process the initial video using the denoising filter with adjusted parameters to obtain a first video.
[0044] When processing the initial video using the denoising filter with adjusted parameters, the initial video is analyzed through pixel neighborhood analysis, and filters of different intensities are applied to the human body area and background area in the initial video to ensure the clarity of the human body contour; and for the fitting scene of the initial video, while retaining the human body edge details in the video, noise is removed, and adaptive window size (5×5 to 9×9) and σ value (0.6-1.2) can be used for processing.
[0045] Thus, the scene-adaptive denoising process of the initial video is completed, and the first video is obtained.
[0046] S204: Correct the main color tone of the scene in the first video using a preset color correction strategy to obtain a corrected second video.
[0047] The process of correcting the main color tone of the scene in the first video using the color correction strategy may include: automatically detecting and correcting the color temperature difference according to the main color tone of the scene, enhancing the local contrast in the video, thereby improving the distinction between the edge of the clothing and the outline of the human body, and when improving the distinction between the edge of the clothing and the outline of the human body, the dynamic adjustment range of the adjustment weight γ is 0.85-1.15; the video can also be color-standardized to ensure the consistency of the color of the clothing in the video under different lighting conditions, thereby completing the correction of the scene-specific color of the fitting room in the video; thereby, the scene-specific color correction processing of the fitting room of the first video is completed, thereby obtaining the second video.
[0048] S205: Use a preset resolution optimization strategy to perform resolution optimization processing on the second video to obtain a third video.
[0049] The content of the resolution optimization processing of the second video includes: dynamically adjusting the optimal resolution according to the proportion of the human body in the picture to avoid information loss caused by fixed resolution; realizing regional perception of non-uniform scaling, maintaining higher resolution for important areas of the human body, and moderately reducing the resolution of background areas; applying an edge-preserving interpolation algorithm to ensure that the edges and texture details of the clothing are not blurred during the resolution transformation; thus, completing the fitting optimization resolution standardization processing of the second video and obtaining the third video.
[0050] S206: Use a preset frame rate stabilization strategy to perform frame rate stabilization processing on the third video, and perform fitting enhancement processing on the processed video based on a preset fitting enhancement strategy to obtain a target video.
[0051] Using the preset frame rate stabilization strategy, the process of performing frame rate stabilization processing on the third video is as follows: developing an adaptive frame rate control algorithm based on human motion speed; based on the algorithm, increasing the sampling rate when the user in the video moves quickly, and reducing the frame rate when the user is stationary or moves slowly; implementing intelligent interpolation technology assisted by motion prediction; focusing on calculations of clothing wrinkles and fluttering areas in the video; using an inter-frame consistency maintenance mechanism to ensure temporal continuity in subsequent clothing rendering and suppress flickering; thereby completing the processing of the human motion perception frame rate and obtaining the processed video.
[0052] The processed video is subjected to fitting enhancement processing, specifically including: enhancing the texture of the user's clothing in the video to improve the recognizability of fabric texture and pattern; sharpening the human body contour in the video to enhance the accuracy of subsequent posture recognition, using a directional sharpening operation with an intensity factor of 0.4-0.8; establishing a clothing region of interest map (ROI map) on the user's portrait in the video to provide priority guidance for subsequent processing; thus, completing the fitting enhancement preprocessing to obtain the target video.
[0053] In the embodiments provided in the embodiments of the present application, the accuracy of subsequent human body contour recognition is improved by preprocessing the video, and the accuracy of human body contour recognition is also improved under complex backgrounds and low light conditions. In addition, the texture detail retention of clothing in the video is improved, which effectively reduces the error rate of subsequent posture detection, reduces subsequent processing delays, improves the real-time processing capability of the device, improves the overall stability of the system, and reduces the abnormal frame rate of the system.
[0054] The process of preprocessing videos in this application is different from the traditional commonly used video preprocessing process. Compared with the existing video preprocessing, the video preprocessing method provided by this application has higher accuracy in human body contour recognition, higher texture detail retention, lower posture detection error rate, less processing delay, and less abnormal frame rate, making the performance of the entire system stronger.
[0055] S102: Perform human posture detection and human posture tracking on the target video to obtain human posture data of the user in each video frame of the target video.
[0056] In an embodiment of the present application, for each video frame of the target video, a preset posture detection module is used to perform human key point detection on the user's human body image to obtain human skeleton data of the video frame, and a preset multi-person separation network is used to perform human instance segmentation on the video frame to determine the user's human body instance in the video frame, and based on the human body instance, the user's skeleton key point data is determined in the human skeleton data; the correlation between the user in the video frame and in adjacent video frames is established, and the motion trajectory of each skeleton key point of the user between the video frame and the adjacent video frames is smoothed to obtain human posture data of the video frame.
[0057] The human body posture data of the video frame includes but is not limited to information on the user's various skeletal key points, association information between the user in the video frame and between connected video frames, and motion trajectory information on the user's various skeletal key points in the video frame and between adjacent video frames.
[0058] The preset posture detection module includes a self-developed dual-path hierarchical posture estimation network (DP-HPENet), which combines a lightweight improved ResNet-50 with a transformer hybrid architecture and can identify 17-25 skeletal key points including the head, shoulders, elbows, wrists, hips, knees and ankles. The system not only efficiently identifies the spatial position of skeletal key points, but also evaluates the reliability of each point through a key point confidence self-calibration mechanism, significantly improving the detection stability in fast-moving scenes. DP-HPENet can be used to detect human key points in each video frame, and combined with an adaptive spatiotemporal attention module to dynamically adjust the receptive field size, thereby improving detection accuracy and reducing the computational complexity of the detection process.
[0059] In another embodiment provided in the present application, the posture detection module also includes a hierarchical occlusion inference module (ST-GCN-HOI) based on a spatiotemporal graph convolutional network, which can accurately distinguish between three situations: self-occlusion, occlusion by others, and occlusion by scene objects, and infer the positions of the occluded skeletal key points. By adopting a differentiated inference strategy, the positions of the skeletal key points are inferred in severe occlusion scenarios, which can improve the accuracy of the inference results.
[0060] It should be noted that when multiple people appear in a video, the resulting human skeleton data contains skeletal data for multiple people. A multi-person segmentation network is used to segment the video frames to identify the user's human instances within the video frames. The skeleton keypoint data corresponding to each user's human instance is then determined within the skeleton data. This skeleton keypoint data contains the positional information of multiple skeletal keypoints of the user.
[0061] It should be noted that the multi-person segmentation network incorporates boundary perception and can be called the boundary-aware adaptive multi-person segmentation network (BA-MPSNet). By using this network, high-precision human instance segmentation can be achieved. Human instance segmentation is a crucial link in AR fitting systems. It is not only the basis for determining the target person (i.e., the user), but also a prerequisite for the subsequent precise fitting of virtual clothing.
[0062] The BA-MPSNet (Boundary-Aware Adaptive Multi-Person Segmentation Network) proposed in this application adopts an improved instance segmentation architecture for high-precision human segmentation. The instance segmentation architecture includes edge-aware mechanism, posture-guided segmentation attention mechanism, adaptive resolution processing, and semantically enhanced segmentation.
[0063] A specific edge enhancement module is introduced into the edge perception mechanism, which enhances the detection capability of human body contour boundaries through multi-level gradient feature extraction, enabling the system to accurately capture the boundaries between the human body and the environment or other people, and the boundary accuracy is improved by 15.3% compared with conventional methods; the posture-guided segmentation attention mechanism uses previously detected bone information to guide the segmentation process, forming a skeleton-segmentation joint optimization loop, so that the segmentation results are highly consistent with the human skeletal structure; in adaptive resolution processing, computing resources are dynamically allocated according to the complexity of the human body area, and high-resolution processing is used in areas with complex details (such as hair, fingers, etc.), while low-resolution processing is used for flat areas; semantic enhanced segmentation not only distinguishes between people and backgrounds, but also identifies different areas such as clothing and skin, providing semantic information for the subsequent precise fitting of virtual clothing.
[0064] By performing human instance segmentation, different individuals can be accurately distinguished in multi-person scenes, avoiding misplacement of fitting effects or application to non-target people, providing accurate human boundary information for virtual clothing, allowing clothing to fit the human body naturally, helping the system understand the occlusion relationship between different parts of the human body, and ensuring that virtual clothing can be correctly rendered even in the case of occlusion. Combining the segmentation results with depth estimation, the system can more accurately understand the position and posture of the human body in 3D space, provide more accurate hand segmentation results for subsequent gesture recognition, and enhance the user's interactive experience with virtual clothing; through high-quality human instance segmentation, the system can still maintain accurate fitting effects in complex backgrounds and multi-person scenes, which is an advantage that traditional virtual fitting systems based on single-person static images cannot achieve.
[0065] It should be noted that when determining the user to be tried on (who can be regarded as the target user), this can be determined through the interaction between the system and the user. When the user opens the AR fitting application, the user can specify the target user through gesture selection (such as pointing to a specific person), voice commands or touch screen operations. This is the most basic recognition method and is suitable for scenarios where the user actively selects.
[0066] Secondly, for complex scenes, the system introduces a boundary-aware adaptive multi-person segmentation network (BA-MPSNet) for high-precision human instance segmentation, and combines it with position priority analysis to determine the target person. BA-MPSNet improves the accuracy of human contours by integrating an edge enhancement module, and can still accurately distinguish different individuals in complex scenes with overlapping boundaries. Specifically, the system will make judgments based on at least one of the following priority rules: (1) Center area priority: people in the center area of the screen are usually determined as target users; (2) Size priority: people who occupy a larger proportion of the screen are given priority as target users; (3) Foreground priority: based on depth estimation, foreground people have a higher priority as target users than background people.
[0067] This application can also keep the target person's ID tracking to avoid the target switching in multi-person cross-scenes. During the continuous interaction process, the system maintains continuous recognition of the target person through a temporal memory mechanism. Even if the target person is temporarily blocked or re-enters after leaving the screen, the system can re-identify the same person based on the previously established feature descriptors (including clothing texture, body shape features, facial features, etc.), achieving continuity of the fitting experience.
[0068] By establishing the association between the user in the video frame and in adjacent video frames, the user's temporal tracking is achieved. The specific process is as follows: the adaptive key point association framework (MF-AKA) based on multi-dimensional feature fusion is used to establish inter-frame character associations. The dynamic feature weight allocation and two-stage association strategy can be used to improve the ID retention rate in complex scenarios and reduce the ID switching rate by 31.4%. At the same time, the hierarchical skeleton-constrained Kalman filter tracking framework (HS-KF) can also be applied for multi-granularity motion prediction.
[0069] It should be noted that dynamic feature weight allocation is one of the core innovations of MF-AKA (Adaptive Keypoint Association Framework for Multi-dimensional Feature Fusion). It is closely related to skeletal keypoints. The dynamic feature weight allocation mechanism of this system is based on real-time scene analysis and adaptively adjusts the weight ratio of each dimension feature according to different types of skeletal keypoints and motion states. The specific implementation process is as follows:
[0070] 1) Multi-dimensional feature extraction: The system first extracts a multi-dimensional feature vector from each key point. The multi-dimensional feature vector includes but is not limited to: spatial position features (x, y coordinates and depth estimation), appearance features (visual descriptors of the area around the key point), motion features (speed and acceleration information of the key point), structural features (relative position relationship of the key point to other key points), and temporal features (motion trajectory features of the key point in the previous frames);
[0071] 2) Scenario Adaptive Analysis: The system evaluates the characteristics of the current scene in real time;
[0072] 3) Movement intensity: detect whether the human body is moving fast or slow;
[0073] 4) Occlusion: Evaluate whether the key point is in an occlusion state;
[0074] 5) Light changes: determine whether the ambient light is stable;
[0075] 6) Perspective change: Detect whether the camera perspective has changed significantly;
[0076] 7) Keypoint-specific weight allocation: Different weighting strategies are used for different skeletal keypoints. For example, for stable keypoints (such as shoulders and hips), the weight of structural features is increased; for fast-moving keypoints (such as wrists and ankles), the weight of motion features is increased; for easily occluded keypoints (such as elbows in side view), the weight of appearance features is increased.
[0077] 8) Dynamic weight update mechanism: The system uses an iterative weighted learning algorithm to continuously optimize the weight allocation strategy based on the success rate and prediction error of skeleton key point tracking. Specifically:
[0078] ;
[0079] in, represents the weight of feature f at time t, α and β are learning rate parameters, and the specific values can be set according to actual needs; represents the tracking success rate, which refers to the ratio of feature f to key points successfully associated during the tracking process; Represents the prediction error, which refers to the position or state deviation when using feature f for prediction.
[0080] Specifically, It is used to measure the effectiveness of this feature in maintaining the identity continuity of human skeleton key points. The higher the value, the more valuable the feature is for stable tracking of specific key points (such as wrists, elbows, etc.). In actual calculation, it is usually the number of times the key points are correctly associated between consecutive frames divided by the total number of tracking frames; Used to measure the accuracy of key point position prediction based on this feature. The lower the value, the more accurately the feature can predict the future position of the key point. It is usually expressed in pixel distance or normalized distance.
[0081] The correlation between this dynamic feature weight distribution and key points is reflected in the following: the system will formulate the best weight strategy for the characteristics of each key point (such as joint type, movement characteristics, and occlusion frequency). For example, for key points with high-speed movement such as the wrist, the system will dynamically increase the weight of the motion features; while for relatively stable key points such as the head, the weight of the appearance features will be increased. Through this dynamic adjustment of feature weights, the system can maintain stable key point tracking performance in complex scenarios, effectively reducing the ID switching rate by 31.4%, which is crucial to ensuring the continuity of the fitting experience.
[0082] The MF-AKA framework can be used to achieve efficient and stable inter-frame character association, that is, to establish the association between the user in the video frame and the adjacent video frames. This application adopts a multi-level association strategy, that is, a "coarse to fine" three-layer association method, as follows:
[0083] 1) Global character association: First, establish the correspondence between character instances in the previous and next frames at the overall level;
[0084] 2) Body Part Association: The human body is divided into main parts such as head, torso, limbs, etc. and sub-region association is performed;
[0085] 3) Fine association of key points: Finally, accurately associate each key point to establish a complete skeleton correspondence;
[0086] 4) Spatiotemporal consistency matching, specifically: short-term association: predicting the possible position of key points in the next frame through Kalman filtering and establishing a matching cost matrix; medium-term association: maintaining a motion trajectory model of 5-15 frames and evaluating the degree of matching through trajectory similarity; long-term association: establishing a long-term identity descriptor based on the person's appearance features to support long-term tracking and re-identification;
[0087] 5) Two-stage association decision: In the first stage, a greedy matching algorithm is used to quickly associate high-confidence matching pairs; in the second stage, for fuzzy matching, a multiple hypothesis tracking (MHT) algorithm is introduced to delay the decision until sufficient evidence is collected;
[0088] 6) Conflict resolution mechanism: Conflict detection based on skeleton structure constraints identifies physically impossible matching results; uses the Hungarian algorithm to optimize association results globally, thereby achieving a global optimal solution; and compares forward and reverse matching results based on a cross-validation strategy to improve association reliability.
[0089] 7) Occlusion and Re-identification Processing: Occlusion Prediction: When potential occlusion is detected, the system pre-activates the appearance descriptor memory; Tracking During Occlusion: For occluded key points, the possible position is inferred based on structural constraints; Re-identification: When a person reappears, the appearance matching and motion consistency are combined to determine;
[0090] 8) Temporal Memory Mechanism: Maintains a cached character feature library containing all recently appeared character features, applies a sliding window strategy to balance real-time performance and stability, and a feature update mechanism to continuously update the character feature representation based on new observations.
[0091] Through the above methods, the system can accurately track the target person in complex multi-person scenes, maintaining stable ID association even in partial occlusion, when people intersect, or when they briefly leave the field of view, providing continuous and consistent human skeleton data for AR fitting. This high-precision inter-frame character association is the foundation for a natural and smooth dynamic fitting experience.
[0092] The core purpose of multi-granularity motion prediction is to improve the system's ability to accurately understand and predict human motion, especially to maintain stable tracking performance at different time scales and motion complexities. In this system, this technology is implemented through the hierarchical skeleton-constrained Kalman filter tracking framework (HS-KF). Its main advantages include: 1) Adapting to different types of motion, multi-granularity prediction can handle multiple types of motion at the same time, including but not limited to: slow continuous motion (such as standing, slow walking), sudden fast motion (such as turning, waving), periodic motion (such as walking, running cyclical movements), non-periodic complex motion (such as various posture changes when trying on clothes); 2) Improving tracking stability. By comprehensively analyzing motion patterns at different time scales, the system can reduce tracking loss caused by sudden motion, improve tracking stability by 43.2%, cope with interference factors such as camera shake, maintain the stability of the skeleton structure, and so on. Maintain accurate motion prediction under low frame rate conditions and adapt to network fluctuations; 3) Improve the realism of physical simulation. Accurate motion prediction is crucial to the physical simulation of clothing. It can accurately predict the future positions of key points, so that virtual clothing can follow human movement more naturally, predict changes in motion acceleration, and correctly simulate inertial effects, such as the fluttering of clothes, changes in wrinkles, etc., capture subtle changes in body posture, and achieve more refined interactions between clothing and the human body; 4) Computing resource optimization. Multi-granularity prediction allows the system to allocate computing resources at different levels. For example, high-precision multi-model prediction is used for important key points (such as joints), and simplified model prediction is used in secondary areas to reduce the overall computing burden; computing resources are dynamically allocated according to the difficulty of prediction to improve system response speed.
[0093] Multi-granularity motion measurement can divide the human skeleton into multiple levels, including global position, trunk, and limbs, with different prediction models used at each level. For the prediction model at the global position level, it predicts overall displacement and rotation and captures large-scale motion. For the prediction model at the trunk level, it maintains the stability of the torso structure and handles posture changes. For the prediction model at the limb level, it handles flexible and changeable arm and leg movements. For the prediction model at the detail level, it captures fine movements of fingers, face, etc. It should be noted that different motion states also use different data models. For example, linear models are used for predictions during uniform motion, acceleration models are used for predictions during variable speed motion, and periodic models are used to identify and predict repetitive actions. Data-driven models can learn complex motion patterns based on historical data. Furthermore, multiple models can be fused for prediction, such as using prediction models at different skeletal levels and corresponding data processing models for prediction. When making multi-model fusion predictions, human biomechanical constraints are introduced to ensure that the prediction results conform to physical laws, and joint angle restrictions are imposed to prevent unnatural postures from being predicted. Speed and acceleration constraints are added to make the motion conform to the range of human motion capabilities, and bone length fixed constraints are added to maintain the consistency of human structure.
[0094] By using multi-granularity motion prediction, the system can achieve accurate understanding of human motion in complex dynamic scenes, provide high-quality skeletal motion data for the AR fitting system, achieve natural coordination between virtual clothing and real human motion, and greatly enhance the realism and user experience of dynamic fitting.
[0095] In the embodiments provided in the present application, when smoothing the motion trajectories of the user's various skeletal key points between video frames and adjacent video frames, the motion-aware adaptive smoothing algorithm (MA-ASA) can be used to intelligently process the trajectories of each skeletal key point, dynamically adjust the smoothing parameters according to different motion states, and combine the second-order derivative constraints to retain key acceleration information, thereby suppressing the improvement of the jitter effect while maintaining the improvement of the detail retention rate.
[0096] In the embodiment provided in this application, the trajectory of the skeleton key points can be obtained through multi-stage processing, for example:
[0097] 1) Initial trajectory construction
[0098] Through the aforementioned DP-HPENet network, the system identifies the spatial positions of 17-25 human key points in each frame of video; uses the temporal tracking module (MF-AKA framework) to establish the correspondence between each skeleton key point between frames, ensuring the identity consistency of the same skeleton key point between different frames; based on the corresponding skeleton key points in these consecutive frames, the system constructs the initial trajectory sequence ,in, Indicates the The spatial position of the keypoint in the frame.
[0099] 2) Trajectory completion and correction
[0100] For missing trajectories due to occlusion or detection failure, the system uses the ST-GCN-HOI occlusion inference system to estimate the position. For low-confidence detection results, the spatiotemporal graph convolutional network is used to make corrections to improve the continuity of the trajectory. The sliding window interpolation method is used to handle short-term missing trajectories, and long-term missing trajectories are inferred in combination with skeleton structure constraints.
[0101] 3) Multi-scale trajectory representation
[0102] Multi-scale trajectory representation includes: micro-trajectory, macro-trajectory and structured trajectory. Among them, micro-trajectory can capture high-frequency, small-amplitude key point movements and is suitable for subtle gesture recognition; macro-trajectory focuses on low-frequency, large-amplitude overall movement trends and is used for posture understanding; structured trajectory can combine multiple key point trajectories into structural units (such as "arm" consists of shoulder, elbow and wrist) to analyze the overall structural movement.
[0103] 4) Trajectory feature extraction
[0104] Trajectory features include position features, velocity features, acceleration features, angle features and collaborative features; position features are used to record the absolute coordinates of key points in each frame; velocity features are used to calculate the displacement changes of key points between adjacent frames; acceleration features are used to analyze velocity changes and capture the start and end of the action; angle features are used to calculate joint angle changes and describe posture changes; collaborative features are used to analyze the relative motion relationship of multiple key points.
[0105] 5) Data smoothing and enhancement
[0106] MA-ASA (motion-aware adaptive smoothing algorithm) can be applied to process the raw trajectory data. The smoothing strength can be dynamically adjusted according to the intensity of the movement. Weaker smoothing is used in fast-moving sections to retain dynamic details; stronger smoothing is used in the stable stage to effectively suppress jitter and noise. Combined with second-order derivative constraints, it ensures that the smoothing process does not lose key acceleration information and retains motion characteristics. The obtained high-quality key point trajectory is the basis for multiple subsequent functions.
[0107] By providing human motion data for physical simulation through the motion trajectory of skeletal key points, virtual clothing can be driven to deform naturally with the human body and support action recognition, enabling the system to understand the user's specific posture (such as rotation, arm lifting, etc.), providing accurate human motion information for dynamic clothing fitting, and assisting the system in predicting human posture in the next few frames, thereby reducing rendering delays. Through the above process, the system can obtain stable, continuous, and accurate key point trajectory data. After being processed by the MA-ASA algorithm, this data not only retains the dynamic characteristics of human motion, but also effectively suppresses detection jitter and noise, providing high-quality basic human motion data for subsequent AR fitting experiences.
[0108] Through a series of processing, human posture data for each video frame is obtained. The human posture data includes human skeleton data, association information of skeletal key points, trajectory information of skeletal key points, and other content. By applying the solution provided in the application, the high-quality human posture data obtained for each video frame is key information for understanding multi-person activities and interactions, ensuring that the system can achieve stable and accurate multi-person posture analysis under frequency-limited video conditions. Through the synergistic effect of the above-mentioned innovative technologies, the overall performance of the system is improved compared to existing technologies, especially showing significant advantages in challenging scenarios such as complex environments, multi-person interactions, and rapid movements.
[0109] S103 : For each video frame, determine human body model parameters of multiple dimensions based on human body posture data of the video frame, and use the human body model parameters to construct a 3D human body model of the user in the video frame.
[0110] Human body posture data includes but is not limited to human body shape information and human body posture information. Human body shape information includes but is not limited to the user's height, weight, fatness and other human body shape characteristics. Human body posture information includes content used to determine the human body's movement posture, such as the position and angle of each skeletal key point.
[0111] In the embodiment provided in the present application, for each video frame, based on the various human body indicators of the preset deformable human body model, the indicator parameters of each human body indicator are extracted from the human body posture data; the various detail description parameters are extracted from the video frame; and the various indicator parameters and the various detail description parameters are determined as various human body model parameters.
[0112] Various human body indicators include, but are not limited to, height, shoulder width, waist circumference, hip circumference, and other indicators related to human body shape. Various detailed description parameters include, but are not limited to, parameters for details such as clothing folds and muscle contours. It should be noted that the indicator parameters for various human body indicators are the same across different video frames.
[0113] In the embodiment provided by this application, the proportions and perspectives of the characters in the video are analyzed, and the depth estimation algorithm is combined to deduce the real index parameters of various human body indicators. This application uses a multi-perspective adaptive depth fusion framework to deduce the index parameters of various human body indicators. Specifically, the human motion information in the video frame sequence is integrated to construct a spatiotemporal depth field. When the user makes different postures, the system captures the depth information from different angles and matches it through the skeleton key points to achieve equivalent multi-perspective depth reconstruction, thereby obtaining the initial depth estimation result. This application also develops a human body morphology prior database, which includes more than 100,000 accurate 3D scan data of different body shapes. The database is used to constrain and optimize the initial depth estimation results, solving the common depth jump problem of traditional depth estimation at the human body contour.
[0114] In the embodiments provided in this application, a parameterized human body model SMPL (Skinned Multi-Person Linear Model) is used to generate a highly personalized 3D human body model. The SMPL model can generate a 3D human body model by adjusting morphological parameters and posture parameters. The SMPL model contains approximately 6,890 vertices and 13,776 faces, which is sufficient to represent the detailed appearance of the human body while maintaining computational efficiency.
[0115] In the process of using various human body model parameters to construct the user's 3D human body model in the video frame, the various indicator parameters in the various human body model parameters are used to construct a basic SMPL model. Then, the surface geometry of the basic SMPL model is adjusted using various detail description parameters to enhance the realism of the model. In the process of adjusting the representation geometry of the basic SMPL model, surface properties such as surface normal vectors, curvature, and deformation hotspots are calculated, and the surface of the basic SMPL model is adjusted using the calculated surface properties. The temporal consistency of the 3D human body model in each video frame is also guaranteed. By ensuring smooth changes in human body model parameters and avoiding sudden changes in model morphology, the reconstructed 3D human body model changes naturally with the user's movements.
[0116] The reconstructed 3D human body model accurately expresses the user's body shape and posture, providing the necessary "digital human body" foundation for subsequent virtual clothing fitting. This digital human body not only contains static shapes, but can also dynamically respond to various posture changes of the user in the video. It is the key prerequisite for this application to achieve realistic AR fitting effects.
[0117] S104: Acquire a virtual garment of a target garment from a preset garment model library; the target garment is the garment to be tried on selected by the user.
[0118] Collect clothing data of the target clothing; and determine a virtual clothing of the target clothing in a clothing model library based on the clothing data.
[0119] The clothing data of the target clothing includes but is not limited to the clothing number, clothing type, style, clothing seller, etc.
[0120] The clothing model library includes multiple virtual garments of clothing. Virtual clothing can be regarded as clothing models. Virtual clothing can be made by professional 3D design software or obtained by 3D scanning of actual clothing. It contains information such as the clothing's geometric shape, material, texture, and physical parameters.
[0121] The clothing number in the clothing data can be used to find a virtual garment of the target clothing in the clothing model library.
[0122] S105. For the 3D human body model of each video frame, obtain clothing rendering information of the 3D human body model, and render virtual clothing on the 3D human body model based on a preset dynamic fitting algorithm and the clothing rendering information to obtain a virtual rendered video frame corresponding to the video frame.
[0123] It should be noted that clothing rendering information includes clothing size information, environment information, action prediction information, visual perception information, augmented reality texture information and scene parameter information.
[0124] Reference Figure 3 , which is a flowchart of rendering a virtual garment on a 3D human body model based on a preset dynamic fitting algorithm and garment rendering information provided in an embodiment of the present application, to obtain a virtual rendered video frame corresponding to the video frame, is specifically described as follows:
[0125] S301: Determine a clothing category of the virtual clothing based on clothing size information of the virtual clothing in the clothing rendering information, and render the virtual clothing at a position corresponding to the clothing category in the 3D human body model according to the clothing size information.
[0126] The clothing size information of the virtual clothing includes information for describing the clothing type of the virtual clothing. The clothing categories include but are not limited to pants, skirts, dresses, short sleeves, suits, etc.
[0127] The clothing size information also includes size data corresponding to the clothing type. For example, if the clothing type is short-sleeved, the size data includes but is not limited to shoulder width, sleeve length, chest circumference, waist circumference, etc. If the clothing type is trousers, the size data includes but is not limited to thigh circumference, trouser length, etc.
[0128] The 3D human body model is divided into 37 functional semantic regions, and a nonlinear mapping relationship is established between each functional semantic region and the virtual garment. The virtual garment is then rendered on the 3D human body model. It should be noted that different parts of the virtual garment correspond to different locations on the 3D human body model, and some locations on the 3D human body model do not correspond to the virtual garment; for example, if the garment type is pants, the corresponding locations on the 3D human body model are the legs and hips.
[0129] When rendering virtual clothing on a 3D human body model, a suitable initial rendering position can be determined on the 3D human body model according to the clothing category of the virtual clothing, and then the virtual clothing can be rendered from the initial rendering position. For example, the initial rendering position of a top on the 3D human body model is above the shoulder; the initial rendering position of pants on the 3D human body model is at the waist. Predefined fitting rules and clothing semantic information can be used to ensure the rationality of the starting state.
[0130] In the method provided in the embodiment of the present application, an adaptive intelligent sizing system (AISS) can be used to determine the clothing category of the virtual clothing based on the clothing size information of the virtual clothing, and render the virtual clothing at a position corresponding to the clothing category in the 3D human body model based on the clothing size information. When rendering the virtual clothing on the 3D human body model, a layered adjustment strategy can be adopted, including a skeleton layer (such as shoulder width, sleeve length), a contour layer (such as chest circumference, waist circumference), a detail layer (such as neckline, pleat distribution), and other multiple layers for multi-level size optimization, identifying and adapting to special body shape features (such as special body shapes such as hunched chest, upright chest, round shoulders, etc.), and making targeted adjustments to the clothing structure, thereby improving the rendering quality of the virtual clothing.
[0131] S302: Constructing a signed distance field of the 3D human body model after rendering the virtual clothing. The signed distance field is used to describe the spatial relationship between the virtual clothing and the 3D human body model.
[0132] In the embodiments provided herein, a signed distance field (SDF) is a three-dimensional function that describes the shortest distance from each vertex of a virtual garment to the surface of a 3D human body model in three-dimensional space, where the three-dimensional space encompasses both the virtual garment and the 3D human body model. Using a signed distance field allows for rapid determination of the spatial relationship between the virtual garment and the 3D human body model, enabling efficient collision detection and fit calculations.
[0133] In the embodiments provided in this application, a spatially adaptive sampling strategy is introduced into the directed distance field. This strategy increases the sampling density in areas with significant changes in human curvature (such as joints and neck) and reduces the sampling density in flat areas, achieving a balance between accuracy and efficiency. The specific calculation expression is:
[0134] ;
[0135] Where p is a spatial point in three-dimensional space, q is a point on the human body surface S, and sign(p) indicates whether point p is inside the body (-1) or outside (+1). This adaptive SDF technique reduces the computational effort compared to traditional uniform sampling methods while maintaining accuracy in key areas.
[0136] S303 : Based on the preset optimization strategy and the signed distance field, the rendering of the virtual clothing on the 3D human body model is adjusted until a preset convergence condition is met, and then the adjustment of the virtual clothing is stopped.
[0137] In the embodiment provided in this application, the optimization strategy is an iterative optimization algorithm, which uses a preset optimization strategy and a directed distance field to adjust the vertex positions of the virtual clothing in the rendering of the 3D human body model, so that the virtual clothing fits the 3D human body model better. The adjustment of the virtual clothing is stopped until the preset convergence condition is met.
[0138] It should be noted that the vertex positions of the virtual garment refer to the coordinates of the mesh vertices constituting the garment in three-dimensional space.
[0139] The iterative adjustment process is as follows:
[0140] Step 1: First, address the penetration problem between the virtual garment and the 3D human body model, and push the garment part that penetrates the 3D human body model to the surface of the 3D human body model; then apply internal constraints of the garment to maintain the continuity and structural integrity of the fabric; finally, balance the external force and internal constraints to find the energy minimization state.
[0141] Furthermore, the system adopts an innovative hierarchical optimization strategy to decompose the fitting problem into three sub-problems to be solved sequentially, as shown in steps 2 to 4.
[0142] Step 2: Global position alignment.
[0143] The virtual garment is rigidly transformed as a whole so that key feature points (such as the neckline and cuffs) are roughly aligned with the corresponding parts of the human body, providing a good initial state for local fitting.
[0144] Step 3: Collision response iteration.
[0145] The system uses the aforementioned distance field to quickly detect the penetration between clothing and the human body, and applies a gradient to each penetration vertex.
[0146] Step 4: Displacement in the degree direction.
[0147] ;
[0148] in, is an adaptive step size parameter, which is dynamically adjusted according to the penetration depth; is the new position of the vertices of the virtual garment; is the original position of the vertices of the virtual garment; is the SDF value of the original position of the vertex of the virtual clothing, is the SDF gradient direction of the original position of the vertex of the virtual clothing.
[0149] Internal constraint maintenance: After each collision response, the system solves the fabric internal constraint equations:
[0150] ;
[0151] Among them, M is the mass matrix, K is the stiffness matrix, d is the deformation vector, For external force.
[0152] The above steps are iterated until a preset convergence condition is reached. The preset convergence condition can be that the maximum displacement is less than a threshold or the maximum number of iterations has been reached. This hierarchical iterative method converges faster than traditional global optimization methods while ensuring physical rationality.
[0153] In the embodiments provided herein, this solution introduces a nonlinear elastic model based on tetrahedral finite elements, which accurately simulates the physical behavior of anisotropic fabrics under large deformations. The system calculates the fabric strain tensor in real time and, when potential abnormal stretching is detected, automatically adjusts the local mesh structure and physical parameters to effectively prevent non-physical deformation or tearing. It also provides an intelligent seam processing framework that implements seam modeling technology. This framework automatically generates pairing constraints by identifying the geometric features of garment opening structures (such as front plackets and side seams). By introducing the concept of "virtual seams," semi-rigid connections are used to simulate different seam strengths, and seam parameters are dynamically adjusted based on fabric properties. For closure structures such as buttons and zippers, the system establishes a parameterized closure model library that automatically selects the appropriate closure state based on detected human posture.
[0154] Furthermore, the clothing rendering information is used to adjust the rendering details of the virtual clothing rendered on the 3D human body model. For the specific adjustment process, refer to S304-S307.
[0155] S304: Adjust material rendering parameters of the virtual clothing on the 3D human body model based on the environment information in the clothing rendering information.
[0156] In the embodiments provided in the present application, environmental information includes but is not limited to environmental light information, light source direction, intensity, color temperature, etc. The present application can use the Environmental Aware Dynamic Material System (EADMS) to dynamically adjust the material rendering parameters of virtual clothing based on environmental information, so that the virtual clothing can present correct visual effects in different lighting environments and enhance the sense of reality; and apply BRDF (bidirectional reflectance distribution function) technology to simulate the changing effects of materials such as silk and embroidery at different viewing angles, thereby solving the problem of unrealistic materials caused by changes in perspective in AR.
[0157] The material rendering parameters of virtual clothing include but are not limited to reflectivity parameters, which are used to control the degree of reflection of the material to light of different wavelengths; highlight parameters, which are used to enhance the highlight effect of materials such as silk in strong light environments; scattering parameters, which are used to adjust the scattering behavior of light inside the fabric, affecting the softness of materials such as wool; and shadow projection intensity parameters, which are used to adjust the shadow performance of wrinkles according to the ambient light intensity.
[0158] By adjusting the material rendering parameters of virtual clothing, the rendering of virtual clothing is made more natural, making the AR fitting experience more realistic and natural.
[0159] S305 , determining mesh attribute data of different regions of the virtual garment on the 3D human body model based on the motion prediction information and visual perception information in the garment rendering information, and adjusting the mesh resources and mesh density of the region based on the mesh attribute data of the different regions.
[0160] The action prediction information includes information on predicting the user's movement trajectory, such as predicting the user's movement trajectory such as raising his hand or turning around; the visual perception information includes but is not limited to the user's visual observation information of clothing during the fitting process, such as the user's attention to the collar and sleeves.
[0161] In the method provided in the embodiments of the present application, the virtual garment can be divided into multiple regions, each of which has different grid attribute data. For example, the deformation probability of each region is determined based on motion prediction information, and the deformation probability of the region is part of the network attribute information of the region. In another example, the visual salience of each region is determined based on visual perception information. The visual salience of a region is used to describe the user's visual attention to the region. The higher the visual salience, the higher the user's visual attention to the region, and vice versa. The visual salience of a region is part of the network attribute information of the region.
[0162] In this application, the dynamic interactive mesh optimization framework (DIMOF) is used to determine the mesh attribute data of different areas of virtual clothing on a 3D human body model based on motion prediction information and visual perception information, and adjust the mesh resources and mesh density of the area based on the mesh attribute data of different areas, that is, increase the mesh density in areas with high deformation probability, and reduce the mesh density in areas with low deformation probability, thereby reducing the computational burden; increase mesh resources in areas with high visual saliency, thereby retaining more details in areas with high line of sight saliency, and reduce mesh resources in areas with low visual saliency, thereby reducing the amount of computation under the same visual quality.
[0163] S306 : Adjusting the texture of the virtual clothing on the 3D human body model based on the augmented reality texture information in the clothing rendering information.
[0164] This application uses an Augmented Reality Texture Projection System (ARTPS) to adjust the texture of virtual clothing on a 3D human model based on augmented reality texture information. It automatically reallocates the UV space of the 3D human model as the human body moves, ensuring that textures do not stretch or compress unnaturally during large movements, reducing texture distortion. It also dynamically allocates texture resolution based on viewing distance and importance, reducing texture memory usage and improving rendering performance while maintaining visual quality.
[0165] The solution presented in this application intelligently generates physically accurate detail effects based on local human morphological features (such as shoulder contours and waist curves) and the physical properties of fabric. This technology transcends the limitations of traditional detail generation methods based on simple geometric transformations and can generate natural wrinkles, undulations, and shadows in real time on mobile devices with limited computing resources, enhancing realism.
[0166] S307 : Adjusting clothing physical simulation parameters of the virtual clothing on the 3D human body model based on the scene parameter information in the clothing rendering information.
[0167] In the embodiments provided in the present application, the clothing physical simulation parameters of the virtual clothing on the 3D human body model are adjusted based on the scene parameter information through the contextual intelligent physical parameter system (CIPPS); the scene parameter information includes but is not limited to the description information of the detected usage scenarios, such as shopping malls, homes, outdoor scenes, etc.; it also includes contact perception dynamic information, such as the contact description information between the human body and clothing, such as the contact description between the clothing and the user when the user sits down.
[0168] The adjusted garment physical simulation parameters include, but are not limited to, material-based parameters, surface interaction parameters, environmental response parameters, and local-specific parameters. Material-based parameters include, but are not limited to, density, which controls the garment's weight and droop; tensile stiffness, which controls the fabric's resistance to stretching; bending stiffness, which determines the ease with which wrinkles form in the garment; and damping coefficient, which influences the rate at which fabric vibrations decay. Surface interaction parameters include, but are not limited to, the static friction coefficient, which controls the garment's adhesion to the human body or other objects; the kinetic friction coefficient, which affects the ease with which the garment slides on the human surface; and the collision rebound coefficient, which determines the extent to which the garment rebounds after a collision. Environmental response parameters include, but are not limited to, the air resistance coefficient, which controls the air resistance experienced by the garment during movement; the gravity response factor, which adjusts the garment's response to gravity; and the wind influence factor, which determines the extent to which the garment flutters in the wind. Local-specific parameters include, but are not limited to, seam stiffness, which controls the physical properties of the garment's seams; wrinkle memory, which simulates the durability of wrinkles in certain fabrics (such as wrinkled fabrics); and shape retention, which simulates the shape-maintaining ability of structured garments like suits.
[0169] Dynamic adjustment of these parameters is the key to achieving a realistic fitting experience. By dynamically adjusting these parameters, the physical behavior of the clothing transitions naturally, responding to environmental changes just like real clothing, greatly enhancing the immersion and realism of virtual fitting.
[0170] The GDTS framework transforms static parameter settings into a dynamic, adaptive system by integrating five major systems: the Adaptive Intelligent Sizing System (AISS), the Environment-Aware Dynamic Material System (EADMS), the Dynamic Interactive Mesh Optimization Framework (DIMOF), the Augmented Reality Texture Projection System (ARTPS), and the Context-Intelligent Physical Parameter System (CIPPS). This improves overall system performance and user satisfaction compared to traditional methods. After being processed by the GDTS framework, the prepared virtual garment contains complete geometric, material, and physical information, providing the necessary data foundation for dynamic fitting and enhancing the realism and naturalness of the final fitting effect.
[0171] A dynamic fit algorithm is used to render virtual garments onto 3D human models, resulting in a more realistic rendering of the garments. The output of the dynamic fit algorithm is a preliminary fit state, which accounts for the geometric relationship between the garment and the human body but does not fully simulate the garment's physical dynamic behavior. This initial fit state serves as the starting point for the next physical simulation, further simulating the garment's dynamic effects under the influence of gravity and human motion.
[0172] Precise dynamic fitting is both the technical challenge and core competitiveness of AR fitting systems, directly determining the naturalness and realism of virtual fitting. An excellent fitting algorithm ensures that virtual clothing appears to be "worn" on the user, rather than simply "attached" to the user.
[0173] S308: Perform physical simulation on the 3D human body model including the virtual clothing using a preset physical simulation technology to obtain a virtual rendering video frame corresponding to the video frame.
[0174] In order to make the virtual fitting experience more realistic and natural, physical simulation technology is used to perform physical simulation on the 3D human body model containing the virtual clothing, so that the virtual clothing can naturally produce realistic dynamic effects as the human body moves.
[0175] In the embodiments provided in this application, physical simulation technology is implemented using an adaptive multi-resolution physical simulation framework, a hybrid solver architecture, a deep learning-based physical parameter automatic calibration system, and a physical boundary processing system unique to the AR environment.
[0176] Among them, an adaptive multi-resolution physical simulation framework is used to dynamically allocate computing resources for the 3D human model containing virtual clothing based on visual importance and deformation activity. For example, high-precision physical grids (≤5mm grid spacing) are used in the visually critical areas (such as collars, cuffs, and hems) and large deformation areas (such as joint bends) of the 3D human model containing virtual clothing; low-precision grids (≥15mm grid spacing) are used in the visually secondary areas (such as the back plane area) of the 3D human model containing virtual clothing. Real-time dynamic subdivision technology is also used to automatically adjust the grid density and physical computing intensity according to the human body's motion state. The system implements an automated importance assessment algorithm:
[0177] ;
[0178] The parameters are dynamically adjusted based on the current viewing angle, motion amplitude, and fabric properties. This method reduces computational complexity compared to traditional uniform grid methods while maintaining visual quality.
[0179] in, Score the importance of the region; is the visibility factor, is the deformation activity, and MaterialProperty is the material property value.
[0180] It's important to note that the visibility factor measures the degree to which a particular area of clothing is currently visible to the user. For example, if a user is trying on a jacket facing a mirror, the visibility factor for the front of the jacket will be high, while the visibility for the back will be low. The system tracks the user's viewing angle in real time and calculates the visibility probability of each area.
[0181] Deformation activity reflects the degree to which a garment area is deforming. For example, when a user bends their arm, noticeable wrinkles form at the elbow of the sleeve, causing the deformation activity of this area to surge. Meanwhile, the deformation activity of the relatively still torso is lower.
[0182] The Material Property value takes into account the fabric's inherent need for detailed simulation. Special materials like silk and wrinkled fabric require more detailed physical simulation to capture their characteristics, so their Material Property values are higher; while ordinary cotton fabrics have relatively lower values.
[0183] α, β, and γ are three weighting parameters that determine the relative importance of their respective parameters. The system dynamically adjusts the specific values of these three weighting parameters based on different scenarios. For example, in static display scenes (such as standing observation), α (visibility weight) will be increased to give more resources to the visual focus area; during dynamic try-ons (such as turning around on a catwalk), β (deformation weight) will be increased to ensure the dynamic effect is realistic; when showcasing clothing made of special fabrics, γ (material weight) will be increased to highlight the material characteristics.
[0184] By applying this technology, computing resources can be allocated rationally, and more computing resources can be allocated to areas of high importance, thereby improving the overall visual quality in the future. For example, when the user looks down at their shoes, the system will immediately increase the importance score of the trouser leg area, allocate more grids and computing resources there, and at the same time reduce the accuracy of the collar of the shirt that is temporarily invisible. This not only ensures the visual quality of key areas, but also greatly saves the overall computing power, so that the phone will not overheat and reduce the frequency due to complex calculations. In actual applications, the system will recalculate this score multiple times per second, so when the user's perspective changes or makes different actions, the computing resource allocation will be seamlessly adjusted to always maintain the best visual effect and performance balance.
[0185] The hybrid physics solver architecture combines the advantages of multiple physical models: an explicit integration solver optimized for mobile devices, specifically designed for mobile GPU architectures; position-based dynamics for large deformations at the structural layer; an energy-based minimization method for generating natural wrinkles at the detail layer; and pre-computed physical response lookup tables for accelerated processing at key points.
[0186] The hybrid solution process is expressed as:
[0187] ;
[0188] Where X represents the garment state vector, including information such as position and velocity. X(t+Δt) represents the garment state at the next moment, which can be understood as the new garment state after a small time step (Δt), including information such as the position and velocity of each vertex. GlobalSolver(X(t)) represents the global calculation result, that is, the calculation result of the global solver. LocalCorrection(X(t)) represents local detail correction. DetailEnhancement(X(t)) represents visual detail enhancement.
[0189] It's important to note that GlobalSolver(X(t)) handles the overall motion and large deformations of the garment, such as the swaying of clothing as the figure turns and the drooping of skirts due to gravity. It uses position-based dynamics (PBD), making it particularly well-suited for real-time calculations on mobile devices. Think of it like first using a large brush to outline the basic contours and movement of the garment. LocalCorrection(X(t)) focuses on areas requiring special attention, such as points of collision between the garment and the figure and interactions between garments. This step corrects for any interplay (clothing penetrating through the figure) or unnatural deformations that may occur during the global calculation. It's somewhat like a painter correcting errors and inconsistencies in the basic contours. DetailEnhancement(X(t)) generates small, yet visually significant details, such as the natural folds of fabric, the subtle wiggle of silk, and subtle variations in texture. It employs an energy-based approach, resulting in highly realistic visuals. It's like the fine brushstrokes a painter adds at the end—small but crucial to the final effect.
[0190] The advantage of this multi-level hybrid computing approach is that it can efficiently handle large-scale deformations (suitable for the performance limitations of mobile devices) while also presenting fine visual details (meeting users' demand for realism). For example, when a user wearing a silk shirt turns around, the system can simultaneously handle the overall flutter of the garment (global correction), the fit of the cuffs and wrists (local correction), and the subtle wrinkles and gloss changes in the fabric (detail enhancement).
[0191] The deep learning-based automatic physical parameter calibration system breaks through the limitations of traditional methods that rely on manual settings. The system uses a physical parameter recognition neural network (PhysicsNet) to automatically infer multiple key physical parameters from a single garment image. It contains a physical property database of more than 10,000 real garment samples, covering common materials such as cotton, silk, wool, and synthetic fibers. Through a bidirectional mapping model between physical parameters and visual appearance, it ensures that the simulation results are consistent with the actual garment performance.
[0192] The physical parameter extraction process is:
[0193] ;
[0194] Where I is a garment image (or a video frame containing a 3D human model of rendered virtual garments), M is a garment mesh, and C is the garment category information. These parameters include key physical properties such as elastic modulus, damping coefficient, friction coefficient, and mass distribution. The system can accurately simulate properties such as the fluidity of silk, the stiffness of denim, and the stretchability of knitted fabrics.
[0195] in, The 17 key physical parameters output include: (1) tensile elastic modulus for controlling the resistance of the fabric to stretching; (2) bending elastic modulus for controlling the stiffness of the fabric when bending; (3) shear elastic modulus for controlling the resistance of the fabric to shear deformation; (4) compression elastic modulus for controlling the reaction of the fabric when under pressure; (5) mass density for the mass of the fabric per unit area; (6) dynamic friction coefficient for controlling the friction characteristics of the fabric when moving; (7) static friction coefficient for controlling the friction characteristics of the fabric when it is stationary; (8) damping coefficient for controlling the vibration attenuation rate; (9) air resistance coefficient for affecting the resistance of the fabric to movement in the air; (10) Thickness parameter that controls the thickness characteristics of the fabric; (11) Anisotropy coefficient used to describe the difference in physical properties of the material in different directions; (12) Wrinkle formation threshold used to control the difficulty of wrinkle formation; (13) Resilience coefficient used to describe the ability of the material to recover its original shape after being compressed; (14) Drape coefficient used to control the shape of the fabric when it droops; (15) Gravity response coefficient used to adjust the sensitivity of the fabric to gravity; (16) Surface roughness used to affect visual effects and touch; (17) Thermal deformation coefficient used to reflect the deformation characteristics of the material at different temperatures; These parameters together constitute a complete fabric physical behavior model, enabling the system to accurately simulate the characteristics of various materials.
[0196] In view of the particularity of the augmented reality environment, this application includes an AR environment physical boundary processing system, which is used to detect and model physical boundaries in the real environment (such as desktops and walls) in real time, process the physical interaction between virtual clothing and real objects (such as the contact deformation of clothes and chairs when sitting down), and consider the visual impact of ambient lighting on the gravity deformation of clothing, and dynamically adjust the shadow and wrinkle performance.
[0197] The boundary constraint processing formula in this system is:
[0198] ;
[0199] Among them, v is the vertex of the clothing (that is, the vertex position of the virtual clothing mentioned above), that is, a small point on the virtual clothing, and b is the detected boundary point of the environment, such as the surface point of real objects such as chairs, tables, walls, etc. in the environment. is a set of boundary points, that is All the real object surface points that are identified by the system and may collide with the virtual clothing are saved in it.
[0200] ||v - b|| means calculating the distance between a clothing vertex v and a boundary point b. In other words, it means the spatial distance between a clothing vertex v and a boundary point b.
[0201] min() is a minimal operation. Here, we need to find the shortest distance from a vertex v to all boundary points. ∀ is the mathematical symbol for "for all," meaning that we need to calculate the distance for each point b in the boundary set and then take the minimum value.
[0202] Constraint(v) is the constraint force ultimately applied to the garment vertex v. Simply put, this formula calculates how far away each point on the garment is from the nearest object in the real world, and then applies appropriate constraints based on this distance to prevent the virtual garment from penetrating real objects.
[0203] By applying this technology, when a user sits on a real chair wearing a virtual dress, the dress will naturally flow across the seat, rather than strangely penetrating the chair. When approaching a real wall, the virtual coat will be squeezed and deformed, rather than passing through it. This innovation enables the system to take real-world environmental factors into account, creating a more realistic mixed reality experience.
[0204] The embodiments provided in this application also achieve efficient simulation of the following basic physical effects: Gravity effect simulation: The system accurately calculates the impact of gravity on different parts of clothing, producing a natural draping effect. For example, the hem of a dress will droop naturally, while the material fitted at the shoulders will remain in a relatively fixed position. Dynamic wrinkle generation: When human movement causes clothing to deform, the system can calculate and generate three typical wrinkles in real time: compression wrinkles (at the bend of joints), drape wrinkles (such as ripples on the hem of a skirt), and stretch wrinkles (the subtle tension texture produced when the fabric is stretched). Multi-layered clothing interaction: The system can handle the physical interaction between multiple layers of clothing, such as the superposition effect when a coat covers a shirt, ensuring that each layer of clothing has reasonable physical performance and does not penetrate each other. Human-clothing interaction: This includes friction calculation, collision response, and inertia effects. When the human body moves quickly, the clothing will produce a delayed follow-up effect due to inertia, enhancing dynamic realism.
[0205] By integrating these innovative technologies, the system's physics simulation module achieves efficient, high-quality dynamic clothing simulation on mobile devices, breaking the traditional trade-off between computing resource limitations and realistic performance. The system maintains a stable 60 FPS performance while delivering near-professional-level clothing physics, providing a technical foundation for AR fitting applications.
[0206] The output of physical simulation is a series of garment model states that exhibit natural dynamic effects. These states are synchronized with human movement and exhibit physical behavior similar to that of real clothing. This dynamic realism is the fundamental difference between AR fitting systems and static image synthesis methods, and is a key factor in enhancing user immersion and experience satisfaction.
[0207] S106: Generate an AR fitting video based on each virtual rendering video frame, and display the AR fitting video to the user.
[0208] AR rendering and fusion are key steps in seamlessly integrating virtual clothing into real videos. The goal is to create a mixed reality effect that is visually indistinguishable from real. This application uses advanced neural rendering technology combined with traditional graphics methods to achieve the following core functions:
[0209] (1) Lighting estimation and matching: The system analyzes the lighting conditions of the real environment from the video frames, including parameters such as light source direction, intensity, color temperature, and ambient light. These lighting parameters are applied to the virtual clothing rendering to ensure that the shadows, highlights, and overall brightness of the clothing are consistent with the real scene. For example, when the user stands in bright natural light, the virtual clothing will show the corresponding daylight effect; under indoor lighting, it will show an appropriate yellow tone and soft shadows.
[0210] (2) Realistic rendering of materials: Different clothing materials (such as cotton, silk, leather, wool, etc.) have unique visual characteristics. The system uses physically based rendering (PBR) technology to simulate a variety of material characteristics. Various material characteristics include but are not limited to: subsurface scattering, which is used to simulate the scattering effect of light after penetrating thin fabrics, such as translucent silk; anisotropic reflection, which is used to simulate the directional gloss of fabrics such as silk and satin; microsurface structure, which is used to simulate the subtle texture of rough fabrics (such as wool and denim); refraction and transparency, which are used to process the perspective effect of thin or mesh materials; edge fusion and occlusion processing, which are used by the system to accurately calculate the depth relationship between virtual clothing and real human bodies and solve complex occlusion problems. When solving complex occlusion problems, when clothing should be partially occluded by the human body (such as arms crossed in front of the chest), the system can correctly handle the occlusion relationship; at the edge of the clothing, a special edge fusion algorithm is applied to eliminate jagged edges and hard boundaries to create a natural transition effect; considering the complex rendering of translucent objects (such as tulle), the clothing itself and the partially occluded human body are displayed at the same time.
[0211] (3) Dynamic detail enhancement: The system uses the physical simulation data calculated in the previous steps to enhance the visual details of the clothing; the visual details include but are not limited to: (a) dynamic wrinkle texture, which generates realistic wrinkle texture maps on the clothing surface based on the deformation data of physical simulation; (b) stress visualization, where the stretched area of the fabric will show appropriate color changes or fabric texture deformation; (c) dynamic shadows, where the self-shadows and mutual shadows between the clothing and the human body will be updated in real time as the movement changes.
[0212] (4) Neural rendering enhancement: The system applies deep learning technology to further improve the rendering quality, which can be improved in the following aspects: (a) Style transfer: Ensure that the visual style of the rendered clothing is consistent with the video image; (b) Detail synthesis: The neural network automatically adds subtle defects and details commonly found in the real world to the rendering results; (c) Super resolution: Intelligently upsample the preliminary rendering results to increase the richness of details; (d) Temporal consistency: Ensure that the rendering effect remains stable between consecutive frames of the video to avoid flickering or jitter; (e) Real-time performance optimization: In order to achieve a smooth experience on mobile devices, the system adopts a multi-level rendering strategy; The multi-level rendering strategy includes: using higher precision rendering for visually important areas (such as the chest and shoulders), using simplified rendering for secondary areas (such as the back and less noticeable areas), and using temporal reprojection technology to reuse the calculation results of the previous frames to reduce the computational burden.
[0213] The output of AR rendering and fusion is a series of visually realistic synthetic video frames. The virtual clothing in these frames blends naturally with the real video scene, giving users the intuitive feeling of "this is me wearing this clothing." This seamless visual effect is a key indicator of the success of AR fitting systems and a decisive factor in user acceptance of this technology.
[0214] In the embodiment provided in the present application, a multimodal interaction processing module is provided. After displaying the AR fitting video to the user, the module can also collect the user's multimodal clothing adjustment information, where the multimodal clothing adjustment information includes at least one of clothing local adjustment information and clothing touch information; the multimodal clothing adjustment information is applied to adjust the rendering state of the virtual clothing on the 3D human body model in the AR fitting video.
[0215] The local adjustment information in the multimodal clothing adjustment information can be the adjustment information input by the user's voice. The multimodal interaction processing module of the present application includes the use of a semantic understanding framework to identify the adjustment information input by the user's voice. The framework includes a domain-specific language (DSL) processing engine that supports understanding and executing professional clothing adjustment terms such as "waist tightening", "adjusting drape", and "strengthening shoulder lines". The multimodal interaction processing module also includes a clothing concept parsing system based on a knowledge graph, which automatically converts abstract semantics into precise three-dimensional mesh deformation parameters. Exemplarily, the semantic mapping algorithm is expressed as: geometric deformation operation of the clothing mesh = semantic converter (user command phrase, clothing type, human body context).
[0216] "User command phrases" can range from professional clothing terms like "tighten the waist," "strengthen the shoulder line," or "increase drape" to everyday expressions like "slimmer fit" or "open the neckline a little wider." The system understands the actual intent of these natural language expressions. "Garment type" refers to whether the person being tried on is a jacket, shirt, or dress. The same "tighten" command on a suit might mean adjusting the shoulder line, while on a skirt it might mean tightening the waist. "Human context" considers the user's body shape and posture to ensure that the geometric deformation is appropriate for the individual being tried on. For example, for a body type with round shoulders, the specific implementation of "strengthen the shoulder line" will differ from that for a standard body type.
[0217] This semantic mapping algorithm precisely quantifies abstract linguistic concepts into specific 3D mesh deformation parameters—which vertices to move, how far, and in what direction. This allows the system to not simply widen the neckline when a user asks, "Make the collar a little wider!" Instead, it considers the collar structure, fabric properties, and neck shape, making adjustments with the precision of a professional tailor. This allows users to interact with the system in the most natural way possible, without having to learn complex technical operations. They can adjust virtual garments with the same ease as communicating with a real tailor.
[0218] The system, trained on a corpus of over 6,000 professional clothing adjustment dialogues, achieves a high level of understanding of specialized clothing terminology. The system can process complex semantic commands, such as "slightly tighten the waist but keep the chest loose," by breaking them down into multiple coordinated mesh deformation operations while maintaining the overall aesthetics of the garment. This model significantly improves the professionalism and accuracy of the virtual fitting process.
[0219] The multimodal interaction processing module also includes a tactile feedback system based on material properties. This system integrates tactile feedback into the AR fitting experience, breaking through the limitations of traditional purely visual interaction: This application develops a material property tactile mapping framework to convert the physical properties of fabrics into perceptible tactile feedback patterns; the tactile parameter calculation formula is:
[0220] HapticProfile(t) = Σ[Wi·Pi(materialProperties)];
[0221] Among them, Pi represents different tactile characteristics, Wi is the dynamic weight coefficient, HapticProfile(t) is the tactile feedback feature (t), and materialProperties is the material property;
[0222] The system has built a tactile feature database containing 21 common fabrics, which can simulate a variety of textures from the smoothness of silk to the roughness of denim, and realize force perception interaction. The amount of pressure applied by the user will affect the intensity of the tactile feedback and the degree of deformation of the clothing. The tactile feedback works in conjunction with the physical simulation system to provide differentiated feedback when the user "touches" special structures such as wrinkles or seams. The system generates fine tactile patterns through the vibration motor of the mobile phone. Experiments show that users can identify different fabric types with an accuracy of 78% based on tactile feedback alone, significantly enhancing the immersion and realism of virtual fitting.
[0223] The multimodal interaction processing module also includes a progressive interaction system for intent prediction, which implements forward-looking progressive interaction based on intent prediction. For example, it uses a temporal convolutional network (TCN) to analyze user interaction sequences in real time and predict the next possible operation intention.
[0224] Prediction model structure:
[0225] IntentProbability = SoftMax(TCN(InteractionSequence <t-n:t>));
[0226] The Chinese expression of the above formula is: Intent probability = Normalization function (Temporal Convolutional Network (Recent Interaction Sequence)).
[0227] It's important to note that the system records the user's recent action sequences, which form the "recent interaction sequence tn:t" section. For example, a user might first adjust their neckline, then look at the side effects, and then adjust their waist. These consecutive actions are recorded to form a time series. This action sequence is then fed into a temporal convolutional network (TCN) for analysis. This network is particularly adept at discovering patterns in time series, identifying user habits and tendencies based on behavior over the past few minutes. After analysis, the network outputs a set of raw values representing the initial probability of the user performing various subsequent actions. However, these raw values require further processing to be meaningful. A "Softmax" function then converts these raw values into a probability distribution, so that the sum of the probabilities of all possible actions is 100%. For example, the system might determine that the probability of the user adjusting their neckline next is 60%, adjusting their sleeve length is 30%, and the probability of any other action is 10%. The resulting "intention probability" is a ranked list of possible actions, and the system prioritizes high-probability actions.
[0228] This application preloads resources based on predictions and precalculates possible deformations, reducing perceived latency from the standard 120ms to 38ms. By developing an interactive fluency optimization engine, discrete operation commands are transformed into a continuous clothing adjustment process. A "partial confirmation" interaction mode is implemented, allowing the system to begin responding before the user completes the entire gesture, significantly improving interaction fluency. The system can recognize and learn user-specific operation patterns, such as the habitual sequence of "adjusting the neckline first, then the sleeve length," and proactively optimize the UI layout to facilitate subsequent operations. Experimental data shows that progressive interaction reduces the time it takes to complete common clothing adjustment tasks.
[0229] The multimodal interaction processing module of this application is a multimodal signal deep fusion framework. Through a multi-level signal integration architecture, gesture, voice, sight and expression data are integrated at the feature level. Through the signal synergy enhancement algorithm, different modalities complement and verify each other to improve the overall recognition accuracy. Through the conflict resolution mechanism, when there is a contradiction between different modal inputs, the system will dynamically assign weights based on historical accuracy and current confidence. Through cross-modal intention Figure 1 Consistency verification ensures that the system understands the user's intent accurately and coherently.
[0230] For example, if a user gazes at the neckline, gestures downward, and says "lower," the system understands that the user wants to lower the neckline, not the entire garment. This deep fusion allows the system to correctly understand user intent even in noisy environments or when gestures are ambiguous.
[0231] By integrating these technologies, the system's multimodal interaction processing module achieves an unprecedentedly natural, precise, and personalized interactive experience in AR virtual fitting scenarios. Compared to traditional methods, this results in improved user satisfaction, higher first-time success rates, and lower error rates. These technologies not only enhance the system's usability but also significantly increase the immersiveness and professionalism of virtual fitting, bringing the virtual fitting experience closer to, or even surpassing, that of physical fitting.
[0232] In the method provided in the embodiment of the present application, when the AR fitting video is displayed to the user, the following processing is also performed:
[0233] (1) Video stream optimization: The system comprehensively optimizes the video frame sequence after AR rendering to ensure the final presentation quality. The optimization aspects include: (a) Frame rate matching: Ensure that the frame rate of the output video is consistent with the original video, usually 24-30 frames per second, to avoid unnatural motion speed; (b) Temporal domain smoothing: Apply temporal filtering algorithms to eliminate inter-frame jitter and flicker, such as Temporal Anti-aliasing (TAA) technology, to improve video stability; (c) Color calibration: Adjust the color output according to the characteristics of the user device display to ensure accurate color restoration; (d) Encoding optimization: Select appropriate video encoding parameters to optimize file size and transmission efficiency while ensuring visual quality.
[0234] (2) Detail optimization processing, including:
[0235] (a) Edge refinement: Apply special edge detection and optimization algorithms to focus on the edge areas where clothing contacts the human body to eliminate possible artifacts; (b) Stability enhancement: Apply additional stability processing to areas that may be unstable (such as fluttering fabrics) to avoid unrealistic fast jitters; (c) Visual consistency correction: Ensure that the color, brightness and texture of the clothing remain consistent throughout the video and do not change suddenly due to scene changes; (d) User interface integration: The system integrates AR fitting content with intuitive user interface elements, including but not limited to: Real-time control options: Display clothing adjustment controls at appropriate locations on the screen, such as color selection and size fine-tuning buttons; Information overlay: Overlay clothing details, prices, materials and other information based on user needs; Split-screen comparison: Supports displaying comparison views of different clothing or different colors of the same clothing on the same screen; Gesture prompts: Provide visual gesture guidance when it is detected that the user has an intention to interact but the operation is inaccurate; (e) Multi-terminal adaptation: The system makes adaptive adjustments based on the characteristics of different display devices, including but not limited to: Mobile phone portrait mode: Optimize (f) Result saving and sharing: the system provides multiple ways to record and share the fitting experience, including but not limited to: video recording: saving the complete fitting process video, including user interaction and action display; selected frame screenshot: automatically or manually capturing static images of the best fitting effect; social media sharing: one-click sharing function, supporting direct publishing to mainstream social platforms; (g) Shopping integration: seamless integration with e-commerce platforms, supporting adding to shopping carts or placing orders directly from fitting results; (h) Data collection and analysis: with the user's permission, the system collects anonymous usage data for improving clothing recommendations and system optimization. The optimization content includes but is not limited to: recording the user's trying time and interaction level on different clothing; analyzing the user's facial feedback and clothing adjustment preferences; and providing more accurate personalized recommendations for subsequent users based on the collected data.
[0236] By transforming AR fitting technology into an intuitive and easy-to-use user experience, users can gain an unprecedented virtual fitting experience on familiar devices. Perfect output processing and user-friendly interface design are key factors in ensuring that technological innovation can be transformed into actual commercial value.
[0237] In the embodiment provided by the present application, a target video containing a user to be tried on is processed to obtain the human posture data of each video frame in the target video; the human posture data of each video frame is used to construct a 3D human body model for each video frame; for the 3D human body model of each video frame, the clothing rendering information of the 3D human body model is obtained, and based on the dynamic fitting algorithm and the clothing rendering information, a virtual garment is rendered on the 3D human body model to obtain a virtual rendered video frame of the video frame; an AR fitting video is generated based on each virtual rendered video frame, and the AR fitting video is displayed to the user. By applying this application, the selected clothing can be tried on in the form of AR, and the fitting effect of the clothing can be displayed to the user, effectively solving the problems of inappropriate size and mismatched style caused by the inability to try on clothing purchased online, reducing the probability of consumer returns and exchanges, and providing users with a good shopping experience.
[0238] Reference Figure 4 , which is a flow chart of another virtual fitting method provided in an embodiment of the present application, is specifically described as follows:
[0239] The input video is obtained and pre-processed, and human posture detection and human posture tracking are performed on the pre-processed video to obtain human posture data of each video frame, and a 3D human body model of each video frame is generated based on the human posture data of each video frame; the virtual clothing of the target clothing selected by the user is determined, and the virtual clothing is rendered onto the 3D human body model of each video frame using a dynamic fitting algorithm, and the virtual clothing on the 3D human body model is simulated using physical simulation technology to obtain an AR fitting video. Furthermore, when the AR fitting video is displayed to the user, when multimodal interaction information is received, the rendering state of the virtual clothing in the AR fitting video is adjusted based on the multimodal interaction information, and the adjusted AR fitting video is displayed to the user.
[0240] This application uses dynamic video processing and real-time physical simulation technology to fundamentally overcome the limitation of traditional virtual fitting, which can only process static images or preset movements. It achieves a natural fitting effect under arbitrary movements and significantly improves the applicability of the technology. By combining neural rendering with traditional graphics methods, it achieves a seamless integration of virtual clothing and real video. The lighting effects and wrinkle changes of the clothing are highly consistent with the ambient lighting, solving the "texture feeling" problem common in existing AR fitting. The dedicated fabric physics engine developed by this application can run efficiently on terminals such as mobile devices while maintaining realistic physical effects. The computational efficiency is more than 300% higher than that of general physics engines, making the real fitting experience possible for the first time on ordinary consumer-grade devices. The 3D human body reconstruction algorithm of this application can accurately infer the user's full body size from a single video, with a measurement accuracy of ±1.5cm, far better than the average error of ±3-5cm of existing technologies, laying the foundation for precise clothing fitting.
[0241] With this application, users can freely move around in the video and view the effects of clothing in real time. The system supports dynamic display of clothing during various daily movements such as rotation, walking, and sitting, filling the gap in static fitting that cannot evaluate the comfort and dynamic aesthetics of clothing, allowing users to immerse themselves in dynamic fitting. During the fitting process, multimodal interaction is natural and intuitive. By integrating multiple interaction methods such as gestures, voice, and expressions, users can adjust clothing through intuitive operations such as "stretching" and "pinching", with an interaction success rate of over 95%, which is 40% higher than the traditional button-based interface. A full range of appearance assessments can also be performed. This application supports 360-degree viewing, allowing users to observe the fitting effect from different angles, overcoming the limitations of traditional fitting mirrors that can only observe limited viewing angles, and providing a more comprehensive evaluation of the wearing effect. In addition, this application provides personalized fitting recommendations. Based on the user's body shape and wearing effect analysis, the system can intelligently recommend more suitable sizes and styles, with an accuracy rate 35% higher than traditional size charts, effectively reducing the trouble of returns and exchanges due to inappropriate sizes.
[0242] By applying the solution provided in this application, consumers can more accurately evaluate the fit of clothing before purchasing through accurate virtual fitting effects. Actual measurements show that the return and exchange rate of clothing e-commerce can be reduced from an average of 30% to below 15%, effectively reducing the return and exchange rate, directly reducing merchants' losses and logistics costs, and improving conversion rates. That is, the interactivity and fun of virtual fitting significantly increase user retention time and purchase intention. Pilot data show that after deploying this system, the conversion rate of clothing categories on e-commerce platforms increased by an average of 28%, far higher than the industry average growth level. It can also reduce operating costs, that is, reduce the demand for fitting space and sample inventory in physical stores. The area of a single store can be reduced by 20% while providing a richer selection of fitting options. At the same time, it significantly reduces clothing loss and labor costs caused by fitting, and provides multi-channel application capabilities. The system is suitable for various scenarios such as online e-commerce, offline smart mirrors, mobile apps, etc., supports omni-channel marketing strategies, provides brands with a unified virtual fitting experience, and enhances brand image consistency.
[0243] Although the present invention depicts each operation in a specific order, this should not be understood as requiring that these operations be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0244] and Figure 1 Correspondingly, the embodiment of the present application provides a virtual fitting device, which is used to support Figure 1 The implementation of the method shown, refer to Figure 5 , is a structural diagram of a virtual fitting device provided in an embodiment of the present application, and is specifically described as follows:
[0245] The first acquisition unit 501 is configured to acquire a target video containing a user to try on clothes;
[0246] A second acquiring unit 502 is configured to perform human posture detection and human posture tracking processing on the target video to acquire human posture data of the user in each video frame of the target video;
[0247] A construction unit 503 is configured to determine, for each video frame, human body model parameters in multiple dimensions based on the human body posture data of the video frame, and construct a 3D human body model of the user in the video frame using the human body model parameters;
[0248] The third acquisition unit 504 is configured to acquire a virtual garment of a target garment from a preset garment model library; the target garment is the garment to be tried on selected by the user;
[0249] A rendering unit 505 is configured to obtain clothing rendering information of the 3D human body model for each video frame, and render the virtual clothing on the 3D human body model based on a preset dynamic fitting algorithm and the clothing rendering information to obtain a virtual rendered video frame corresponding to the video frame;
[0250] The display unit 506 is configured to generate an AR fitting video based on each of the virtual rendering video frames, and display the AR fitting video to the user.
[0251] In another embodiment provided in the present application, the device further includes: an acquisition unit; the acquisition unit is used to acquire multimodal clothing adjustment information of the user, wherein the multimodal clothing adjustment information includes at least one of clothing local adjustment information and clothing touch information; and the multimodal clothing adjustment information is applied to adjust the rendering state of the virtual clothing on the 3D human body model in the AR fitting video.
[0252] In another embodiment provided by the present application, the first acquisition unit 501 of the device performs the process of acquiring the target video containing the user to be fitted with clothes, including:
[0253] Get an initial video of the user who wants to try on clothes;
[0254] Determining scene complexity in the initial video, and adjusting filter parameters of a preset denoising filter based on the scene complexity;
[0255] Processing the initial video using a denoising filter with adjusted parameters to obtain a first video;
[0256] Correcting the main color tone of the scene in the first video using a preset color correction strategy to obtain a corrected second video;
[0257] Using a preset resolution optimization strategy, performing resolution optimization processing on the second video to obtain a third video;
[0258] The third video is subjected to frame rate stabilization processing using a preset frame rate stabilization strategy, and the processed video is subjected to fitting enhancement processing based on a preset fitting enhancement strategy to obtain a target video.
[0259] In another embodiment provided by the present application, the second acquisition unit 502 of the device performs the human posture detection and human posture tracking processing on the target video to obtain the human posture data of the user in each video frame of the target video, including:
[0260] For each video frame of the target video, a preset posture detection module is used to perform human key point detection on the user's human body image to obtain the human skeleton data of the video frame, and a preset multi-person separation network is used to perform human instance segmentation on the video frame to determine the human body instance of the user in the video frame, and based on the human body instance, the skeleton key point data of the user is determined in the human skeleton data; the correlation between the user in the video frame and the adjacent video frames is established, and the motion trajectory of each skeleton key point of the user between the video frame and the adjacent video frames is smoothed to obtain the human posture data of the video frame.
[0261] In another embodiment provided by the present application, the construction unit 503 of the apparatus performs the process of determining the human body model parameters of multiple dimensions based on the human body posture data of each video frame, including:
[0262] For each of the video frames, based on various human body indicators of a preset deformable human body model, extracting an indicator parameter of each of the human body indicators from the human body posture data;
[0263] Extracting various detail description parameters from the video frame;
[0264] The index parameters and the detail description parameters are determined as human body model parameters.
[0265] In another embodiment provided by the present application, the third acquisition unit 504 of the device executes the process of acquiring the virtual clothing of the target clothing from the preset clothing model library, including:
[0266] Collect clothing data of target clothing;
[0267] Based on the clothing data, a virtual clothing of the target clothing is determined in the clothing model library.
[0268] In another embodiment provided by the present application, the rendering unit 505 of the device executes the preset dynamic fitting algorithm and the clothing rendering information to render the virtual clothing on the 3D human body model to obtain a virtual rendered video frame corresponding to the video frame, including:
[0269] determining a clothing category of the virtual clothing based on clothing size information of the virtual clothing in the clothing rendering information, and rendering the virtual clothing at a position corresponding to the clothing category in the 3D human body model according to the clothing size information;
[0270] Constructing a signed distance field of the 3D human body model after rendering the virtual clothing, wherein the signed distance field is used to describe the spatial relationship between the virtual clothing and the 3D human body model;
[0271] Adjusting the rendering of the virtual garment on the 3D human body model based on a preset optimization strategy and the signed distance field until a preset convergence condition is met, and then stopping adjusting the virtual garment;
[0272] Applying the clothing rendering information to adjust rendering details of the virtual clothing rendered on the 3D human body model;
[0273] A preset physical simulation technology is used to perform physical simulation on the 3D human body model including the virtual clothing to obtain a virtual rendering video frame corresponding to the video frame.
[0274] In another embodiment provided by the present application, the rendering unit 505 of the apparatus performs the step of applying the clothing rendering information to adjust the rendering details of the virtual clothing rendered on the 3D human body model, including:
[0275] Adjusting material rendering parameters of the virtual clothing on the 3D human body model based on the environmental information in the clothing rendering information;
[0276] determining mesh attribute data of different regions of the virtual garment on the 3D human body model based on the motion prediction information and the visual perception information in the garment rendering information, and adjusting mesh resources and mesh density of the regions based on the mesh attribute data of the different regions;
[0277] Adjusting the texture of the virtual clothing on the 3D human body model based on the augmented reality texture information in the clothing rendering information;
[0278] Based on the scene parameter information in the clothing rendering information, clothing physical simulation parameters of the virtual clothing on the 3D human body model are adjusted.
[0279] An embodiment of the present invention further provides a storage medium, which includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the above-mentioned virtual fitting method.
[0280] The embodiment of the present invention further provides an electronic device, the structural diagram of which is shown in FIG. Figure 6 As shown, it specifically includes a memory 601 and one or more instructions 602, wherein the one or more instructions 602 are stored in the memory 601 and are configured to be executed by one or more processors 603 to perform the above-mentioned virtual fitting method.
[0281] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of the relevant regions.
[0282] The specific implementation processes and derivative methods of the above embodiments are all within the protection scope of the present invention.
[0283] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.
[0284] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0285] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A virtual fitting method, characterized in that: include: Obtain a target video containing a user to try on clothes; Performing human posture detection and human posture tracking processing on the target video to obtain human posture data of the user in each video frame of the target video; For each of the video frames, determining human body model parameters in multiple dimensions based on the human body posture data of the video frame, and constructing a 3D human body model of the user in the video frame using the respective human body model parameters; Obtaining a virtual garment of a target garment from a preset garment model library; the target garment is the garment to be tried on selected by the user; For each 3D human body model in the video frame, obtaining clothing rendering information of the 3D human body model, and rendering the virtual clothing on the 3D human body model based on a preset dynamic fitting algorithm and the clothing rendering information, to obtain a virtual rendered video frame corresponding to the video frame; generating an AR fitting video based on each of the virtual rendered video frames, and displaying the AR fitting video to the user; The step of rendering the virtual garment on the 3D human body model based on a preset dynamic fitting algorithm and the garment rendering information to obtain a virtual rendered video frame corresponding to the video frame includes: determining a clothing category of the virtual clothing based on clothing size information of the virtual clothing in the clothing rendering information, and rendering the virtual clothing at a position corresponding to the clothing category in the 3D human body model according to the clothing size information; Constructing a signed distance field of the 3D human body model after rendering the virtual garment, wherein the signed distance field is used to describe the spatial relationship between the virtual garment and the 3D human body model, and the signed distance field is a three-dimensional spatial function used to describe the shortest distance from each vertex position of the virtual garment to the surface of the 3D human body model in three-dimensional space; Adjusting the rendering of the virtual garment on the 3D human body model based on a preset optimization strategy and the signed distance field until a preset convergence condition is met, and then stopping adjusting the virtual garment; Applying the clothing rendering information to adjust rendering details of the virtual clothing rendered on the 3D human body model, the clothing rendering information including clothing size information, environment information, motion prediction information, visual perception information, augmented reality texture information, and scene parameter information, the motion prediction information including information predicting the user's motion trajectory, and the visual perception information including information on the user's visual observation of the clothing during the fitting process; A preset physical simulation technology is used to perform physical simulation on the 3D human body model including the virtual clothing to obtain a virtual rendering video frame corresponding to the video frame.
2. The method according to claim 1, characterized in that Also includes: Collecting multimodal clothing adjustment information of the user, wherein the multimodal clothing adjustment information includes at least one of clothing local adjustment information and clothing touch information; The multimodal clothing adjustment information is applied to adjust the rendering state of the virtual clothing on the 3D human body model in the AR fitting video.
3. The method according to claim 1, characterized in that The step of obtaining a target video containing a user to try on clothes includes: Get an initial video of the user who wants to try on clothes; Determining scene complexity in the initial video, and adjusting filter parameters of a preset denoising filter based on the scene complexity; Processing the initial video using a denoising filter with adjusted parameters to obtain a first video; Correcting the main color tone of the scene in the first video using a preset color correction strategy to obtain a corrected second video; Using a preset resolution optimization strategy, performing resolution optimization processing on the second video to obtain a third video; The third video is subjected to frame rate stabilization processing using a preset frame rate stabilization strategy, and the processed video is subjected to fitting enhancement processing based on a preset fitting enhancement strategy to obtain a target video.
4. The method according to claim 1, wherein The performing of human posture detection and human posture tracking processing on the target video to obtain human posture data of the user in each video frame of the target video includes: For each video frame of the target video, a preset posture detection module is used to perform human key point detection on the user's human body image to obtain the human skeleton data of the video frame, and a preset multi-person separation network is used to perform human instance segmentation on the video frame to determine the human body instance of the user in the video frame, and based on the human body instance, the skeleton key point data of the user is determined in the human skeleton data; the correlation between the user in the video frame and the adjacent video frames is established, and the motion trajectory of each skeleton key point of the user between the video frame and the adjacent video frames is smoothed to obtain the human posture data of the video frame.
5. The method according to claim 1, wherein For each of the video frames, determining human body model parameters in multiple dimensions based on human body posture data of the video frame includes: For each of the video frames, based on various human body indicators of a preset deformable human body model, extracting an indicator parameter of each of the human body indicators from the human body posture data; Extracting various detail description parameters from the video frame; The index parameters and the detail description parameters are determined as human body model parameters.
6. The method according to claim 1, characterized in that The applying the clothing rendering information to adjust rendering details of the virtual clothing rendered on the 3D human body model includes: Adjusting material rendering parameters of the virtual clothing on the 3D human body model based on the environmental information in the clothing rendering information; determining mesh attribute data of different regions of the virtual garment on the 3D human body model based on the motion prediction information and the visual perception information in the garment rendering information, and adjusting mesh resources and mesh density of the regions based on the mesh attribute data of the different regions; Adjusting the texture of the virtual clothing on the 3D human body model based on the augmented reality texture information in the clothing rendering information; Based on the scene parameter information in the clothing rendering information, clothing physical simulation parameters of the virtual clothing on the 3D human body model are adjusted.
7. A virtual fitting device, characterized in that: include: A first acquisition unit is configured to acquire a target video containing a user to be fitted with clothing; A second acquisition unit is configured to perform human posture detection and human posture tracking processing on the target video to obtain human posture data of the user in each video frame of the target video; a construction unit, configured to determine, for each video frame, human body model parameters in multiple dimensions based on human body posture data of the video frame, and construct a 3D human body model of the user in the video frame using the human body model parameters; A third acquiring unit is configured to acquire a virtual garment of a target garment from a preset garment model library; the target garment is the garment to be tried on selected by the user; a rendering unit configured to obtain, for each of the 3D human body models in the video frame, clothing rendering information of the 3D human body model, and render the virtual clothing on the 3D human body model based on a preset dynamic fitting algorithm and the clothing rendering information, to obtain a virtual rendered video frame corresponding to the video frame; a display unit, configured to generate an AR fitting video based on each of the virtual rendering video frames, and display the AR fitting video to the user; The rendering unit is specifically configured to: determining a clothing category of the virtual clothing based on clothing size information of the virtual clothing in the clothing rendering information, and rendering the virtual clothing at a position corresponding to the clothing category in the 3D human body model according to the clothing size information; Constructing a signed distance field of the 3D human body model after rendering the virtual garment, wherein the signed distance field is used to describe the spatial relationship between the virtual garment and the 3D human body model, and the signed distance field is a three-dimensional spatial function used to describe the shortest distance from each vertex position of the virtual garment to the surface of the 3D human body model in three-dimensional space; Adjusting the rendering of the virtual garment on the 3D human body model based on a preset optimization strategy and the signed distance field until a preset convergence condition is met, and then stopping adjusting the virtual garment; Applying the clothing rendering information to adjust rendering details of the virtual clothing rendered on the 3D human body model, the clothing rendering information including clothing size information, environment information, motion prediction information, visual perception information, augmented reality texture information, and scene parameter information, the motion prediction information including information predicting the user's motion trajectory, and the visual perception information including information on the user's visual observation of the clothing during the fitting process; A preset physical simulation technology is used to perform physical simulation on the 3D human body model including the virtual clothing to obtain a virtual rendering video frame corresponding to the video frame.
8. A storage medium, characterized in that: The storage medium includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the virtual fitting method according to any one of claims 1 to 6.
9. An electronic device, characterized in that: The device comprises a memory and one or more instructions, wherein the one or more instructions are stored in the memory and configured to be executed by one or more processors to implement the virtual fitting method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Virtual fitting method and device, electronic equipment and medium
CN113129450A
Display device and virtual fitting system and method
CN116523579A
Method and system for processing character motion trail in monocular video motion capture
CN118486074A