Immersive wearing-free motion capture digital human body feeling interaction system
Through the dual-computing motherboard architecture and dynamic timeout mechanism optimization algorithm, the problems of motion capture data loss and special posture recognition in wearable motion capture system are solved, and the digital human body interaction effect with low latency and high naturalness are achieved.
Patent Information
- Application Number
- CN202510604936.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing wearable motion capture digital human body-sensing interactive system is prone to problems in hardware performance, software settings and network stability, resulting in frame drops, lags, and incoherent movements, making it difficult to accurately identify special postures and actions, affecting the interaction effect.
The dual data checksum dynamic timeout mechanism is adopted, combined with static attitude detection and dynamic timeout processing, and the dual computing motherboard architecture optimization algorithm is used to realize the reliability and smooth processing of motion capture data, support the deployment of multiple devices, and use bone control and attitude analysis modules to ensure coherence and natural movements.
With the low hardware computing power requirement, low latency and high naturalness digital human body-sensory interaction is achieved, solving the problems of motion capture data loss and special posture recognition, ensuring the consistency and nature of the movement.
Smart Images

Figure CN120508208A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction technology, and in particular to an immersive wearable motion capture digital human body sensory interaction system. Background Art
[0002] As the wave of digitalization sweeps across the globe, interactive experiences with virtual humans are becoming increasingly popular across various fields. For a long time, virtual human actuation relied primarily on wearable motion capture devices, requiring users to wear sensors or markers to capture their movements. However, with continuous technological breakthroughs, a new, wearable, motion-capture-free digital human real-time actuation solution has emerged. Based on AI image recognition technology, this solution uses an RGB monocular camera to capture motion and transmits the motion data in real time to a 3D virtual human real-time interaction platform for real-time rendering. This ensures that the virtual human is synchronized with real-life movements, enabling smooth, real-time actuation and interaction.
[0003] The wearable motion-capture digital human solution is not only a technological innovation but also a new milestone in the user experience of virtual-reality interaction. As the technology continues to mature and its application scenarios expand, the wearable motion-capture digital human solution will flourish in even more fields, from immersive cultural tourism experiences to engaging community wellness exercises; from the stunning performance of commercial brands to the ultimate entertainment of somatosensory gaming, it will provide users with unprecedented surprises and enjoyment through a more natural, fluid, and personalized interactive approach.
[0004] However, in actual use, since motion capture data will be affected by factors such as hardware performance, software settings, and network stability, insufficient hardware performance, improper software settings, or unstable network may cause frame drops in the motion capture data, resulting in problems such as stuttering and discontinuous and unnatural movements in the digital human body interaction, affecting the digital human body interaction effect. In addition, for certain special human postures and movements, existing technologies are often difficult to accurately identify and respond to, resulting in delays and errors in the interaction. Summary of the Invention
[0005] The purpose of the present invention is to provide an immersive wearable motion capture digital human body sensory interaction system to solve the problems raised in the above background technology.
[0006] To achieve the above-mentioned object, the present invention provides the following technical solutions: an immersive wearable motion capture digital human body sensory interaction system, comprising a scene simulation layer, a character simulation layer, and an action simulation layer;
[0007] The scene simulation layer is responsible for constructing and simulating the entire simulation environment or background, providing a realistic background for subsequent characters and actions, making the simulation more realistic and immersive. It includes the following modules:
[0008] Terrain simulation module: responsible for restoring the real terrain, using oblique photography technology to obtain the real terrain and surrounding environment, and obtaining map data (such as vector slices and terrain elevation) through the Mapbox API, and building the terrain in combination with UE5's Landscape system;
[0009] Terrain Processing Module: Leveraging the Quixel Megascans asset library, we integrate millions of scanned real-world materials (such as rocks and soil) and apply them directly to terrain. Combined with custom masks (such as moss distribution and weathering effects), we achieve high-precision surface details. Specifically, we use node-based tools to define rules (such as vegetation distribution and rock placement) to automatically generate large-scale terrain.
[0010] Shadow processing module: Provides delicate shadow transitions through virtual shadow mapping technology, achieving consistent light and shadow at near and far distances;
[0011] Dynamic Weather Module: Combines UE5's Niagara particle system and volumetric fog to simulate the effects of rain, snow, sandstorms and other weather on the terrain, enhancing the sense of immersion.
[0012] The character simulation layer is responsible for simulating the characters in the simulation and realizing a virtual human model with film-level realism. It includes the following modules:
[0013] Character simulation module: uses MetaHumans' preset universal character templates as a starting point for creation, and supports exporting to DCC software such as Maya for in-depth customization;
[0014] Hair simulation module: Use MayaXGenGroom tool to realize complex hairstyles and use linear wire deformer to drive hair dynamic simulation;
[0015] The action simulation layer is responsible for simulating the specific actions of the character, including action capture, editing, synthesis, and rendering, ensuring that the character's actions are smooth and natural, and match the environment and tasks. It is also responsible for handling the interaction between the character and the environment, such as collision detection and physical feedback. It includes the following modules:
[0016] Motion capture module: responsible for detecting human posture and gesture key points, providing a flexible framework for processing and transmitting human posture and gesture key point data, and supporting flexible deployment on multiple devices (including CPU, GPU, NPU, BPU, etc.);
[0017] Receiving and processing module: This module uses the UDP protocol to receive motion capture data from external devices, including 33 3D body key points, 21 3D key points for each hand, and 478 3D facial key points. The module then implements data parsing, coordinate conversion, and quaternion calculation in the UDPSubsystem, using an adaptive interpolation mechanism to smooth the data and effectively suppress jitter.
[0018] Posture analysis module: responsible for detecting specific movement patterns in real time;
[0019] Timeout detection module: responsible for valid data detection and timeout processing, using the hand data timeout reset mechanism to solve the data loss in motion capture;
[0020] Skeletal Control Module: This module calculates the relationship between the palm orientation and the forearm axis to achieve a natural twisting effect, and enables overall character movement based on changes in pelvic position. It supports smooth transitions and anomaly detection, and ensures a smooth transition to the default posture when data is lost through the intelligent application and automatic reset of hand data.
[0021] Furthermore, the terrain simulation module uses a Landscaping plug-in to import GIS data to create Landscape, WorldComposition or WorldPartition. The Landscaping plug-in supports procedural generation of roads and vegetation based on landuse data, and combines with Mapbox to realize online map textures (the relevant authorization will comply with the Unreal Engine's own End User License Agreement (EULA), Autodesk's official Maya XGen Groom commercial authorization, Mapbox's relevant terms of service and other authorization terms).
[0022] Furthermore, the specific operation of the deep customization is as follows: the user superimposes custom geometry based on MetaHuman's mesh and rig data, and achieves 1:1 matching through the Wrap3 tool.
[0023] Furthermore, the hair simulation module reduces the number of CVs by constructing curve templates and cylinder winding techniques, preventing Unreal Engine crashes due to high-density hair, balancing performance and detail, and enabling cross-platform dynamic effect import through Alembic caching. Furthermore, by using Maya's XGen combined with GS CurveTools to create layered hair bundles, FiberShop to generate hair clips and bake textures, the Physion plug-in combined with DER physics simulation, Niagara GPU simulation, and Epic's Metahuman shader, hair thickness, density, and simulation parameters (such as bend damping) are adjusted to ensure the compatibility of complex hairstyles across different hardware devices.
[0024] Furthermore, the motion capture module is deployed separately on a computing mainboard, calls the camera to collect video, obtains the motion capture data and sends it to another computing mainboard via UDP for subsequent processing and application, and the two computing mainboards are connected by a network cable.
[0025] Furthermore, the motion capture module uses a cascade detection method to improve accuracy and efficiency, and key information is directly transmitted between modules. The core process is as follows:
[0026] 1) The camera collects video frames and preprocesses them;
[0027] 2) Key facial feature detection: The model outputs estimates of 478 3D facial feature points;
[0028] 3) Posture detection: Detect the entire image to obtain the human body area;
[0029] 4) Pose keypoint extraction: Based on the cropped image of the human body region, 33 3D human body keypoints are inferred;
[0030] 5) Hand region detection: Use the posture results to determine the wrist position and crop the high-resolution hand region;
[0031] 6) Hand key point extraction: Detect 21 3D key points of each hand in the ROI;
[0032] 7) Key point smoothing: filtering and smoothing the detection results;
[0033] 8) Result formatting and UDP transmission: Format the data and send it to the target address.
[0034] Furthermore, the human body region cropped image is a 2D image, and its 3D coordinate inference is specifically achieved by the following method:
[0035] Ⅰ. 3D coordinate inference implementation:
[0036] Realize the mapping from monocular camera input to 3D motion capture;
[0037] Depth estimation is added during model training to encode relative depth into the z coordinate;
[0038] The x and y coordinates are normalized to between 0 and 1, and the z coordinate represents the relative depth;
[0039] II. Model training data source and optimization:
[0040] Use open source datasets containing depth information;
[0041] Customized datasets collected using depth cameras;
[0042] Ensure data accuracy through manual calibration;
[0043] Align and unify the depth data of multiple datasets to establish a unified depth reference system;
[0044] Perform data enhancement and cleaning;
[0045] III. Model architecture upgrade:
[0046] Upgrade the original 2D key point definition to 3D key points including depth estimation;
[0047] Optimize loss function design to balance 2D and 3D prediction accuracy.
[0048] Furthermore, the motion capture module uses the human body detection results of the previous frame to predict the position of the human body in the current frame, thereby reducing the frequent calls to full-image detection and lowering the computational overhead. Specific strategies include:
[0049] I. Predict the displacement of the human body in the current frame based on the detection results of the previous frame, and make reasonable estimates using velocity and acceleration information;
[0050] II. In most cases, the prediction results are used to skip the expensive full-frame detection, and the detection is re-triggered only when the prediction deviates significantly.
[0051] Furthermore, the motion capture module adopts different strategies for tracking key points and hands to ensure smooth and stable detection results:
[0052] Ⅰ. Human body key point stability: Adopting the One-Euro filtering algorithm, it filters high-frequency noise while maintaining low latency and reduces detection jitter;
[0053] II. Hand keypoint stabilization: Using the hand ROI information from the previous frame, the hand ROI is adjusted through a more stable detection framework and real-time smoothing algorithm to ensure that the hand details and position remain stable.
[0054] The performance optimization of the motion capture module specifically includes:
[0055] Ⅰ. The model is quantified specifically for edge computing devices, supports NPU or BPU acceleration, and reduces CPU load;
[0056] Ⅱ. Using data pipeline processing, each processing stage is executed in parallel;
[0057] III. Implement an asynchronous processing framework to eliminate blocking between IO and computing;
[0058] IV. Optimize memory allocation and data transfer to reduce unnecessary copies.
[0059] Furthermore, the posture analysis module specifically includes:
[0060] T-posture correction: By analyzing the angle of the arm relative to the spine, it determines whether the user is in a T-posture for automatic calibration and posture reset;
[0061] Body rotation analysis: Analyze the difference in upper and lower body rotation to determine whether it is full-body rotation or partial rotation, so as to apply digital human motion performance more naturally.
[0062] The present invention provides an immersive wearable motion capture digital human body sensory interaction system, which has the following beneficial effects:
[0063] The present invention adopts double data verification, combined with static posture detection and dynamic timeout mechanism, to ensure the reliability of motion capture data, and smoothly processes the entire process from data reception to skeleton application, so that the movement can be coherent and natural, and cooperates with intelligent forearm twisting, and can automatically calculate the forearm twisting angle based on the palm direction, solving the problem of unnatural elbow rotation in traditional IK systems. At the same time, the system can perform adaptive displacement detection, intelligently judge the overall displacement through changes in pelvic position, and support multiple coordinate systems, so that the overall movement of the body can be smoothly detected and applied without stuttering, especially being able to accurately identify and respond to special postures (such as T posture). In addition, the runtime performance is greatly improved by pre-caching the bone index and transformation matrix.
[0064] This system is based on Coolyue Technology's original dual-computing motherboard hardware architecture. Through full-stack algorithm optimization from the bottom layer of the code to the application layer, it achieves better interactive effects with lower hardware computing power requirements, thereby realizing low-latency and highly natural digital human body interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 This is a logical block diagram of an immersive, wearable motion capture digital human body interaction system according to the present invention. DETAILED DESCRIPTION
[0066] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0067] like Figure 1 As shown, an immersive wearable motion capture digital human body interaction system includes:
[0068] Scene simulation layer
[0069] Construct and simulate the entire simulation environment or background to provide a realistic background for characters and actions, making the simulation more realistic and immersive. In this embodiment, the scene simulation layer includes the following modules:
[0070] Terrain simulation module: This module uses oblique photography to capture real-world terrain and surrounding environments, and uses the Mapbox API to obtain map data (such as vector tiles and terrain elevations), integrating it with UE5's Landscape system to build terrain. Specifically, the terrain simulation module uses the Landscaping plugin to import GIS data to create Landscapes, WorldCompositions, or WorldPartitions. The Landscaping plugin supports procedural road generation and vegetation generation based on landuse data, while integrating with Mapbox to implement online map textures.
[0071] Terrain Processing Module: Through the Quixel Megascans asset library, millions of scanned real-world materials (such as rocks and soil) are integrated and applied directly to the terrain. Combined with custom masks (such as moss distribution and weathering effects), high-precision surface details are achieved. Specifically, node-based tools are used to define rules (such as vegetation distribution and rock placement) to automatically generate large-scale terrain. For example, after setting parameters, UE5 can dynamically generate diverse landforms such as forests and deserts, and realize randomization of level layouts.
[0072] Shadow processing module: Provides delicate shadow transitions through virtual shadow mapping technology, achieving consistent light and shadow at near and far distances;
[0073] Dynamic Weather Module: Combines UE5's Niagara particle system and volumetric fog to simulate the effects of rain, snow, sandstorms and other weather on the terrain, enhancing the sense of immersion.
[0074] Character simulation layer
[0075] Simulate the characters in the simulation to achieve a virtual human model with film-level realism. In this embodiment, the character simulation layer includes the following modules:
[0076] Character Simulation Module: This module uses MetaHumans' preset universal character templates as a starting point for creation, and supports export to DCC software such as Maya for in-depth customization. Users can overlay custom geometry based on MetaHuman's mesh and rig data, and achieve 1:1 matching through the Wrap3 tool. This modular design retains MetaHuman's standardized animation compatibility while allowing technicians to break through preset limitations.
[0077] Hair Simulation Module: Complex hairstyles are achieved through the MayaXGenGroom tool, and linear wire deformers are used to drive dynamic hair simulation. Specifically, the number of CVs is reduced through the construction of curve templates and cylinder winding techniques to avoid Unreal Engine crashes due to high-density hair, thereby balancing performance and detail. Cross-platform dynamic effect import is achieved through Alembic caching. In addition, by using Maya's XGen combined with GS CurveTools to create layered hair bundles, FiberShop to generate hairpins and bake textures, using the Physion plug-in combined with DER physics simulation, Niagara GPU simulation, and using Epic's Metahuman shader, etc.; hair thickness, density, and simulation parameters (such as bend damping) are adjusted to ensure the compatibility of complex hairstyles on different hardware devices.
[0078] Motion simulation layer
[0079] Simulate the specific actions of the human character, including action capture, editing, synthesis, and rendering, to ensure that the character's actions are smooth and natural, and match the environment and task, as well as handle the interaction between the character and the environment, such as collision detection and physical feedback. In this embodiment, the action simulation layer includes the following modules:
[0080] Motion capture module: Human body posture and gesture key point detection, providing a flexible framework to process and transmit human body posture and gesture key point data, and supporting flexible deployment on multiple devices (including CPU, GPU, NPU, BPU, etc.).
[0081] The motion capture module is deployed independently on one computing board. It uses the camera to capture video and sends the acquired motion capture data via UDP to another computing board for subsequent processing and application. The two computing boards are connected by a network cable. This distributed architecture design enables the system to better balance the computing load and ensure the real-time performance of the motion capture system.
[0082] The motion capture module uses a cascade detection method to improve accuracy and efficiency. Key information is directly transmitted between modules. For example, the results of posture detection are used to intercept the human body area. Subsequent posture key point extraction is based on key point inference in this area. The core process is as follows:
[0083] 1) The camera collects video frames and preprocesses them;
[0084] 2) Key facial feature detection: The model outputs estimates of 478 3D facial feature points;
[0085] 3) Posture detection: Detect the entire image to obtain the human body area;
[0086] 4) Pose keypoint extraction: Based on the cropped image of the human body region, 33 3D human body keypoints are inferred;
[0087] 5) Hand region detection: Use the posture results to determine the wrist position and crop the high-resolution hand region;
[0088] 6) Hand key point extraction: Detect 21 3D key points of each hand in the ROI;
[0089] 7) Key point smoothing: filtering and smoothing the detection results;
[0090] 8) Result formatting and UDP transmission: Format the data and send it to the target address.
[0091] The human body region cropped image is a 2D image, and its 3D coordinate inference is achieved in the following way:
[0092] Ⅰ. 3D coordinate inference implementation:
[0093] Realize the mapping from monocular camera input to 3D motion capture;
[0094] Depth estimation is added during model training to encode relative depth into the z coordinate;
[0095] The x and y coordinates are normalized to between 0 and 1, and the z coordinate represents the relative depth;
[0096] II. Model training data source and optimization:
[0097] Use open source datasets containing depth information;
[0098] Customized datasets collected using depth cameras;
[0099] Ensure data accuracy through manual calibration;
[0100] Align and unify the depth data of multiple datasets to establish a unified depth reference system;
[0101] Perform data enhancement and cleaning;
[0102] III. Model architecture upgrade:
[0103] Upgrade the original 2D key point definition to 3D key points including depth estimation;
[0104] Optimize loss function design to balance 2D and 3D prediction accuracy.
[0105] The motion capture module uses the human body detection results of the previous frame to predict the position of the human body in the current frame, reducing the frequent calls to full-image detection and lowering the computational overhead. Specific strategies include:
[0106] I. Predict the displacement of the human body in the current frame based on the detection results of the previous frame, and make reasonable estimates using velocity and acceleration information;
[0107] II. In most cases, the prediction results are used to skip the expensive full-frame detection, and the detection is re-triggered only when the prediction deviates significantly.
[0108] To ensure smooth and stable detection results, the motion capture module adopts different strategies for human key points and hand tracking:
[0109] Ⅰ. Human body key point stability: Adopting the One-Euro filtering algorithm, it filters high-frequency noise while maintaining low latency and reduces detection jitter;
[0110] II. Hand keypoint stabilization: Using the hand ROI information from the previous frame, the hand ROI is adjusted through a more stable detection framework and real-time smoothing algorithm to ensure that the hand details and position remain stable.
[0111] The performance optimization of the motion capture module includes:
[0112] Ⅰ. The model is quantified specifically for edge computing devices, supports NPU or BPU acceleration, and reduces CPU load;
[0113] Ⅱ. Using data pipeline processing, each processing stage is executed in parallel;
[0114] III. Implement an asynchronous processing framework to eliminate blocking between IO and computing;
[0115] IV. Optimize memory allocation and data transfer to reduce unnecessary copies.
[0116] Receiving and processing module: Uses the UDP protocol to receive motion capture data from external devices, including 33 3D body key points, 21 3D key points for each hand, and 478 3D facial key points. It implements data parsing, coordinate conversion, and quaternion calculation in the UDPSubsystem, and uses an adaptive interpolation mechanism to smooth the data and effectively suppress jitter.
[0117] Posture Analysis Module: Detects specific motion patterns in real time, including:
[0118] T-posture correction: By analyzing the angle of the arm relative to the spine, it determines whether the user is in a T-posture for automatic calibration and posture reset;
[0119] Body rotation analysis: Analyze the difference in upper and lower body rotation to determine whether it is full-body rotation or partial rotation, so as to apply digital human motion performance more naturally.
[0120] Timeout detection module: data valid detection and timeout processing, using the hand data timeout reset mechanism to solve the data loss in motion capture.
[0121] Skeletal Control Module: This module calculates the relationship between the palm orientation and the forearm axis to achieve a natural twisting effect, resolving the unnatural elbow twisting problem in traditional IK systems. It also enables overall character movement based on pelvic position changes, supports smooth transitions and anomaly detection, and ensures a smooth transition to the default pose when data is lost through intelligent application and automatic resetting of hand data.
[0122] In this embodiment, the technology of the skeleton-driven motion capture and displacement determination system is implemented as follows:
[0123] 1. Calculation method of the spatial relationship between palm and forearm
[0124] Based on the skeletal point position data, this system uses the principles of analytical geometry to calculate the precise spatial relationship between the palm and forearm, achieving high-precision wrist flexion and extension and forearm rotation control.
[0125] 1.1 Calculation of forearm pronation and supination angles
[0126] The forearm's supination / pronation angles are calculated using the following steps:
[0127] 1) Define the forearm axial vector vec(v_arm) as the unit vector from the elbow to the wrist;
[0128] 2) Define the palm normal vector vec(n_palm) calculated by the palm key points;
[0129] 3) Define the palm-up direction vec(u_palm) as: vec(u_palm) = vec(v_arm) × vec(n_palm);
[0130] 4) Calculate the ideal upward direction vec(u_ideal) as:
[0131] vec(u_ideal)=vec(u_ref)-(vec(u_ref)·vec(v_arm))·vec(v_arm) where vec(u_ref) is the reference upward direction (usually the torso upward direction);
[0132] 5) Calculate the pronation and supination angles θ_sup as:
[0133] θ_sup=cos^(-1)(vec(u_palm)·vec(u_ideal));
[0134] 6) Angle direction determination: If (vec(u_palm)×vec(u_ideal))·vec(v_arm)<0, then θ_sup=-θ_sup;
[0135] 1.2 Calculation of wrist flexion and extension angle
[0136] The flexion / extension angle of the wrist is calculated using the following steps:
[0137] 1) Define the palm direction vector vec(d_hand) as the unit vector from the wrist to the base of the middle finger;
[0138] 2) Project the palm direction vector to the vertical plane of the forearm: vec(d_proj) = vec(d_hand) - (vec(d_hand) · vec(v_arm)) · vec(v_arm);
[0139] 3) Calculate the flexion-extension rotation axis vec(a_flex): vec(a_flex) = vec(v_arm) × vec(d_proj);
[0140] 4) Calculate the flexion and extension angle θ_flex:
[0141] θ_flex=cos^(-1)(vec(d_proj)·vec(u_palm))
[0142] 5) Angle direction determination: If vec(d_proj)·(vec(a_flex)×vec(u_palm))<0, then θ_flex=-θ_flex
[0143] 1.3 Skeleton Rotation Quaternion Synthesis
[0144] Convert the calculated angle to a quaternion rotation and apply it to the corresponding bone:
[0145] 1) Forearm rotation quaternion: Q_forearm = Q(vec(v_arm),θ_sup·0.6) where Q(vec(v),θ) represents the quaternion of θ radians rotation around the axis vec(v);
[0146] 2) Wrist flexion and extension quaternion: Q_wrist = Q(vec(a_flex),θ_flex·0.7);
[0147] 3) Update skeleton pose:
[0148] Q_final_forearm=Q_forearm·Q_current_forearmQ_final_wrist=Q_wrist·Q_current_wrist
[0149] 2. Displacement recognition system based on cadence analysis
[0150] This system adopts a dynamic time window cadence analysis method to achieve high-precision walking displacement recognition and speed calculation.
[0151] 2.1 Step Detection Algorithm
[0152] 1) Define the ground reference height: Z_ground = Z_pelvis - h_offset, where h_offset is the predefined vertical offset from the pelvis to the ground (usually 90 cm);
[0153] 2) Foot lift status determination: Left foot lifted: (Z_left_foot - Z_ground) > h_threshold Right foot lifted: (Z_right_foot - Z_ground) > h_threshold, where h_threshold is the foot lift height threshold (usually 15 cm);
[0154] 3) Step detection conditions:
[0155] 2.2 Time Window Cadence Analysis
[0156] 1) Maintain the step timestamp queue Q = {t_1, t_2, ..., t_n};
[0157] 2) When a new step is detected, add the current time t_current to the queue;
[0158] 3) Remove timestamps from the queue that exceed the time window: if(t_current-t)>T_windowthenremovetfromQ where T_window is the time window value (usually 0.5 seconds);
[0159] 4) Calculate the step frequency: f_step = |Q| / T_window where |Q| is the number of timestamps in the queue;
[0160] 5) Calculate the target movement speed: v_target = f_step·S_step, where S_step is the speed coefficient corresponding to each step (usually 60 cm / s);
[0161] 6) Smooth velocity changes: v_current = FInterp(v_current, v_target, Δt, α) where α is the smoothing coefficient (usually 5.0) and FInterp is the smoothing interpolation function;
[0162] 2.3 Body Orientation Calculation
[0163] Body orientation is calculated using the following steps:
[0164] 1) Get the position vectors of the pelvis and the left and right clavicles: vec(v_left) = vec(p_left_clavicle) -
[0165] vec(p_pelvis)vec(v_right)=vec(p_right_clavicle)-vec(p_pelvis)
[0166] 2) Calculate the normal vector: vec(n) = vec(v_right) × vec(v_left)
[0167] 3) Projection to the horizontal plane: vec(n_horizontal) = norm(n_x,n_y,0)
[0168] 4) Set the movement direction: vec(d_move) = vec(n_horizontal)
[0169] 3. Pelvic rotation cumulative calculation
[0170] This system implements a cumulative calculation method of pelvic rotation based on arm posture, which effectively solves the problem of insufficient pelvic rotation in motion capture.
[0171] 3.1 Arm Posture Detection
[0172] 1) Left arm 90 degree posture judgment: θ_left_arm = ∠(vec(p_left_shoulder)-vec(p_left_hip),vec(p_left_shoulder)-vec(p_left_elbow))LeftArm90 = |θ_left_arm-90°|<θ_threshold;
[0173] 2) Right arm 90-degree posture determination: θ_right_arm=∠(vec(p_right_shoulder)-vec(p_right_hip),vec(p_right_shoulder)-vec(p_right_elbow))RightArm90=|θ_right_arm-90°|<θ_threshold;
[0174] 3.2 Pelvic rotation accumulation
[0175] 1) When the left arm is at 90 degrees, rotate the pelvis to the right: Q_Δ_left = Q(vec(v_forward), +3°) Q_pelvis = Q_Δ_left·Q_pelvis;
[0176] 2) When the right arm is at 90 degrees, rotate the pelvis to the left: Q_Δ_right = Q(vec(v_forward), -3°) Q_pelvis = Q_Δ_right·Q_pelvis;
[0177] 4. Calculation of body plane normal vector
[0178] This system proposes a robust body plane normal vector calculation method based on four-point construction, which can be effectively applied to scenarios such as wrist position adjustment.
[0179] 4.1 Calculation of body plane normal vector
[0180] 1) Get four key points: left shoulder vec (p_LS), right shoulder vec (p_RS), left hip vec (p_LH), right hip vec (p_RH)
[0181] 2) Calculate the diagonal vector:
[0182] vec(d_1)=vec(p_LS)-vec(p_RH)vec(d_2)=vec(p_RS)-vec(p_LH);
[0183] 3) Calculate the plane normal vector: vec(n_body) = norm(vec(d_1) × vec(d_2));
[0184] 4) Ensure the normal vector is facing forward: if vec(n_body)·vec(v_forward)<0, then vec(n_body)=-vec(n_body);
[0185] 4.2 Adjustment of wrist position based on visibility
[0186] 1) Get the wrist visibility value v_wrist;
[0187] 2) When visibility is lower than the threshold, adjust the wrist position in the opposite direction of the body plane normal vector: vec(p_wrist_adjusted) = vec(p_wrist) - d_adjustment · vec(n_body), where d_adjustment is the adjustment distance (usually 15 cm).
[0188] The embodiments of the present invention are presented for purposes of illustration and description and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments are chosen and described in order to better illustrate the principles of the invention and its practical application and to enable those skilled in the art to understand the invention and design various embodiments with various modifications as suited for specific applications.
Claims
1. An immersive wearable motion capture digital human body interaction system, characterized by: Including scene simulation layer, character simulation layer and action simulation layer; The character simulation layer is responsible for simulating the characters in the simulation and realizing a virtual human model with film-level realism. It includes the following modules: Character simulation module: uses the preset universal character templates provided by MetaHumans as the starting point for creation, and supports exporting to DCC software for in-depth customization; Hair simulation module: Use MayaXGenGroom tool to achieve complex hairstyles and use linear wire deformer to drive hair dynamic simulation; The action simulation layer is responsible for simulating the specific actions of the character and handling the interaction between the character and the environment. It includes the following modules: Motion capture module: responsible for human posture and gesture key point detection, providing a flexible framework for processing and transmitting human posture and gesture key point data, and supporting flexible deployment on multiple devices; Receiving and processing module: uses UDP protocol to receive motion capture data from external devices, implements data parsing, coordinate conversion and quaternion calculation in UDPSubsystem, and uses adaptive interpolation mechanism to smoothly process data; Posture analysis module: responsible for real-time detection of specific movement patterns; Timeout detection module: responsible for valid data detection and timeout processing, using the hand data timeout reset mechanism to solve the data loss in motion capture; Skeletal control module: This module calculates the relationship between the palm orientation and the forearm axis to achieve a natural twisting effect, and realizes overall character movement based on changes in pelvic position, supporting smooth transitions and anomaly detection.
2. The immersive wearable motion capture digital human body interaction system according to claim 1, characterized in that: The terrain simulation module uses the Landscaping plug-in to import GIS data to create Landscape, WorldComposition or WorldPartition. The Landscaping plug-in supports procedural generation of roads and vegetation based on landuse data, and is combined with Mapbox to implement online map textures.
3. The immersive wearable motion capture digital human body interaction system according to claim 1, characterized in that: The scene simulation layer is responsible for constructing and simulating the entire simulation environment or background, providing a realistic background for characters and actions, and includes the following modules: Terrain simulation module: responsible for restoring the real terrain, using oblique photography technology to obtain the real terrain and surrounding environment, and obtaining map data through Mapbox API, and building the terrain in combination with UE5's Landscape system; Terrain Processing Module: Integrates real-world materials through the Quixel Megascans asset library and applies them directly to terrain. Combined with custom masks, it achieves high-precision surface details. Specifically, it uses node-based tools to define rules and automatically generate large-scale terrain. Shadow processing module: provides shadow transition through virtual shadow mapping technology to achieve light and shadow consistency at long and short distances; Dynamic Weather Module: Combines UE5's Niagara particle system and volumetric fog to simulate the impact of weather on terrain and enhance immersion.
4. The immersive wearable motion capture digital human body interaction system according to claim 1, characterized in that: The specific operation of the deep customization is as follows: users overlay custom geometry based on MetaHuman's mesh and binding data, and achieve 1:1 matching through the Wrap3 tool.
5. The immersive wearable motion capture digital human body interaction system according to claim 1, characterized in that: The hair simulation module reduces the number of CVs by constructing curve templates and cylinder winding technology, avoiding the crash of the Unreal Engine due to high-density hair, and realizes cross-platform dynamic effect import through Alembic cache.
6. The immersive wearable motion capture digital human body interaction system according to claim 1, characterized in that: The motion capture module is deployed separately on a computing motherboard, calls the camera to collect video, obtains the motion capture data and sends it to another computing motherboard via UDP for subsequent processing and application, and the two computing motherboards are connected by a network cable; The motion capture module adopts a cascade detection method, and key information is directly transmitted between modules. The core process is as follows: 1) The camera collects video frames and preprocesses them; 2) Key facial feature detection: The model outputs estimates of 478 3D facial feature points; 3) Posture detection: Detect the entire image to obtain the human body area; 4) Pose keypoint extraction: Based on the cropped image of the human body region, 33 3D human body keypoints are inferred; 5) Hand region detection: Use the posture results to determine the wrist position and crop the high-resolution hand region; 6) Hand key point extraction: Detect 21 3D key points of each hand in the ROI; 7) Key point smoothing: filtering and smoothing the detection results; 8) Result formatting and UDP transmission: Format the data and send it to the target address.
7. The immersive wearable motion capture digital human body interaction system according to claim 6, characterized in that: The human body region cropped image is a 2D image, and its 3D coordinate inference is specifically achieved by the following method: 3D coordinate inference implementation: Realize the mapping from monocular camera input to 3D motion capture; Depth estimation is added during model training to encode relative depth into the z coordinate; The x and y coordinates are normalized to between 0 and 1, and the z coordinate represents the relative depth; Model training data source and optimization: Use open source datasets containing depth information; Customized datasets collected using depth cameras; Ensure data accuracy through manual calibration; Align and unify the depth data of multiple datasets to establish a unified depth reference system; Perform data enhancement and cleaning; Model architecture upgrade: Upgrade the original 2D key point definition to 3D key points including depth estimation; Optimize loss function design to balance 2D and 3D prediction accuracy.
8. The immersive wearable motion capture digital human body interaction system according to claim 7, characterized in that: The motion capture module uses the human body detection results of the previous frame to predict the position of the human body in the current frame to reduce the frequent calls to full-image detection and reduce computational overhead. Specific strategies include: I. Predict the displacement of the human body in the current frame based on the detection results of the previous frame, and make reasonable estimates using velocity and acceleration information; II. In most cases, the prediction results are used to skip the expensive full-frame detection, and the detection is re-triggered only when the prediction deviates significantly.
9. The immersive wearable motion capture digital human body interaction system according to claim 8, characterized in that: The motion capture module adopts different strategies for human key points and hand tracking: Ⅰ. Stable human key points: Using a Euro filtering algorithm to filter high-frequency noise while maintaining low latency and reduce detection jitter; II. Hand keypoint stabilization: Using the hand ROI information from the previous frame, the hand ROI is adjusted through a detection framework and real-time smoothing algorithm to ensure that the hand details and position remain stable. The performance optimization of the motion capture module specifically includes: Ⅰ. The model is quantified specifically for edge computing devices, supports NPU or BPU acceleration, and reduces CPU load; Ⅱ. Using data pipeline processing, each processing stage is executed in parallel; III. Implement an asynchronous processing framework to eliminate blocking between IO and computing; IV. Optimize memory allocation and data transfer to reduce unnecessary copies.
10. The immersive wearable motion capture digital human body interaction system according to claim 1, characterized in that: The posture analysis module specifically includes: T-posture correction: By analyzing the angle of the arm relative to the spine, it determines whether the user is in a T-posture for automatic calibration and posture reset; Body rotation analysis: Analyze the difference in upper and lower body rotation to determine whether it is full-body rotation or partial rotation, so as to apply digital human motion performance more naturally.
Citation Information
Patent Citations
Human-computer interaction method and device based on virtual reality
CN116382486A
Universal 3D digital human real-time motion capture method and system based on monocular camera
CN119417958A
Digital twin WEB application air gesture interaction method and system
CN119723653A
Human-computer interaction method for dummy ape game
CN1731316A
Key point detection model training method and apparatus and virtual character driving method and apparatus
WO2024060978A1