Interaction method and system
Through data collection and preprocessing, a three-dimensional skeletal model is generated and mapped into the virtual environment, solving the problems of expensive motion capture equipment and poor tracking effects, achieving full-body motion capture and natural interaction, and enhancing the immersive experience in the metaverse.
Patent Information
- Application Number
- CN202410804417.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-20
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-06-20
AI Technical Summary
Motion capture equipment is expensive and virtual reality devices do not have full-body tracking, resulting in poor immersive experience for users in the metaverse.
Through modules such as data acquisition, preprocessing, posture estimation, data recognition, organization and transmission, optimization, cloud integration, compensation and user interaction, combined with headsets, handles, sensors and additional cameras to collect data, a three-dimensional skeleton model is generated and mapped to the virtual environment to provide visual and tactile feedback and achieve full-body motion capture effects.
It improves the user's immersive experience in the metaverse, reduces equipment costs by optimizing and compensating accuracy, and enables full-body tracking and natural interaction.
Smart Images

Figure CN120219672B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of interactive technology, and more particularly to an interactive method and system. Background Art
[0002] The Metaverse is a virtual world that integrates virtual reality, augmented reality, blockchain, and other advanced technologies. In this virtual world, users can interact, socialize, work, entertain, and conduct various economic activities through digital avatars. The Metaverse includes Q-version character models and two-dimensional character models. If users want a better immersive experience, they need to purchase motion capture equipment. The motion capture equipment obtains limb posture data and feeds it back to the Metaverse to achieve a full-body motion capture effect. However, motion capture equipment is expensive, and virtual reality equipment does not have full-body tracking effects, which indirectly leads to a poor user immersion experience. Summary of the Invention
[0003] The purpose of the present invention is to address the shortcomings of the existing technology and provide an interactive method and system, aiming at data collection and preprocessing, establishing a virtual model from posture estimation to data recognition, and achieving image and sensor fusion motion capture effects by optimizing and compensating accuracy and feeding back to the metaverse.
[0004] To this end, this application provides an interactive system, including the following modules:
[0005] Data collection module, which collects user action and environmental data;
[0006] Preprocessing module, which preprocesses the collected data to improve data quality;
[0007] The posture estimation module detects the user's key point positions based on the data and generates a 3D skeleton model;
[0008] Data identification module, based on data and modeling model identification;
[0009] The transmission module includes a first-class port and a second-class port, and uses the first-class port or the second-class port for transmission according to the transmission data;
[0010] Optimization module, which performs preliminary optimization based on model parameters;
[0011] Cloud integration module, overall optimization based on the model;
[0012] Compensation module, based on model error compensation after optimization;
[0013] User interaction module, which provides an interactive interface and enables natural interaction through 3D models;
[0014] Data mapping module, which maps the user's 3D posture and movements into the virtual environment;
[0015] Feedback module, providing visual and tactile feedback.
[0016] In some specific embodiments, the data collection module is used to collect user motion and environmental data, collect user current motion data based on the head display, sensors and handles, and collect user current motion based on additional cameras and head display cameras;
[0017] VR headset camera, with a built-in high-frame-rate camera, captures real-time video streams of the user's upper body and hands;
[0018] Handle sensor: built-in inertial measurement unit, real-time transmission of hand motion data;
[0019] Additional cameras: Wide-angle lenses and high-frame-rate cameras are installed in the four corners of the room to capture full-body video streams;
[0020] The preprocessing module preprocesses the collected data to improve data quality, enhances the collected image data, and ensures data synchronization and spatial calibration;
[0021] Based on image processing, including image denoising, image contrast adjustment, image cropping, image scaling and data enhancement;
[0022] Image denoising: Use a Gaussian filter to smooth the image and remove high-frequency noise, and a median filter to remove impulse noise in the image;
[0023] Image contrast adjustment: Improve image contrast by equalizing the image histogram, and use adaptive histogram equalization to enhance contrast in local areas to avoid over-enhancement;
[0024] Image cropping: Cropping based on preset areas, such as key areas of the head and hands, and dynamically adjusting the cropping area based on the detected target position;
[0025] Image Scaling: Use bilinear interpolation to maintain image quality during scaling, and use bicubic interpolation for higher accuracy requirements;
[0026] Data augmentation: Randomly rotate the image by a certain angle, flip the image horizontally or vertically, randomly scale the image, and randomly change the brightness, contrast, and saturation of the image;
[0027] Data synchronization and calibration, including time synchronization, spatial calibration, data standardization and data quality testing;
[0028] Time synchronization: Ensure the temporal consistency of data collected by different sensors by adding a timestamp to each data sample, aligning data from different sensors based on the timestamp, and interpolating data with incomplete timestamp matching to ensure data synchronization;
[0029] Spatial calibration: Ensures spatial consistency between different sensors, especially in the fusion process of multiple cameras and IMU sensors. Use calibration plates to calibrate the internal and external parameters of the camera to determine the intrinsic and external parameters of the camera. In a multi-camera system, hand-eye calibration is used to determine the position of the camera relative to the robot or fixed structure. Static and dynamic calibration of IMU sensors is performed to determine the sensor bias and scale factor.
[0030] Data standardization: Unify data formats and scales to ensure that data from different sources are processed at the same scale, normalize data to a unified range, eliminate dimensional differences, and standardize data to a mean of 0 and a variance of 1 to eliminate mean and variance differences between different data sources;
[0031] Data quality detection: Detect and filter low-quality data, ensure the reliability of input data, use statistical methods or machine learning models to detect and filter abnormal data, check data integrity, and filter out data samples that are missing key parts.
[0032] In some specific embodiments, the posture estimation module detects the key point positions of the user based on the data and generates a three-dimensional skeleton model;
[0033] The motion capture module fuses data from different sensors to generate the user's three-dimensional full-body motion;
[0034] The space building module generates an adaptive 3D virtual space based on the user's 3D skeleton model and motion data;
[0035] Model building module, which integrates multiple data and builds a complete virtual environment model;
[0036] Data fusion module, which fuses data from multiple sensors to generate consistent 3D model data;
[0037] The action prediction module predicts the user's next action based on the fused 3D model and by analyzing historical data;
[0038] The model correction module is used to correct errors and inconsistencies that occur during the data fusion process.
[0039] In some specific embodiments, the posture estimation module detects the key point positions of the user based on the data and generates a three-dimensional skeletal model, using existing public datasets to collect data that meets the specific requirements of the system;
[0040] Label the key points of the human body for the image data, use the pre-trained model for preliminary labeling, and then manually correct it;
[0041] Generate a 3D skeleton model using a 2D pose estimation model and a 3D pose estimation model;
[0042] The 2D pose estimation model uses HRNet to process images with rich details, and the 3D pose estimation model uses VideoPose3D to convert 2D key point sequences into 3D poses;
[0043] HRNet receives an input image and converts it into a format suitable for network processing through a series of preprocessing steps (such as normalization, scaling, and cropping). It then extracts initial features through several 3×3 convolutional layers and uses the ReLU activation function for nonlinear transformation. The feature maps generated by the convolutional layers are used to maintain high resolution and enter the high-resolution branch for processing. The network passes the feature maps to multiple parallel convolutional layers with different resolutions. By adjusting the stride and pooling operations of the convolutional layers, feature maps of different resolutions are generated. At the end of each stage, HRNet fuses the feature maps of different resolutions through feature fusion units. These fusion units use interpolation upsampling and convolution downsampling methods to align and fuse the multi-resolution feature maps together, thereby maintaining high-resolution information and enhancing feature expression capabilities.
[0044] HRNet continuously enhances and fuses multi-resolution features through repeated operations of multiple fusion units, which include exchange units and multi-layer convolution;
[0045] Exchange unit: At each stage, feature maps are fused through the exchange unit to form a new multi-resolution feature map;
[0046] Multi-layer convolution: Through multi-layer convolution operations, features are further extracted and fused;
[0047] The output of HRNet is a keypoint heatmap, where each keypoint corresponds to a channel. Each pixel value in the heatmap represents the probability of that location being the corresponding keypoint. In the final layer, a convolutional layer is used to convert the high-resolution feature map into a multi-channel heatmap, where each channel corresponds to a keypoint. The Softmax activation function is used to normalize the heatmap to represent a probability distribution.
[0048] In each heat map channel, the maximum value is found. This position is the predicted key point position. Gaussian filtering is used to refine the heat map to obtain sub-pixel key point positions. The key point positions are then smoothed to reduce jitter and noise. The key point positions are constrained according to the human skeletal structure to ensure that the generated skeletal data is reasonable.
[0049] VideoPose3D receives a sequence of 2D key points within a time window as input. The model uses a fixed-length time window at each time step to capture the changes in 2D key points over a period of time. The input format is ,in is the time window length, is the number of key points, each key point contains 2 coordinates ;
[0050] A 3D convolution layer is used to perform convolution operations in the time dimension to extract spatiotemporal features. The 3D convolution calculation involves sliding filters in the time and space dimensions to generate feature maps. Residual blocks are used to enhance the network's expressiveness and training depth to prevent gradient disappearance.
[0051] Through multi-layer 3D convolution and pooling operations, the extracted spatiotemporal features are input into the fully connected layer or regression layer to predict the position of 3D key points. The output format is ,in is the time window length, is the number of key points, each key point contains 3 coordinates ;
[0052] The mean square error is used as the loss function to minimize the error between the predicted 3D key points and the real 3D key points. The predicted 3D key point sequence is smoothed and filtered to reduce the jitter and noise of the prediction results. The position of the 3D key points is constrained according to the human skeletal structure.
[0053] The motion capture module combines visual data and inertial sensor data to generate 3D motion capture data through 3D point cloud processing and skeletal model matching technology;
[0054] Use timestamps to align visual data and IMU data to ensure they are fused at the same time. When there is a time difference, linear interpolation is used to synchronize the data. A Kalman filter is used to predict the attitude at the next moment using the attitude transfer matrix. Combined with the measurement matrix, the predicted attitude is adjusted using the Kalman gain to minimize the error.
[0055] Use nonlinear attitude transfer function and measurement function, perform attitude prediction and update after linearization;
[0056] The posture distribution is represented by a set of particles, which is applicable to complex nonlinear and non-Gaussian systems. The particle weights are updated according to the measurement values, and a new particle set is generated by resampling.
[0057] Fusion of visual and IMU data generates a 3D skeleton model to represent the user's posture and movements. Kalman filtering or mean filtering is used to smooth the positions of 3D key points to reduce jitter and noise. Human skeletal structure constraints are applied to ensure that the generated 3D skeleton data is reasonable.
[0058] The space building module builds an accurate and dynamic three-dimensional virtual environment by integrating multi-source data and modeling technology to adapt to the user's actual movements;
[0059] By integrating environmental perception, 3D modeling, dynamic adjustment and virtual environment optimization technologies, accurate and dynamic 3D virtual environments can be constructed, utilizing point cloud processing and registration, surface reconstruction and texture mapping, real-time updating and environmental adaptation, rendering optimization and physical interaction;
[0060] The model building module uses timestamps to align pose estimation and motion capture data, ensuring fusion at the same time point and calibrating data from different sensors;
[0061] A Kalman filter is used to fuse multi-source data to generate a consistent 3D skeleton model and motion data, and a corresponding 3D virtual environment is generated based on the user's movements and range of activity.
[0062] Adjust the size and layout of the virtual environment in real time based on user actions to ensure that the user's range of activities is fully reflected;
[0063] Dynamically generate virtual objects based on scene requirements and set physical properties for virtual objects;
[0064] Map the user's skeleton model to the virtual character's skeleton system and adjust the bone weights according to the character model;
[0065] Applying the data generated by the motion capture module to the virtual character to dynamically adjust the virtual character's posture and movements based on the user's real-time movements;
[0066] Use Unity to set physical properties for objects in the virtual environment to achieve real physical interaction;
[0067] The data fusion module uses Unity to set physical properties for objects in the virtual environment to achieve real physical interaction;
[0068] The data fusion module fuses data from different sensors to generate consistent 3D skeleton models and motion data through data synchronization, multimodal data fusion, and result optimization and post-processing steps;
[0069] Using Kalman filter, extended Kalman filter and particle filter technology to ensure the high accuracy and stability of the fused data;
[0070] The action prediction module collects a large amount of historical action data for training and prediction, processes the collected data to make it suitable for model training, selects an appropriate time window length to capture the dynamic characteristics of the action, extracts features from the time series for use in the prediction model, and uses the LSTM network to capture the long-term and short-term dependencies in the time series for action prediction. It uses stacked multi-layer LSTM units to increase the model's expressiveness, and uses a fully connected layer to convert the LSTM output into a prediction result.
[0071] The model correction module measures the error between the predicted data and the actual data, calculates the Euclidean distance of key point positions, compares time series data, measures the similarity between series, and dynamically adjusts the threshold based on historical data and the current environment;
[0072] The Kalman filter is used to correct the prediction results in real time, the attitude transfer matrix is used to predict the attitude at the next moment, and the predicted attitude is corrected in combination with the actual measurement data.
[0073] In some embodiments, the data identification module identifies and extracts feature data from the model output based on the data and the modeling model;
[0074] Data comparison module, which performs comparison processing based on the identified data;
[0075] Error marking module, marking the error between model output and real data;
[0076] The audit module performs audit processing based on the audit content preset in the preset module;
[0077] A preset module performs real-time review and processing based on the marking of error data.
[0078] In some specific embodiments, the data identification module, based on data and modeling model identification, prepares for subsequent error analysis and comparison by identifying and extracting characteristic data from the model output;
[0079] Based on the output data and real data after the model processing is completed, key feature points are extracted from the data, such as key points of human posture. The key points are extracted using a pre-trained posture estimation model, the key point position data is converted into vector form, and the identified feature data is annotated to provide a benchmark for subsequent use;
[0080] The data comparison module performs comparison processing based on the identified data to ensure the temporal consistency of the model output data and the real data, ensures that the predicted data and the real data are compared in the same coordinate system, and calculates the error between the model output and the real data;
[0081] MAE is used to calculate the average value of all key point errors, specifically:
[0082] First: True Value and predicted values , the true value is the actual observed data, and the predicted value It is the data obtained through model prediction;
[0083] Second: absolute error. Absolute error refers to the absolute difference between the true value and the predicted value, that is, ;
[0084] Third: sum, sum the absolute errors of all samples;
[0085] Fourth: Take the average and divide the summed absolute error by the number of samples , get the mean absolute error;
[0086] Evaluate the accuracy and performance of the model. The accuracy indicator uses Recall, which is the proportion of correctly predicted key points in the actual data.
[0087] The error marking module marks the error between the model output and the real data, marks the key points that exceed the error threshold, and dynamically adjusts the error threshold based on the data;
[0088] The review module reviews the error marking results to determine whether the model accuracy meets the requirements;
[0089] Conduct internal review based on preset module presets;
[0090] The preset module sets the audit content for the audit module and performs real-time audit processing based on the marking of error data;
[0091] When the accuracy meets the requirements, it is transmitted to the No. 1 classification port;
[0092] When the accuracy is not enough, it is transmitted to the binary classification port.
[0093] In some embodiments, the collating transmission module audits the transmitted data based on a preset, classifies the data through the error marking module, and transmits the data to the corresponding port;
[0094] The sorting and transmission module includes a first-classification port and a second-classification port;
[0095] A classification port transmits data based on the sorting transmission module and transmits it to the optimization module;
[0096] The binary classification port transmits data based on the sorting transmission module and transmits it to the posture estimation module;
[0097] The optimization module performs preliminary optimization on the model prediction results, reduces data jitter and noise, and transmits the optimized data to the cloud integration module for overall optimization;
[0098] Smooth the model prediction results to reduce data jitter and noise, and use the Kalman filter to smooth the data;
[0099] Perform preliminary data optimization to improve data stability, use filters to remove noise from the data, normalize the data to ensure consistency in the data range of each dimension, and transmit the data to the cloud via the network;
[0100] The cloud integration module uses cloud optimization algorithms to accelerate the optimization process based on the preliminary optimization data transmitted by the optimization module;
[0101] Perform overall optimization in the cloud, further improving the model's prediction accuracy through large-scale computing resources and Bayesian optimization;
[0102] Deserialize the received serialized data based on the preliminary optimization data transmitted from the optimization module;
[0103] Further optimize the data through large-scale computing resources and advanced optimization algorithms, and use Bayesian optimization to adjust model hyperparameters;
[0104] The compensation module performs error compensation on the data optimized by the cloud integration module, receives the optimized data transmitted from the cloud integration module, performs error compensation on the optimized data, reduces the prediction error, and performs dynamic error compensation on the time series data;
[0105] Dynamic error compensation, using a Kalman filter to model the error and update it in real time;
[0106] According to the dynamic error model, the predicted data is adjusted in real time to perform error compensation. The error compensation amount is iteratively updated at each time step to gradually reduce the error.
[0107] In some specific embodiments, the user interaction module is used to provide an intuitive and natural interaction method so that the user can perform various operations in the virtual environment;
[0108] Provides a calibration tool for user full-body tracking, helping users adjust and calibrate the position and angle of the sensor. It allows users to adjust tracking accuracy and sensitivity settings to balance tracking accuracy and system performance. It also provides a virtual posture adjustment interface, allowing users to manually fine-tune the posture of the virtual character.
[0109] Displays system attitude and key parameters, including sensor attitude, tracking accuracy, and network latency, provides user guidance and help information, and records user feedback and operation history.
[0110] In some specific embodiments, the feedback module mainly provides real-time data feedback and user interaction feedback, displays the user's three-dimensional skeleton model and key point positions in real time on the user interface, and reminds the user through visual prompts when the system detects errors or anomalies;
[0111] When a tactile device is used, tactile feedback is provided, and a vibration prompt is given when the user's action exceeds the predetermined range or an error occurs;
[0112] Display real-time data panels in the user interface, providing key information such as skeletal keypoint positions, pose estimation accuracy, motion capture error, and regularly generate statistical reports on user behavior and system performance;
[0113] Data mapping module: The main function of the data mapping module is to map the processed 3D skeleton data to the virtual character in the metaverse to achieve full-body tracking effect;
[0114] Data mapping includes data conversion, data transmission and virtual role mapping;
[0115] First, data conversion, by converting the processed 3D skeletal data into the coordinate system required by the Metaverse, and converting the data format into a compatible format according to the requirements of the Metaverse platform;
[0116] Second, data transmission: data is transmitted to the Metaverse platform in real time through the network interface, and a cache mechanism is used during the transmission process to ensure the continuity and real-time nature of the data.
[0117] Third, virtual character mapping, which binds three-dimensional skeleton data to the virtual character's skeletal system, drives the character's movements, and updates the virtual character's animation according to real-time data to ensure synchronization and smoothness of the movements.
[0118] The interactive method provided in this application further comprises the following steps:
[0119] S100 collects and pre-processes data based on the headset, controller, sensors, and additional cameras;
[0120] S200, performing posture estimation, capture, space establishment, model construction and data fusion processing based on pre-processed data;
[0121] S300, predict and improve response based on fusion data and historical action data, and correct model accuracy;
[0122] S400, based on correction data identification, comparison, error marking and review and arrangement;
[0123] S500, returning or transmitting based on the collated data and the audited data content;
[0124] S600: Optimize and compensate based on the transmitted data, and transmit to the user interaction module;
[0125] 700. The user interaction module outputs the transmission data to the data mapping transmission and the feedback transmission respectively based on the transmission data.
[0126] The interactive system provided by the present application collects data based on a head display, handle, sensor and additional camera, and performs preprocessing, performs posture estimation, capture, space establishment, model construction and data fusion processing based on the preprocessed data, predicts and improves response based on the fused data and historical motion data, and corrects the model accuracy, identifies, compares, marks errors and reviews and organizes the corrected data, returns or transmits the organized data based on the content of the reviewed data, optimizes and compensates based on the transmitted data, and transmits it to the user interaction module, which outputs the transmitted data to the data mapping transmission and feedback transmission respectively, uses posture estimation and capture, and through data identification and comparison, finally optimizes and compensates through organization and outputs it to the data mapping module for transmission, thereby achieving a full-body tracking effect using sensors, cameras, handles and head displays. BRIEF DESCRIPTION OF THE DRAWINGS
[0127] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0128] Figure 1 An overall flow chart of an interactive method and system provided in an embodiment of the present application.
[0129] Figure 2 A flowchart of a posture estimation method and system provided in an embodiment of the present application.
[0130] Figure 3 A data recognition flow chart of an interactive method and system provided in an embodiment of the present application.
[0131] Figure 4 A flowchart of an interactive method and system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0132] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0133] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0134] Please refer to Figure 1 、 Figure 2 and Figure 3 , which shows the process of an embodiment of the interactive system according to the present disclosure.
[0135] like Figure 1 、 Figure 2 and Figure 3 As shown, the interactive system includes the following modules:
[0136] Data collection module, which collects user action and environmental data;
[0137] Preprocessing module, which preprocesses the collected data to improve data quality;
[0138] The posture estimation module detects the user's key point positions based on the data and generates a 3D skeleton model;
[0139] Data identification module, based on data and modeling model identification;
[0140] The transmission module includes a first-class port and a second-class port, and uses the first-class port or the second-class port for transmission according to the transmission data;
[0141] Optimization module, which performs preliminary optimization based on model parameters;
[0142] Cloud integration module, overall optimization based on the model;
[0143] Compensation module, based on model error compensation after optimization;
[0144] User interaction module, which provides an interactive interface and enables natural interaction through 3D models;
[0145] Data mapping module, which maps the user's 3D posture and movements into the virtual environment;
[0146] Feedback module, providing visual and tactile feedback.
[0147] Among them, the data acquisition module and pre-processing module in the above content are specifically:
[0148] The data acquisition module is used to collect user motion and environmental data. It collects user current motion data based on the headset, sensors, and controllers, and collects user current motion data based on additional cameras and the headset camera.
[0149] VR headset camera, with a built-in high-frame-rate camera, captures real-time video streams of the user's upper body and hands;
[0150] Handle sensor: built-in inertial measurement unit, real-time transmission of hand motion data;
[0151] Additional cameras: Wide-angle lenses and high-frame-rate cameras are installed in the four corners of the room to capture full-body video streams;
[0152] The preprocessing module preprocesses the collected data to improve data quality, enhances the collected image data, and ensures data synchronization and spatial calibration;
[0153] Based on image processing, including image denoising, image contrast adjustment, image cropping, image scaling and data enhancement;
[0154] Image denoising: Use a Gaussian filter to smooth the image and remove high-frequency noise, and a median filter to remove impulse noise in the image;
[0155] Image contrast adjustment: Improve image contrast by equalizing the image histogram, and use adaptive histogram equalization to enhance contrast in local areas to avoid over-enhancement;
[0156] Image cropping: Cropping based on preset areas, such as key areas like the head and hands, and dynamically adjusting the cropping area based on the detected target position;
[0157] Image Scaling: Use bilinear interpolation to maintain image quality during scaling, and use bicubic interpolation for higher accuracy requirements;
[0158] Data augmentation: Randomly rotate the image by a certain angle, flip the image horizontally or vertically, randomly scale the image, and randomly change the brightness, contrast, and saturation of the image;
[0159] Data synchronization and calibration, including time synchronization, spatial calibration, data standardization and data quality testing;
[0160] Time synchronization: Ensure the temporal consistency of data collected by different sensors by adding a timestamp to each data sample, aligning data from different sensors based on the timestamp, and interpolating data with incomplete timestamp matching to ensure data synchronization;
[0161] Spatial calibration: Ensures spatial consistency between different sensors, especially in the fusion process of multiple cameras and IMU sensors. Use calibration plates to calibrate the internal and external parameters of the camera to determine the intrinsic and external parameters of the camera. In a multi-camera system, hand-eye calibration is used to determine the position of the camera relative to the robot or fixed structure. Static and dynamic calibration of IMU sensors is performed to determine the sensor bias and scale factor.
[0162] Data standardization: Unify data formats and scales to ensure that data from different sources are processed at the same scale, normalize data to a unified range, eliminate dimensional differences, and standardize data to a mean of 0 and a variance of 1 to eliminate mean and variance differences between different data sources;
[0163] Data quality detection: Detect and filter low-quality data, ensure the reliability of input data, use statistical methods or machine learning models to detect and filter abnormal data, check data integrity, and filter out data samples that are missing key parts.
[0164] Among them, in the above content, the posture estimation module includes the following modules:
[0165] The posture estimation module detects the user's key point positions based on the data and generates a 3D skeleton model;
[0166] The motion capture module fuses data from different sensors to generate the user's three-dimensional full-body motion;
[0167] The space building module generates an adaptive 3D virtual space based on the user's 3D skeleton model and motion data;
[0168] Model building module, which integrates multiple data and builds a complete virtual environment model;
[0169] Data fusion module, which fuses data from multiple sensors to generate consistent 3D model data;
[0170] The action prediction module predicts the user's next action based on the fused 3D model and by analyzing historical data;
[0171] The model correction module is used to correct errors and inconsistencies that occur during the data fusion process.
[0172] Among them, the posture estimation module, motion capture module, space establishment module, model construction module, data fusion module, motion prediction module and model correction module are specifically:
[0173] The posture estimation module detects the user's key point positions based on data and generates a 3D skeleton model. It uses existing public datasets to collect data that meets the specific needs of the system.
[0174] Label the key points of the human body for the image data, use the pre-trained model for preliminary labeling, and then manually correct it;
[0175] Generate a 3D skeleton model using a 2D pose estimation model and a 3D pose estimation model;
[0176] The 2D pose estimation model uses HRNet to process images with rich details, and the 3D pose estimation model uses VideoPose3D to convert 2D key point sequences into 3D poses;
[0177] HRNet receives an input image and converts it into a format suitable for network processing through a series of preprocessing steps (such as normalization, scaling, and cropping). It then extracts initial features through several 3×3 convolutional layers and uses the ReLU activation function for nonlinear transformation. The feature maps generated by the convolutional layers are used to maintain high resolution and enter the high-resolution branch for processing. The network passes the feature maps to multiple parallel convolutional layers with different resolutions. By adjusting the stride and pooling operations of the convolutional layers, feature maps of different resolutions are generated. At the end of each stage, HRNet fuses the feature maps of different resolutions through feature fusion units. These fusion units use interpolation upsampling and convolution downsampling methods to align and fuse the multi-resolution feature maps together, thereby maintaining high-resolution information and enhancing feature expression capabilities.
[0178] HRNet continuously enhances and fuses multi-resolution features through repeated operations of multiple fusion units, which include exchange units and multi-layer convolution;
[0179] Exchange unit: At each stage, feature maps are fused through the exchange unit to form a new multi-resolution feature map;
[0180] Multi-layer convolution: Through multi-layer convolution operations, features are further extracted and fused;
[0181] The output of HRNet is a keypoint heatmap, where each keypoint corresponds to a channel. Each pixel value in the heatmap represents the probability of that location being the corresponding keypoint. In the final layer, a convolutional layer is used to convert the high-resolution feature map into a multi-channel heatmap, where each channel corresponds to a keypoint. The Softmax activation function is used to normalize the heatmap to represent a probability distribution.
[0182] In each heat map channel, the maximum value is found. This position is the predicted key point position. Gaussian filtering is used to refine the heat map to obtain sub-pixel key point positions. The key point positions are then smoothed to reduce jitter and noise. The key point positions are constrained according to the human skeletal structure to ensure that the generated skeletal data is reasonable.
[0183] VideoPose3D receives a sequence of 2D key points within a time window as input. The model uses a fixed-length time window at each time step to capture the changes in 2D key points over a period of time. The input format is ,in is the time window length, is the number of key points, each key point contains 2 coordinates ;
[0184] A 3D convolution layer is used to perform convolution operations in the time dimension to extract spatiotemporal features. The 3D convolution calculation involves sliding filters in the time and space dimensions to generate feature maps. Residual blocks are used to enhance the network's expressiveness and training depth to prevent gradient disappearance.
[0185] Through multi-layer 3D convolution and pooling operations, the extracted spatiotemporal features are input into the fully connected layer or regression layer to predict the position of 3D key points. The output format is ,in is the time window length, is the number of key points, each key point contains 3 coordinates ;
[0186] The mean square error is used as the loss function to minimize the error between the predicted 3D key points and the real 3D key points. The predicted 3D key point sequence is smoothed and filtered to reduce the jitter and noise of the prediction results. The position of the 3D key points is constrained according to the human skeletal structure.
[0187] The motion capture module combines visual data and inertial sensor data to generate 3D motion capture data through 3D point cloud processing and skeletal model matching technology;
[0188] Use timestamps to align visual data and IMU data to ensure they are fused at the same time. When there is a time difference, linear interpolation is used to synchronize the data. A Kalman filter is used to predict the attitude at the next moment using the attitude transfer matrix. Combined with the measurement matrix, the predicted attitude is adjusted using the Kalman gain to minimize the error.
[0189] Use nonlinear attitude transfer function and measurement function, perform attitude prediction and update after linearization;
[0190] The posture distribution is represented by a set of particles, which is applicable to complex nonlinear and non-Gaussian systems. The particle weights are updated according to the measurement values, and a new particle set is generated by resampling.
[0191] Fusion of visual and IMU data generates a 3D skeleton model to represent the user's posture and movements. Kalman filtering or mean filtering is used to smooth the positions of 3D key points to reduce jitter and noise. Human skeletal structure constraints are applied to ensure that the generated 3D skeleton data is reasonable.
[0192] The space building module builds an accurate and dynamic three-dimensional virtual environment by integrating multi-source data and modeling technology to adapt to the user's actual movements;
[0193] By integrating environmental perception, 3D modeling, dynamic adjustment and virtual environment optimization technologies, accurate and dynamic 3D virtual environments can be constructed, utilizing point cloud processing and registration, surface reconstruction and texture mapping, real-time updating and environmental adaptation, rendering optimization and physical interaction;
[0194] The model building module uses timestamps to align pose estimation and motion capture data, ensuring fusion at the same time point and calibrating data from different sensors;
[0195] A Kalman filter is used to fuse multi-source data to generate a consistent 3D skeleton model and motion data, and a corresponding 3D virtual environment is generated based on the user's movements and range of activity.
[0196] Adjust the size and layout of the virtual environment in real time based on user actions to ensure that the user's range of activities is fully reflected;
[0197] Dynamically generate virtual objects based on scene requirements and set physical properties for virtual objects;
[0198] Map the user's skeleton model to the virtual character's skeleton system and adjust the bone weights according to the character model;
[0199] Applying the data generated by the motion capture module to the virtual character to dynamically adjust the virtual character's posture and movements based on the user's real-time movements;
[0200] Use Unity to set physical properties for objects in the virtual environment to achieve real physical interaction;
[0201] The data fusion module uses Unity to set physical properties for objects in the virtual environment to achieve real physical interaction;
[0202] The data fusion module fuses data from different sensors to generate consistent 3D skeleton models and motion data through data synchronization, multimodal data fusion, and result optimization and post-processing steps;
[0203] Using Kalman filter, extended Kalman filter and particle filter technology to ensure the high accuracy and stability of the fused data;
[0204] The action prediction module collects a large amount of historical action data for training and prediction, processes the collected data to make it suitable for model training, selects an appropriate time window length to capture the dynamic characteristics of the action, extracts features from the time series for use in the prediction model, and uses the LSTM network to capture the long-term and short-term dependencies in the time series for action prediction. It uses stacked multi-layer LSTM units to increase the model's expressiveness, and uses a fully connected layer to convert the LSTM output into a prediction result.
[0205] The model correction module measures the error between the predicted data and the actual data, calculates the Euclidean distance of key point positions, compares time series data, measures the similarity between series, and dynamically adjusts the threshold based on historical data and the current environment;
[0206] The Kalman filter is used to correct the prediction results in real time, the attitude transfer matrix is used to predict the attitude at the next moment, and the predicted attitude is corrected in combination with the actual measurement data.
[0207] Among them, in the above content, the data identification module includes the following modules:
[0208] Data identification module, based on data and modeling model identification, and extracts feature data from model output;
[0209] Data comparison module, which performs comparison processing based on the identified data;
[0210] Error marking module, marking the error between model output and real data;
[0211] The audit module performs audit processing based on the audit content preset in the preset module;
[0212] A preset module performs real-time review and processing based on the marking of error data.
[0213] Among them, the data identification module, data comparison module, error marking module, review module and preset module in the above content are specifically:
[0214] The data identification module, based on data and modeling model identification, prepares for subsequent error analysis and comparison by identifying and extracting characteristic data from the model output;
[0215] Based on the output data and real data after the model processing is completed, key feature points are extracted from the data, such as key points of human posture. The key points are extracted using a pre-trained posture estimation model, the key point position data is converted into vector form, and the identified feature data is annotated to provide a benchmark for subsequent use;
[0216] The data comparison module performs comparison processing based on the identified data to ensure the temporal consistency of the model output data and the real data, ensures that the predicted data and the real data are compared in the same coordinate system, and calculates the error between the model output and the real data;
[0217] MAE is used to calculate the average value of all key point errors, specifically: Public school: is the number of samples, It is The true value of the sample, It is The predicted value of the sample, Indicates the The absolute error of the sample; One: the true value and predicted values , the true value is the actual observed data, and the predicted value It is the data obtained through model prediction;
[0218] Second: absolute error. Absolute error refers to the absolute difference between the true value and the predicted value, that is, ;
[0219] Third: sum, sum the absolute errors of all samples;
[0220] Fourth: Take the average and divide the summed absolute error by the number of samples , get the mean absolute error;
[0221] Evaluate the accuracy and performance of the model. The accuracy indicator uses Recall, which is the proportion of correctly predicted key points in the actual data.
[0222] The error marking module marks the error between the model output and the real data, marks the key points that exceed the error threshold, and dynamically adjusts the error threshold based on the data;
[0223] The review module reviews the error marking results to determine whether the model accuracy meets the requirements;
[0224] Conduct internal review based on preset module presets;
[0225] The preset module sets the audit content for the audit module and performs real-time audit processing based on the marking of error data;
[0226] When the accuracy meets the requirements, it is transmitted to the No. 1 classification port;
[0227] When the accuracy is not enough, it is transmitted to the binary classification port.
[0228] Among them, the above content includes the transmission module, optimization module, cloud integration module and compensation module, specifically:
[0229] The transmission module organizes and reviews the transmitted data based on the presets, classifies the data through the error marking module, and transmits the data to the corresponding port;
[0230] The sorting and transmission module includes a first-classification port and a second-classification port;
[0231] A classification port transmits data based on the sorting transmission module and transmits it to the optimization module;
[0232] The binary classification port transmits data based on the sorting transmission module and transmits it to the posture estimation module;
[0233] The optimization module performs preliminary optimization on the model prediction results, reduces data jitter and noise, and transmits the optimized data to the cloud integration module for overall optimization;
[0234] Smooth the model prediction results to reduce data jitter and noise, and use the Kalman filter to smooth the data;
[0235] Perform preliminary data optimization to improve data stability, use filters to remove noise from the data, normalize the data to ensure consistency in the data range of each dimension, and transmit the data to the cloud via the network;
[0236] The cloud integration module uses cloud optimization algorithms to accelerate the optimization process based on the preliminary optimization data transmitted by the optimization module;
[0237] Perform overall optimization in the cloud, further improving the model's prediction accuracy through large-scale computing resources and Bayesian optimization;
[0238] Deserialize the received serialized data based on the preliminary optimization data transmitted from the optimization module;
[0239] Further optimize the data through large-scale computing resources and advanced optimization algorithms, and use Bayesian optimization to adjust model hyperparameters;
[0240] The compensation module performs error compensation on the data optimized by the cloud integration module, receives the optimized data transmitted from the cloud integration module, performs error compensation on the optimized data, reduces the prediction error, and performs dynamic error compensation on the time series data;
[0241] Dynamic error compensation, using Kalman filter to model the error and update it in real time
[0242] The prediction steps are as follows:
[0243]
[0244]
[0245] Second, the update steps are as follows:
[0246]
[0247]
[0248]
[0249] Where: It is at the moment Predicted Moment Pose estimation, is the covariance matrix of the pose estimate, is the attitude transfer matrix, is the control matrix, is the control vector, is the process noise covariance matrix, is the measurement noise covariance matrix, is the Karl-Brühl gain, is the measurement matrix, is the measured value;
[0250] According to the dynamic error model, the predicted data is adjusted in real time to perform error compensation. The error compensation amount is iteratively updated at each time step to gradually reduce the error.
[0251] Among them, in the above content, the user interaction module is specifically:
[0252] A user interaction module is used to provide an intuitive and natural interaction method so that users can perform various operations in the virtual environment;
[0253] Provides a calibration tool for user full-body tracking, helping users adjust and calibrate the position and angle of the sensor. It allows users to adjust tracking accuracy and sensitivity settings to balance tracking accuracy and system performance. It also provides a virtual posture adjustment interface, allowing users to manually fine-tune the posture of the virtual character.
[0254] Displays system attitude and key parameters, including sensor attitude, tracking accuracy, and network latency, provides user guidance and help information, and records user feedback and operation history.
[0255] 10. Among them, the feedback module and data mapping module in the above content are specifically:
[0256] Feedback module: The main function of the feedback module is to provide real-time data feedback and user interaction feedback. It displays the user's 3D skeleton model and key point positions in real time on the user interface. When the system detects an error or anomaly, it alerts the user through visual prompts.
[0257] When a tactile device is used, tactile feedback is provided, and a vibration prompt is given when the user's action exceeds the predetermined range or an error occurs;
[0258] Display real-time data panels in the user interface, providing key information such as skeletal keypoint positions, pose estimation accuracy, motion capture error, and regularly generate statistical reports on user behavior and system performance;
[0259] Data mapping module: The main function of the data mapping module is to map the processed 3D skeleton data to the virtual character in the metaverse to achieve full-body tracking effect;
[0260] Data mapping includes data conversion, data transmission and virtual role mapping;
[0261] First, data conversion, by converting the processed 3D skeletal data into the coordinate system required by the Metaverse, and converting the data format into a compatible format according to the requirements of the Metaverse platform;
[0262] Second, data transmission: data is transmitted to the Metaverse platform in real time through the network interface, and a cache mechanism is used during the transmission process to ensure the continuity and real-time nature of the data.
[0263] Third, virtual character mapping, which binds three-dimensional skeleton data to the virtual character's skeletal system, drives the character's movements, and updates the virtual character's animation according to real-time data to ensure synchronization and smoothness of the movements.
[0264] like Figure 4 As shown, the interaction method includes the following steps:
[0265] Step 100: Collect data based on the head display, controller, sensor, and additional camera, and perform pre-processing;
[0266] Step 200, performing posture estimation, capture, space establishment, model construction and data fusion processing based on pre-processed data;
[0267] Step 300: Predict and improve response based on fusion data and historical action data, and correct model accuracy;
[0268] Step 400, based on the correction data identification, comparison, error marking and review and arrangement;
[0269] Step 500, based on sorting data and reviewing the data content, returning or transmitting;
[0270] Step 600, performing optimization and compensation based on the transmitted data, and transmitting the data to the user interaction module;
[0271] In step 700 , the user interaction module outputs the transmission data to the data mapping transmission and the feedback transmission respectively based on the transmission data. Example 1
[0272] like Figure 1 As shown, the data acquisition module collects data from the head display, handle, sensor and additional camera, and transmits it to the preprocessing module. The preprocessing module preprocesses the collected data and transmits it to the posture estimation module. The posture estimation module processes the data and transmits it to the data recognition module. The data recognition module processes the data and transmits it to the sorting and transmission module. The sorting and transmission module uses a single-classification port or a binary classification port for data transmission based on the data content.
[0273] The data content information of the data identification module is arranged and transmitted to the optimization module when a classification port is used for transmission. The data is optimized by the optimization module and transmitted to the cloud integration module. The cloud integration module optimizes the optimized data and transmits it to the compensation module. The compensation module communicates with the compensation module to accelerate the accuracy compensation. After the compensation is completed, the data is transmitted to the user interaction module. The user interaction module transmits the data to the feedback module and the data mapping module for output.
[0274] The sorting and transmission module is based on the data content information at the data identification module. When a binary port transmission is adopted, the data will be transmitted back to the posture estimation module, and the data will be processed again by the posture estimation module and transmitted to the data identification module. The data will be processed by the data identification module and transmitted to the sorting and transmission module for redistribution and transmission. Example 2
[0275] like Figure 2 As shown, based on the pre-processed transmitted data, the posture estimation module processes the data and transmits it to the motion capture module. The motion capture module captures and processes the data and transmits it to the space establishment module. The space establishment module processes the data and establishes the corresponding space and transmits it to the model construction module. The model construction module establishes a model based on the data and transmits it to the data fusion module. The data fusion module performs fusion processing based on multiple data and transmits it to the action prediction module. The action prediction module processes the data and predicts the user's next action and transmits the data to the model correction module. The model correction module processes and corrects the data and transmits it to the data recognition module. The data recognition module processes the data and transmits it to the sorting and transmission module. The sorting and transmission module sorts the data and distributes it for transmission. Example 3
[0276] like Figure 3 As shown, the data identification module performs identification processing based on the transmitted data and transmits it to the data comparison module. The data comparison module performs comparative analysis based on the data and transmits it to the error marking module. The error marking module marks the error points of the data and transmits it to the review module. The review module performs review based on the review content set by the preset module and transmits it to the sorting and transmission module for distribution and transmission.
[0277] The electronic devices in the embodiments of the present disclosure may include but are not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0278] For example, an embodiment of the present disclosure includes a computer program product including a computer program carried on a non-transitory computer-readable medium, the computer program including program codes for executing the method shown in the flowchart.
[0279] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An interactive system, characterized in that: Includes the following modules: Data collection module, which collects user action and environmental data; Preprocessing module, which preprocesses the collected data to improve data quality; The posture estimation module detects the user's key point positions based on the data and generates a 3D skeleton model; Data identification module, based on data and modeling model identification; The transmission module includes a first-class port and a second-class port, and uses the first-class port or the second-class port for transmission according to the transmission data; Optimization module, which performs preliminary optimization based on model parameters; Cloud integration module, overall optimization based on the model; Compensation module, based on model error compensation after optimization; User interaction module, which provides an interactive interface and enables natural interaction through 3D models; Data mapping module, which maps the user's 3D posture and movements into the virtual environment; Feedback module, providing visual and tactile feedback; Organize the transmission module, optimization module, cloud integration module and compensation module, specifically: The transmission module organizes and reviews the transmitted data based on the presets, classifies the data through the error marking module, and transmits the data to the corresponding port; The sorting and transmission module includes a first-classification port and a second-classification port; A classification port transmits data based on the sorting transmission module and transmits it to the optimization module; The binary classification port transmits data based on the sorting transmission module and transmits it to the posture estimation module; The optimization module performs preliminary optimization on the model prediction results, reduces data jitter and noise, and transmits the optimized data to the cloud integration module for overall optimization; Smooth the model prediction results to reduce data jitter and noise, and use the Kalman filter to smooth the data; Perform preliminary data optimization to improve data stability, use filters to remove noise from the data, normalize the data to ensure consistency in the data range of each dimension, and transmit the data to the cloud via the network; The cloud integration module uses cloud optimization algorithms to accelerate the optimization process based on the preliminary optimization data transmitted by the optimization module; Perform overall optimization in the cloud, further improving the model's prediction accuracy through large-scale computing resources and Bayesian optimization; Deserialize the received serialized data based on the preliminary optimization data transmitted from the optimization module; Further optimize the data through large-scale computing resources and advanced optimization algorithms, and use Bayesian optimization to adjust model hyperparameters; The compensation module performs error compensation on the data optimized by the cloud integration module, receives the optimized data transmitted from the cloud integration module, performs error compensation on the optimized data, reduces the prediction error, and performs dynamic error compensation on the time series data; Dynamic error compensation, using a Kalman filter to model the error and update it in real time; According to the dynamic error model, the predicted data is adjusted in real time to perform error compensation. The error compensation amount is iteratively updated at each time step to gradually reduce the error.
2. The interactive system according to claim 1, wherein: Data acquisition module and preprocessing module, specifically: The data acquisition module is used to collect user motion and environmental data. It collects user current motion data based on the headset, sensors, and controllers, and collects user current motion data based on additional cameras and the headset camera. VR headset camera, with a built-in high-frame-rate camera, captures real-time video streams of the user's upper body and hands; Handle sensor: built-in inertial measurement unit, real-time transmission of hand motion data; Additional cameras: Wide-angle lenses and high-frame-rate cameras are installed in the four corners of the room to capture full-body video streams; The preprocessing module preprocesses the collected data to improve data quality, enhances the collected image data, and ensures data synchronization and spatial calibration; Based on image processing, including image denoising, image contrast adjustment, image cropping, image scaling and data enhancement; Image denoising: Use a Gaussian filter to smooth the image and remove high-frequency noise, and a median filter to remove impulse noise in the image; Image contrast adjustment: Improve image contrast by equalizing the image histogram, and use adaptive histogram equalization to enhance contrast in local areas to avoid over-enhancement; Image cropping: Cropping is performed based on preset areas, including key areas of the head and hands, and the cropping area is dynamically adjusted based on the detected target position; Image Scaling: Use bilinear interpolation to maintain image quality during scaling, and use bicubic interpolation for higher accuracy requirements; Data augmentation: Randomly rotate the image by a certain angle, flip the image horizontally or vertically, randomly scale the image, and randomly change the brightness, contrast, and saturation of the image; Data synchronization and calibration, including time synchronization, spatial calibration, data standardization and data quality testing; Time synchronization: Ensure the temporal consistency of data collected by different sensors by adding a timestamp to each data sample, aligning data from different sensors based on the timestamp, and interpolating data with incomplete timestamp matching to ensure data synchronization; Spatial calibration: Ensures spatial consistency between different sensors. During the fusion of multiple cameras and IMU sensors, calibration plates are used to calibrate the camera's internal and external parameters to determine the camera's intrinsic and extrinsic parameters. In a multi-camera system, hand-eye calibration is used to determine the camera's position relative to the robot or fixed structure. Static and dynamic calibration of IMU sensors is performed to determine the sensor's bias and scale factor. Data standardization: Unify data formats and scales to ensure that data from different sources are processed at the same scale, normalize data to a unified range, eliminate dimensional differences, and standardize data to a mean of 0 and a variance of 1 to eliminate mean and variance differences between different data sources; Data quality detection: Detect and filter low-quality data, ensure the reliability of input data, use statistical methods or machine learning models to detect and filter abnormal data, check data integrity, and filter out data samples that are missing key parts.
3. The interactive system according to claim 1, wherein: The posture estimation module includes the following modules: The posture estimation module detects the user's key point positions based on the data and generates a 3D skeleton model; The motion capture module fuses data from different sensors to generate the user's three-dimensional full-body motion; The space building module generates an adaptive 3D virtual space based on the user's 3D skeleton model and motion data; Model building module, which integrates multiple data and builds a complete virtual environment model; Data fusion module, which fuses data from multiple sensors to generate consistent 3D model data; The action prediction module predicts the user's next action based on the fused 3D model and by analyzing historical data; The model correction module is used to correct errors and inconsistencies that occur during the data fusion process.
4. The interactive system according to claim 3, wherein: Posture estimation module, motion capture module, space establishment module, model construction module, data fusion module, motion prediction module and model correction module, specifically: The posture estimation module detects the user's key point positions based on data and generates a 3D skeleton model. It uses existing public datasets to collect data that meets the specific needs of the system. Label the key points of the human body for the image data, use the pre-trained model for preliminary labeling, and then manually correct it; Generate a 3D skeleton model using a 2D pose estimation model and a 3D pose estimation model; The 2D pose estimation model uses HRNet to process images with rich details, and the 3D pose estimation model uses VideoPose3D to convert 2D key point sequences into 3D poses; HRNet receives the input image and converts it into a format suitable for network processing through a series of preprocessing steps. It extracts initial features through several 3×3 convolutional layers and uses the ReLU activation function for nonlinear transformation. The feature map generated by the convolutional layer is used to maintain high resolution and enters the high-resolution branch for processing. The network passes the feature map to multiple parallel convolutional layers with different resolutions. By adjusting the stride and pooling operation of the convolutional layer, feature maps of different resolutions are generated. At the end of each stage, HRNet fuses the feature maps of different resolutions through the feature fusion unit. These fusion units use interpolation upsampling and convolution downsampling methods to align and fuse multi-resolution feature maps, thereby maintaining high-resolution information and enhancing feature expression capabilities; HRNet continuously enhances and fuses multi-resolution features through repeated operations of multiple fusion units, which include exchange units and multi-layer convolution; Exchange unit: At each stage, feature maps are fused through the exchange unit to form a new multi-resolution feature map; Multi-layer convolution: Through multi-layer convolution operations, features are further extracted and fused; The output of HRNet is a keypoint heatmap, where each keypoint corresponds to a channel; each pixel value in the heatmap represents the probability of that location being the corresponding keypoint. In the last layer, a convolutional layer is used to convert the high-resolution feature map into a heatmap with multiple channels, where each channel corresponds to a keypoint. The Softmax activation function is used to normalize the heatmap to represent a probability distribution. In each heat map channel, the maximum value is found. This position is the predicted key point position. Gaussian filtering is used to refine the heat map to obtain sub-pixel key point positions. The key point positions are then smoothed to reduce jitter and noise. The key point positions are constrained according to the human skeletal structure to ensure that the generated skeletal data is reasonable. VideoPose3D receives a sequence of 2D key points within a time window as input. The model uses a fixed-length time window at each time step to capture the changes in 2D key points over a period of time. The input format is ,in is the time window length, is the number of key points, each key point contains 2 coordinates ; A 3D convolution layer is used to perform convolution operations in the time dimension to extract spatiotemporal features. The 3D convolution calculation involves sliding filters in the time and space dimensions to generate feature maps. Residual blocks are used to enhance the network's expressiveness and training depth to prevent gradient disappearance. Through multi-layer 3D convolution and pooling operations, the extracted spatiotemporal features are input into the fully connected layer or regression layer to predict the position of 3D key points. The output format is ,in is the time window length, is the number of key points, each key point contains 3 coordinates ; The mean square error is used as the loss function to minimize the error between the predicted 3D key points and the real 3D key points. The predicted 3D key point sequence is smoothed and filtered to reduce the jitter and noise of the prediction results. The position of the 3D key points is constrained according to the human skeletal structure. The motion capture module combines visual data and inertial sensor data to generate 3D motion capture data through 3D point cloud processing and skeletal model matching technology; Use timestamps to align visual data and IMU data to ensure they are fused at the same time. When there is a time difference, linear interpolation is used to synchronize the data. A Kalman filter is used to predict the attitude at the next moment using the attitude transfer matrix. Combined with the measurement matrix, the predicted attitude is adjusted using the Kalman gain to minimize the error. Use nonlinear attitude transfer function and measurement function, perform attitude prediction and update after linearization; The posture distribution is represented by a set of particles, which is applicable to complex nonlinear and non-Gaussian systems. The particle weights are updated according to the measurement values, and a new particle set is generated by resampling. Fusion of visual and IMU data generates a 3D skeleton model to represent the user's posture and movements. Kalman filtering or mean filtering is used to smooth the positions of 3D key points to reduce jitter and noise. Human skeletal structure constraints are applied to ensure that the generated 3D skeleton data is reasonable. The space building module builds an accurate and dynamic three-dimensional virtual environment by integrating multi-source data and modeling technology to adapt to the user's actual movements; By integrating environmental perception, 3D modeling, dynamic adjustment and virtual environment optimization technologies, accurate and dynamic 3D virtual environments can be constructed, utilizing point cloud processing and registration, surface reconstruction and texture mapping, real-time updating and environmental adaptation, rendering optimization and physical interaction; The model building module uses timestamps to align pose estimation and motion capture data, ensuring fusion at the same time point and calibrating data from different sensors; A Kalman filter is used to fuse multi-source data to generate a consistent 3D skeleton model and motion data, and a corresponding 3D virtual environment is generated based on the user's movements and range of activity. Adjust the size and layout of the virtual environment in real time based on user actions to ensure that the user's range of activities is fully reflected; Dynamically generate virtual objects based on scene requirements and set physical properties for virtual objects; Map the user's skeleton model to the virtual character's skeleton system and adjust the bone weights according to the character model; Applying the data generated by the motion capture module to the virtual character to dynamically adjust the virtual character's posture and movements based on the user's real-time movements; Use Unity to set physical properties for objects in the virtual environment to achieve real physical interaction; The data fusion module uses Unity to set physical properties for objects in the virtual environment to achieve real physical interaction; The data fusion module fuses data from different sensors to generate consistent 3D skeleton models and motion data through data synchronization, multimodal data fusion, and result optimization and post-processing steps; Using Kalman filter, extended Kalman filter and particle filter technology to ensure the high accuracy and stability of the fused data; The action prediction module collects a large amount of historical action data for training and prediction, processes the collected data to make it suitable for model training, selects an appropriate time window length to capture the dynamic characteristics of the action, extracts features from the time series for use in the prediction model, and uses the LSTM network to capture the long-term and short-term dependencies in the time series for action prediction. It uses stacked multi-layer LSTM units to increase the model's expressiveness, and uses a fully connected layer to convert the LSTM output into a prediction result. The model correction module measures the error between the predicted data and the actual data, calculates the Euclidean distance of key point positions, compares time series data, measures the similarity between series, and dynamically adjusts the threshold based on historical data and the current environment; The Kalman filter is used to correct the prediction results in real time, the attitude transfer matrix is used to predict the attitude at the next moment, and the predicted attitude is corrected in combination with the actual measurement data.
5. The interactive system according to claim 1, wherein: The data identification module includes the following modules: Data identification module, based on data and modeling model identification, and extracts feature data from model output; Data comparison module, which performs comparison processing based on the identified data; Error marking module, marking the error between model output and real data; The audit module performs audit processing based on the audit content preset in the preset module; A preset module performs real-time review and processing based on the marking of error data.
6. The interactive system according to claim 5, wherein: The data identification module, data comparison module, error marking module, review module and preset module are as follows: The data identification module, based on data and modeling model identification, prepares for subsequent error analysis and comparison by identifying and extracting characteristic data from the model output; Based on the output data and real data after the model processing is completed, key feature points are extracted from the data, such as key points of human posture. The key points are extracted using a pre-trained posture estimation model, the key point position data is converted into vector form, and the identified feature data is annotated to provide a benchmark for subsequent use; The data comparison module performs comparison processing based on the identified data to ensure the temporal consistency of the model output data and the real data, ensures that the predicted data and the real data are compared in the same coordinate system, and calculates the error between the model output and the real data; MAE is used to calculate the average error of all key points; Evaluate the accuracy and performance of the model. The accuracy indicator uses Recall, which is the proportion of correctly predicted key points in the actual data. The error marking module marks the error between the model output and the real data, marks the key points that exceed the error threshold, and dynamically adjusts the error threshold based on the data; The review module reviews the error marking results to determine whether the model accuracy meets the requirements; Conduct internal review based on preset module presets; The preset module sets the audit content for the audit module and performs real-time audit processing based on the marking of error data; When the accuracy meets the requirements, it is transmitted to the No. 1 classification port; When the accuracy is not enough, it is transmitted to the binary classification port.
7. The interactive system according to claim 1, wherein: User interaction module, specifically: A user interaction module is used to provide an intuitive and natural interaction method so that users can perform various operations in the virtual environment; Provides a calibration tool for user full-body tracking, helping users adjust and calibrate the position and angle of the sensor. It allows users to adjust tracking accuracy and sensitivity settings to balance tracking accuracy and system performance. It also provides a virtual posture adjustment interface, allowing users to manually fine-tune the posture of the virtual character. Displays system attitude and key parameters, including sensor attitude, tracking accuracy, and network latency, provides user guidance and help information, and records user feedback and operation history.
8. The interactive system according to claim 1, wherein: Feedback module and data mapping module, specifically: Feedback module: The main function of the feedback module is to provide real-time data feedback and user interaction feedback. It displays the user's 3D skeleton model and key point positions in real time on the user interface. When the system detects an error or anomaly, it alerts the user through visual prompts. When a tactile device is used, tactile feedback is provided, and a vibration prompt is given when the user's action exceeds the predetermined range or an error occurs; Display real-time data panels in the user interface, providing key information such as skeletal keypoint positions, pose estimation accuracy, motion capture error, and regularly generate statistical reports on user behavior and system performance; Data mapping module: The main function of the data mapping module is to map the processed 3D skeleton data to the virtual character in the metaverse to achieve full-body tracking effect; Data mapping includes data conversion, data transmission and virtual role mapping; First, data conversion, by converting the processed 3D skeletal data into the coordinate system required by the Metaverse, and converting the data format into a compatible format according to the requirements of the Metaverse platform; Second, data transmission: data is transmitted to the Metaverse platform in real time through the network interface, and a cache mechanism is used during the transmission process to ensure the continuity and real-time nature of the data. Third, virtual character mapping, which binds three-dimensional skeleton data to the virtual character's skeletal system, drives the character's movements, and updates the virtual character's animation according to real-time data to ensure synchronization and smoothness of the movements.
9. An interactive method, which is implemented based on the system according to claim 1, characterized in that: The steps include: S100 collects and pre-processes data based on the headset, controller, sensors, and additional cameras; S200, performing posture estimation, capture, space establishment, model construction and data fusion processing based on pre-processed data; S300, predict and improve response based on fusion data and historical action data, and correct model accuracy; S400, based on correction data identification, comparison, error marking and review and arrangement; S500, returning or transmitting based on the collated data and the audited data content; S600: Optimize and compensate based on the transmitted data, and transmit to the user interaction module; S700: The user interaction module outputs the transmission data to the data mapping transmission and the feedback transmission respectively based on the transmission data.
Citation Information
Patent Citations
Human motion tracing method based on variable structure multi-model
CN101894278A
Hand virtual-real interaction system
CN114756130A