An experimental simulation method and system based on MediaPipe and virtual reality
Through the virtual reality technology combined with the Unity engine and MediaPipe model, the problem of device dependence and operation difficulties in traditional virtual experiments is solved, and stable gesture recognition and low-cost virtual experiment simulation are realized under different lighting conditions.
Patent Information
- Application Number
- CN202510599815.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-12
AI Technical Summary
Traditional virtual experiments require special wearable devices or handles, which are difficult to operate and have high equipment costs, making it difficult to meet the needs of large-scale teaching.
Virtual experimental scenes are built through the Unity engine, the MediaPipe model is used to identify hand key points, and gesture images are processed through three-dimensional calibration and adaptive histogram equalization algorithm, and hand movements are simulated in combination with the UDP communication protocol to reduce finger occlusion and improve gesture recognition stability.
Steady identification of gesture features under different lighting conditions reduces dependence on special equipment, improves operational convenience and reduces costs, and adapts to large-scale teaching needs.
Smart Images

Figure CN120122831B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly to an experimental simulation method and system based on MediaPipe and virtual reality. Background Art
[0002] Under the traditional experimental teaching mode, many thorny problems will emerge when students conduct actual operation experiments in person. For example, from the perspective of economic cost, during the operation process, due to the unfamiliarity with the experimental procedures and inaccurate techniques of students, it is extremely easy to cause damage and additional losses to the experimental equipment. From the safety aspect, students are at risk of potential safety accidents due to their unfamiliar operation of the equipment. Moreover, affected by the limitation of the site space size in traditional experimental teaching, it is difficult to accommodate a large number of students to conduct experiments simultaneously, and the number of experimental equipment also cannot meet the large-scale teaching needs. Often, multiple people share a set of equipment, which compresses the actual operation time of students and greatly reduces the teaching efficiency. Therefore, virtual experiment simulation has begun to enter the teaching field.
[0003] Current virtual experiments usually require special wearable devices or hardware such as handles to be realized. Accurately simulating hand movements or touch operations during the experimental process through sensing technology is difficult to operate, and the cost of these hardware devices is high, and they also require maintenance and installation, which is cumbersome and complex and will increase the work of the staff. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide an experimental simulation method and system based on MediaPipe and virtual reality, aiming to solve the technical problems in the prior art that traditional virtual experiments require special wearable devices or handles to be realized, and are difficult to operate and have high equipment costs.
[0005] On the one hand, the present invention provides an experimental simulation method based on MediaPipe and virtual reality, and the method includes:
[0006] Rendering a 3D model through the Unity engine to construct a virtual experimental scene and experimental equipment;
[0007] Collecting the gesture images of the user, and adjusting the mapping of the gesture images through three-dimensional calibration;
[0008] Performing spatial component conversion on the gesture images, and preprocessing the gesture images through an adaptive histogram equalization algorithm to obtain initial images;
[0009] Using a MediaPipe model to identify the hand key points in the initial images, and correcting the position coordinates of the hand key points according to the positional relationship of the hand key points;
[0010] Deeply encapsulate the data information of the hand key points and transmit it to the Unity engine for parsing through the UDP communication protocol, and simulate hand action demonstrations in the Unity engine.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the experimental simulation method based on MediaPipe and virtual reality provided by the present invention, by performing spatial component conversion on the gesture image and preprocessing the gesture image through the adaptive histogram equalization algorithm, it helps to more stably identify gesture features under different lighting conditions, improve the quality of the gesture image, lay a foundation for subsequent hand key point calculation, introduce a formula based on the relationship between adjacent key points to correct the hand key points, reduce the occurrence of finger mutual occlusion, and improve the hand key points, thus solving the technical problems in the prior art that traditional virtual experiments require special wearable devices or handles to implement, and are difficult to operate and have high equipment costs.
[0012] According to one aspect of the above technical solution, the steps of collecting the user's gesture image and adjusting the mapping of the gesture image through three-dimensional calibration specifically include:
[0013] Based on the changes in the spatial position and shooting angle of the camera, convert the three-dimensional points in the world coordinate system into three-dimensional points in the camera coordinate system. The formula is as follows:
[0014] ,
[0015] where, is the three-dimensional point data in the world coordinate system, are respectively 's three-dimensional coordinates, is the three-dimensional point in the camera coordinate system, are respectively 's three-dimensional coordinates, is the rotation matrix, is the translation vector;
[0016] Then project the conversion from the camera coordinate system to the gesture image coordinate system, which is represented by the internal parameter matrix. The formula is as follows:
[0017] ,
[0018] where, is the internal parameter matrix, , are respectively the focal lengths of the camera in the x-axis and y-axis directions, is the tilt factor.
[0019] According to one aspect of the above technical solution, the steps of performing spatial component conversion on the gesture image specifically include:
[0020] Extract the color components from the gesture image. The color components include the R component, the G component, and the B component, and their values range from 0 to 1;
[0021] Convert the color components into color space information. The formula is as follows:
[0022] ,
[0023] where V is the brightness information, C is the color difference, S is the saturation, and H is the hue.
[0024] According to one aspect of the above technical solution, the steps of using the MediaPipe model to identify the hand key points in the initial image and correcting the position coordinates of the hand key points according to the positional relationship of the hand key points specifically include:
[0025] Use the MediaPipe model to identify the hand key points in the initial image, and calculate the distance and angle between adjacent hand key points. The formula is as follows:
[0026] ,
[0027] where is the distance between adjacent hand key points, is the angle between adjacent hand key points, , are the coordinates of the hand key point axis and axis respectively, , are the coordinates of the hand key point axis and axis respectively;
[0028] When it is detected that the distance and the angle exceed the preset range, it means that the gesture is occluded, and correct the position coordinates of the hand key points.
[0029] According to one aspect of the above technical solution, the steps of calculating the key vector between the hand key point
[0030] to the hand key point and correcting the position coordinates of the hand key points when it is detected that the distance and the angle exceed the preset range specifically include: Calculate the key vector between the hand key point
[0031] ,
[0032] where is the key vector;
[0033] Construct a rotation matrix based on the said angle, and the calculation formula is as follows:
[0034] ,
[0035] wherein, is the angle difference, is the rotation matrix, is the minimum angle value, is the maximum angle value;
[0036] Based on the said rotation matrix and the said key vector, calculate the rotation vector to be rotated, and the formula is as follows:
[0037] ,
[0038] wherein, is the rotation vector;
[0039] Based on the said rotation vector, preliminarily correct the position coordinates of the said hand key point , and the formula is as follows:
[0040] ,
[0041] wherein, , are the position coordinates of the hand key point after preliminary correction, , are respectively of component, component.
[0042] According to one aspect of the above technical solution, the method further includes:
[0043] Normalize the position coordinates ( ( , )) of the hand key point
[0044] after preliminary correction to obtain the current coordinates; and Simulate the annealing process, set the initial temperature, and perform iterative cooling according to the preset cooling strategy. In each round of iteration, add random offset perturbations to the current coordinates in the
[0045] direction to obtain new coordinates;
[0046] When the preset number of iterations is reached, obtain the current coordinates and the corresponding fitness in each iteration, and calculate the hand key points by inverse operation of the current coordinates with the minimum fitness. The final corrected position coordinates of the hand key points.
[0047] According to one aspect of the above technical solution, the calculation formula of the fitness function is:
[0048] ,
[0049] where F is the fitness of the hand key points, , are weight coefficients, is the distance from the new coordinates of the hand key point to the position coordinates of the hand key point , is the angle from the new coordinates of the key point to the position coordinates of the hand key point , is the distance from the position coordinates of the hand key point in the correct gesture demonstration to the position coordinates of the hand key point , is the angle from the position coordinates of the hand key point in the correct gesture demonstration to the position coordinates of the hand key point .
[0050] According to one aspect of the above technical solution, the data information of the hand key points is deeply encapsulated and transmitted to the Unity engine for parsing through the UDP communication protocol. The steps of simulating the hand action demonstration in the Unity engine specifically include:
[0051] Deeply encapsulate the data information of the hand key points through the ROS message type, and the data information includes position coordinates and color space information;
[0052] Transmit through the UDP communication protocol, and after cyclic redundancy check, send it to the Unity engine for parsing;
[0053] The Unity engine performs graphic processing on the parsed data information, changes the virtual experimental scene and experimental equipment through matrix transformation, and simulates the hand action demonstration. The formula is as follows:
[0054] ,
[0055] where is the data information after matrix transformation, is the translation matrix, is the rotation matrix, is a scaling matrix.
[0056] According to one aspect of the above technical solution, the method further includes:
[0057] Taking each hand key point as an independent variable and other hand key points as dependent variables, introducing non-linear features, and establishing a collaborative regression model;
[0058] Solving the regression coefficients of the collaborative regression model by the least squares method, iteratively optimizing the regression coefficients, and obtaining the collaborative motion coefficients between different hand key points;
[0059] Integrating a series of obtained collaborative motion coefficients of different hand key points to obtain a collaborative matrix, and dynamically optimizing the collaborative matrix through a regularization algorithm;
[0060] According to the optimized collaborative matrix, introducing a time delay model to obtain a collaborative angle;
[0061] Assisting in correcting the position coordinates of the hand key points through the collaborative angle.
[0062] Another aspect of the present invention lies in providing an experimental simulation system based on MediaPipe and virtual reality. The experimental simulation system based on MediaPipe and virtual reality is used to implement the above experimental simulation method based on MediaPipe and virtual reality. The system includes:
[0063] A virtual scene construction module for rendering a 3D model through the Unity engine to construct a virtual experimental scene and experimental equipment;
[0064] A gesture image acquisition module for acquiring the gesture images of the user and adjusting the mapping of the gesture images through three-dimensional calibration;
[0065] An image preprocessing module for performing spatial component conversion on the gesture images and preprocessing the gesture images through an adaptive histogram equalization algorithm to obtain initial images;
[0066] A coordinate correction module for using the MediaPipe model to identify the hand key points in the initial image and correcting the position coordinates of the hand key points according to the positional relationship of the hand key points;
[0067] A simulated experiment demonstration module for deeply encapsulating the data information of the hand key points and transmitting it to the Unity engine for parsing through the UDP communication protocol, and simulating the hand motion demonstration in the Unity engine. Description of the Drawings
[0068] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:
[0069] Figure 1 It is a schematic flowchart of an experimental simulation method based on MediaPipe and virtual reality in the first embodiment of the present invention;
[0070] Figure 2 It is a structural block diagram of an experimental simulation system based on MediaPipe and virtual reality in the second embodiment of the present invention;
[0071] Explanation of the reference signs in the drawings:
[0072] Virtual scene construction module 100, gesture image acquisition module 200, image preprocessing module 300, coordinate correction module 400, simulation experiment demonstration module 500. Detailed implementation manners
[0073] To make the objectives, features, and advantages of the present invention more obvious and understandable, the following provides a detailed description of the specific implementation manners of the present invention in conjunction with the drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0074] Embodiment 1
[0075] Please refer to Figure 1 , a method for experimental simulation based on MediaPipe and virtual reality provided by the first embodiment of the present invention, the method includes steps S10 - step S14:
[0076] Step S10, rendering a 3D model through the Unity engine to construct a virtual experimental scene and experimental equipment;
[0077] Among them, the virtual experimental scene and experimental equipment are presented in a highly refined three-dimensional model form within the user's field of vision. The projection device selects a professional product with ultra-high resolution (such as reaching 4K or above) and high brightness to ensure that the model interface can maintain excellent clarity and color restoration in different lighting environments, accurately replicating the visual effects of the real scene.
[0078] Furthermore, the displayed model interface is built on the powerful graphics rendering framework of the Unity engine. With the help of the Scriptable Render Pipeline (SRP) of the Unity engine, real-time and efficient rendering operations are performed on elements such as the complex geometric structure of the 3D model, high-precision material textures, and realistic lighting effects (including physically based rendering, PBR technology for light and shadow simulation), generating an extremely realistic and highly immersive virtual scene. Therefore, the model interface integrates an advanced interactive logic architecture, and users can perform natural and smooth interactive operations with various devices in the scene through subsequent gesture recognition interaction technology.
[0079] Before the step of rendering the 3D model through the Unity engine to construct the virtual experimental scene and experimental equipment, it also includes:
[0080] Based on deep learning speech recognition of the Transformer model, start the required virtual experimental scene and experimental equipment through voice commands.
[0081] It should be noted that when the user starts the required virtual experimental scene and experimental equipment through voice commands, that is, using speech recognition technology to convert the user's speech into machine-readable commands. The speech recognition technology uses a speech recognition model based on deep learning, such as a model based on the Transformer architecture. Its core principle is to capture and analyze the features in the speech signal through the multi-head attention mechanism. This method can more accurately recognize speech commands under various accents and language habits compared to traditional speech recognition methods, greatly improving the convenience and accuracy of system startup. During the startup process, the system will automatically load relevant algorithm models and configuration files to ensure the normal operation of subsequent gesture recognition and interaction functions.
[0082] Step S11, collect the user's gesture images and adjust the mapping of the gesture images through three-dimensional calibration.
[0083] For example, wear a professional motion capture marker set for the subject to collect data of gesture images. Design a series of rich and diverse and representative hand movements.
[0084] By way of example rather than limitation, in addition to common actions such as making a fist, stretching fingers, and bending specific fingers, it also includes grasping objects of different shapes (such as circles, squares, triangles) and sizes (such as circular objects with a diameter of 2 cm - 10 cm), and performing fine gesture operations (such as pinching small objects, rotating knobs, etc.). Each action is repeated no less than 20 times to obtain sufficient and stable data samples.
[0085] Specifically, during the data acquisition process, a synchronization device is used to ensure the synchronization of the motion capture system and the video device that records the actions of the subject. The shooting angle of the video device should cover all directions of the hand to facilitate subsequent multi-dimensional verification and analysis of the data.
[0086] The collected raw joint position data is filtered to remove noise and outliers. A method combining Kalman filtering and median filtering is adopted. First, Kalman filtering is used for preliminary smoothing of the data, and then median filtering is used to remove possible remaining outliers. The filtering parameters are dynamically adjusted according to different action types and data characteristics to achieve the best filtering effect.
[0087] Specifically, based on the spatial position and the change of the shooting angle of the camera, the three-dimensional points in the world coordinate system are transformed into three-dimensional points in the camera coordinate system. The formula is as follows:
[0088] ,
[0089] where, is the three-dimensional point data in the world coordinate system, are respectively the three-dimensional coordinates of, is the three-dimensional point in the camera coordinate system, are respectively the three-dimensional coordinates of, is the rotation matrix, is the translation vector;
[0090] Then, the projection of the camera coordinate system onto the gesture image coordinate system is represented by the intrinsic matrix. The formula is as follows:
[0091] ,
[0092] where, is the intrinsic matrix, , are respectively the focal lengths of the camera in the x-axis and y-axis directions, is the tilt factor.
[0093] It should be noted that the spatial position and the shooting angle of the camera are precisely adjusted through a three-dimensional calibration algorithm based on machine vision, and combined with the triangulation principle of multi-view vision to obtain the best shooting perspective, effectively avoiding problems such as image acquisition quality degradation caused by occlusion, environmental reflection, and other factors.
[0094] By way of example and not limitation, for instance, an industrial-grade high-definition camera with high frame rate (frame rate set to 120fps or above) and low-latency characteristics is equipped to collect the user's gesture images in real time. The frame rate of the camera is optimized through strict timing control and synchronization algorithms to ensure that it can accurately capture the user's fast and subtle gesture movements.
[0095] Further, before the step of collecting the user's gesture images and adjusting the mapping of the gesture images through three-dimensional calibration, it further includes:
[0096] Based on speech synthesis technology and text synthesis technology, voice or text reminder information is generated before each experimental step.
[0097] Among them, according to the status of the current experimental step and the semantic understanding analysis result of the virtual experimental scenario, voice reminders are generated through advanced speech synthesis technology, or the corresponding operation gesture guidance information is displayed in text form using a UI system based on vector graphics rendering on the display interface. Speech synthesis adopts the end-to-end text-to-speech (TTS) technology based on deep learning, such as the advanced models like FastSpeech2 based on the Transformer architecture. This model can generate natural and fluent speech close to real human voices through a multi-modal fusion attention mechanism and an improved variant of the long short-term memory (LSTM) network. The text reminder is through the UI system of the Unity engine, using a graphic drawing algorithm based on Bezier curves and an adaptive layout algorithm to clearly and visually guide the user with icons and text to understand the current executable gesture operations, helping the user quickly understand and master the way to interact with the virtual experimental scenario and experimental equipment. These operation guidance information is dynamically updated through real-time semantic analysis and decision tree algorithms according to the state transition of different experimental steps and the dynamic changes of the virtual experimental scenario, ensuring that the user can always clearly know the operation process of the next experimental step.
[0098] Step S12, perform spatial component conversion on the gesture image, and preprocess the gesture image through an adaptive histogram equalization algorithm to obtain an initial image;
[0099] Specifically, extract color components from the gesture image, and the color components include R component, G component, and B component, and their value ranges are between 0 and 1;
[0100] Convert the color components into color space information, and the formula is as follows:
[0101] ,
[0102] Among them, V is the brightness information, C is the color difference, S is the saturation, and H is the hue.
[0103] It should be noted that after obtaining the gesture image collected by the camera, the conversion of the spatial components helps to more stably identify the gesture features under different lighting conditions.
[0104] Secondly, preprocessing operations are performed to improve the quality of the gesture image and lay a foundation for subsequent hand key point calculation, such as histogram equalization and Gaussian filtering.
[0105] Furthermore, the gesture image is preprocessed through an adaptive histogram equalization algorithm to obtain an initial image.
[0106] It should be noted that the adaptive histogram equalization algorithm divides the gesture image into multiple small blocks, performs histogram equalization on each small block separately, and at the same time limits the enhancement amplitude of the contrast to avoid noise amplification caused by over-enhancement. It can be implemented using the cv2.createCLAHE() function in OpenCV in Python.
[0107] Step S13, use the MediaPipe model to identify the hand key points in the initial image, and correct the position coordinates of the hand key points according to the positional relationship of the hand key points;
[0108] Specifically, use the MediaPipe model to identify the hand key points in the initial image, calculate the distance and angle between adjacent hand key points, and the formulas are as follows:
[0109] ,
[0110] Among them, is the distance between adjacent hand key points, is the angle between adjacent hand key points, , are the axis and axis coordinates of the hand key point respectively, , are the axis and axis coordinates of the hand key point respectively;
[0111] When it is detected that the distance and the angle exceed the preset range, it means that there is an occlusion situation for the gesture, and the position coordinates of the hand key points are corrected.
[0112] Calculate the key vector between the hand key point and the hand key point , and the formula is as follows:
[0113] ,
[0114] Among them, is the key vector;
[0115] According to the said angle, a rotation matrix is constructed, and the calculation formula is as follows:
[0116] ,
[0117] Among them, is the angle difference, is the rotation matrix, is the minimum angle value, is the maximum angle value;
[0118] Based on the said rotation matrix and the said key vector, the rotation vector required for rotation is calculated, and the formula is as follows:
[0119] ,
[0120] Among them, is the rotation vector;
[0121] Based on the said rotation vector, the position coordinates of the said hand key points are initially corrected, and the formula is as follows:
[0122] ,
[0123] Among them, , are the position coordinates of the hand key point after initial correction, , are respectively of component, component.
[0124] It should be noted that although the MediaPipe model can detect hand key points well, in the case of finger occlusion, the detection results may have errors. To solve this problem, a formula based on the relationship between adjacent key points is introduced for initial correction. The movement and posture of fingers usually follow certain physiological structures and movement laws. The occlusion situation can be judged by analyzing the relationships such as the distance and angle between adjacent key points, and initial correction can be carried out. For example, there are certain physiological limitations in the activity angles of human finger joints. The minimum angle of the metacarpophalangeal joint of the index finger and middle finger is -π / 6, and the maximum angle is π / 3. By initially correcting the finger coordinates in the occlusion situation, the accuracy of the coordinates is improved.
[0125] The said method further includes:
[0126] Normalize the position coordinates ( , , ) of the preliminarily corrected hand key points to obtain the current coordinates;
[0127] Specifically, the numerical value of the coordinate normalization process is in the range of [0, 1].
[0128] Add random offsets to the normalized position coordinates in the and directions to perturb them and obtain the current coordinates;
[0129] By way of example and not limitation, and are the random offsets in the and directions respectively, is a uniform normal distribution, , is a small positive number set according to the actual situation to control the perturbation amplitude.
[0130] ,
[0131] Simulate the annealing process, set the initial temperature, and perform iterative cooling according to the preset cooling strategy. In each round of iteration, add random offsets to the current coordinates in the and directions to perturb them and obtain new coordinates;
[0132] Among them, the preset cooling strategy can be expressed as , is a constant, slightly less than 1, is the temperature of the th round of iteration, is the temperature of the th round of iteration.
[0133] Calculate the fitness corresponding to the new coordinates and compare it with the fitness of the current coordinates, calculate the difference, and obtain the current coordinates in each round of iteration according to the difference;
[0134] Specifically, the calculation formula of the fitness function is:
[0135] ,
[0136] Among them, F is the fitness of the hand key points, , are the weight coefficients, is the new coordinate of the hand key point [[ID=A2]]to the hand key point The distance of the position coordinates is the key point of the new coordinates to the hand key point The angle of the position coordinates is the hand key point for correct gesture demonstration of the position coordinates to the hand key point The distance of the position coordinates is the hand key point for correct gesture demonstration of the position coordinates to the hand key point The angle of the position coordinates
[0137] In addition, the method further includes:
[0138] Determine whether the difference is less than a preset difference;
[0139] If so, determine the new coordinates as the current coordinates;
[0140] If not, calculate the coordinate probability of the new coordinates. When the coordinate probability exceeds the preset coordinate probability, determine the new coordinates as the current coordinates. When the coordinate probability does not exceed the preset coordinate probability, delete the new coordinates.
[0141] Among them, the coordinate probability can be expressed as: , is the coordinate probability, is the difference in the th round of iteration, is the temperature in the
[0142] When the preset number of iterations is reached, obtain the current coordinates and the corresponding fitness in each round of iteration, and calculate the final corrected position coordinates of the hand key point through inverse operation for the current coordinates with the minimum fitness
[0143] The method further includes:
[0144] Record the change trend of the fitness in the past several rounds of iteration. If the fitness decreases slowly or no longer decreases continuously for multiple rounds, terminate the iteration in advance or adjust the cooling strategy;
[0145] When the iteration is terminated in advance, obtain the current coordinates and the corresponding fitness in each round of iteration, and calculate the final corrected position coordinates of the hand key point through inverse operation for the current coordinates with the minimum fitness
[0146] To further improve the accuracy of the initial correction, a simulated annealing process, a perturbation strategy, and a hand keypoint fitness detection are added to obtain the final corrected position coordinates, thereby obtaining the most reasonable and highest-precision position coordinates.
[0147] In the actual scenario, the torsion of the metacarpophalangeal joint will produce a non-linear coupling effect. Therefore, only linear correction is not enough. To achieve non-linear correction, a collaborative regression model is introduced, specifically:
[0148] Taking each hand keypoint as an independent variable and other hand keypoints as dependent variables, non-linear features are introduced to establish a collaborative regression model;
[0149] For example, considering the non-linear features of hand joint movement, polynomial terms and interaction terms are introduced to improve the fitting accuracy of the model.
[0150] Among them, the data of each hand keypoint collected is divided into a training set, a validation set, and a test set according to a certain ratio (such as 70%, 15%, 15%). During the division process, the method of stratified sampling is adopted to ensure that the data distribution of different action types and subjects in each dataset is uniform. The data is normalized, and the data of each hand keypoint is mapped to the range of [0,1] or [-1,1] to accelerate the training speed of the neural network and improve the stability of the model. The maximum-minimum normalization method is used to determine the normalization parameters according to the maximum and minimum values of the training set data, and then the validation set and test set data are normalized according to the same parameters. To increase the diversity of the training data, data augmentation techniques are adopted.
[0151] In addition to performing operations such as random rotation, translation, and scaling on the hand keypoint data, noise interference can also be introduced, and different lighting conditions can be simulated. Through these operations, new training samples are generated to improve the robustness of the model. During the data augmentation process, the parameters of the augmentation operations are randomly controlled. For example, the rotation angle is randomly selected between -30° and 30°, and the translation distance is randomly selected between -5mm and 5mm to ensure that the generated training samples have sufficient diversity.
[0152] In addition, the number of neurons in the input layer is equal to the number of hand keypoints to ensure that all the angle information of the hand keypoints can be received. An input normalization layer is added to the input layer to further normalize the input data to improve the robustness of the model. The input normalization layer uses the method of batch normalization to normalize the input data during each training, so that the mean of the input data is 0 and the variance is 1. The middle layer contains multiple hidden layers, and the number of neurons in each hidden layer can be adjusted according to the experimental results.
[0153] Generally speaking, the number of neurons can be set in descending order, for example, 128 neurons in the first hidden layer and 64 neurons in the second hidden layer. When setting the number of neurons, a grid search method is used to optimize and find the optimal combination. Batch normalization layers and ReLU activation functions are added between each hidden layer to accelerate network convergence and introduce nonlinear features. Batch normalization layers can reduce internal covariate shift and improve network training efficiency. The ReLU activation function introduces nonlinearity, enabling the network to learn more complex features.
[0154] The number of neurons in the output layer is equal to the number of hand keypoints. The output layer accounts for nonlinear coupling effects and contains hand keypoint data. An output scaling layer is added to the output layer to restore the output data to the original hand keypoint data range. The output scaling layer performs the inverse operation based on the normalization parameters, restoring the normalized output data to the actual hand keypoint data. The mean squared error (MSE) loss function is selected to measure the difference between the hand keypoint data output by the network and the actual hand keypoint data. The MSE loss function can intuitively reflect the magnitude of the error in the network output. To improve the model's robustness to outliers, the Huber loss function is used. The Huber loss function uses the MSE loss function for small errors and the absolute value loss function for large errors, effectively reducing the impact of outliers on model training.
[0155] Solving the regression coefficient of the collaborative regression model by the least square method, iteratively optimizing the regression coefficient to obtain the collaborative motion coefficient between different hand key points;
[0156] In order to improve the accuracy of the solution, the iterative optimization method is used to update the regression coefficients multiple times.
[0157] For example, for the a-th hand key point and the b-th hand key point, the coefficient obtained by regression analysis is C ab . Summarize a series of C ab , get the collaborative matrix C and optimize it.
[0158] Integrating a series of collaborative motion coefficients obtained for different hand key points to obtain a collaborative matrix, and dynamically optimizing the collaborative matrix through a regularization algorithm;
[0159] For example, regularization algorithms such as L1 or L2 regularization are introduced to prevent overfitting and improve the generalization ability of the model. The parameters of the regularization algorithm are selected by the method of cross-validation to find the optimal regularization strength. The collaborative matrix is updated dynamically. As new data is continuously collected, the coefficients in the collaborative matrix are adjusted in real time to adapt to the changes in the hand movement pattern. The incremental learning method is adopted. When new data arrives, only the affected coefficients are updated to reduce the computational amount.
[0160] According to the optimized collaborative matrix, a time-delay model is introduced to obtain the collaborative angle;
[0161] Considering that there may be a delay in the collaborative effect, a time-delay parameter is introduced to correct the change in the collaborative angle. Through experimental analysis, the time delay of the collaborative movement between different hand key points is obtained, and a time-delay model is established. For example, when the angle change of the a-th hand key point occurs at time t, the change in the collaborative angle of the b-th hand key point may only be reflected at time t + γ, where γ is the time-delay parameter.
[0162] The position coordinates of the hand key points are assisted and corrected by the collaborative angle.
[0163] In step S14, the data information of the hand key points is deeply encapsulated and transmitted to the Unity engine for parsing through the UDP communication protocol, and the hand movement is simulated and demonstrated in the Unity engine.
[0164] Specifically, the data information of the hand key points is deeply encapsulated in the ROS message type, and the data information includes position coordinates and color space information;
[0165] Among them, the data information is deeply encapsulated according to the strict message specifications of the Robot Operating System (ROS). The ROS message types use the Interface Definition Language (IDL) based on metadata description to define the data structure and format, so as to ensure the efficient and reliable transmission and sharing capabilities of data between different system components. For example, for the data information, through a custom ROS message type, the IDL syntax is used to accurately describe the field information such as the serial number of key points, three-dimensional coordinate values (represented in the Cartesian coordinate system), and coordinate confidence based on uncertainty estimation. Through the ROS message publishing and subscribing mechanism, using a distributed communication protocol based on the publish / subscribe mode, the encapsulated data is published to the message bus in the ROS network topology, and other nodes can obtain the data by subscribing to the corresponding message topics. This data encapsulation method based on a standardized data model and communication protocol greatly improves the scalability and compatibility of the system and facilitates seamless integration with other heterogeneous software systems.
[0166] It is transmitted through the UDP communication protocol, and after cyclic redundancy check, it is sent to the Unity engine for parsing;
[0167] Specifically, the User Datagram Protocol (UDP) is used as the underlying protocol for data transmission. As a connectionless transport layer protocol, UDP has a very low transmission delay characteristic due to its simple protocol stack design and efficient data transmission mechanism, and is particularly suitable for application scenarios with extremely high real-time requirements. In this system, the UDP protocol is responsible for transmitting the encapsulated ROS messages from the image acquisition and calculation node to the node where the Unity engine is located. The transmission process of UDP mainly involves three key steps: data packing, sending, and receiving. The sending end encapsulates the ROS message into a UDP data packet according to the data format specifications of the UDP protocol, adds the destination IP address (using the IPv6 protocol to adapt to future network development needs) and port number, and then sends it out through a network interface based on socket programming; the receiving end listens at the specified UDP port using an event-driven asynchronous I / O model. After receiving the data packet, it unpacks the data packet according to the checksum mechanism of the UDP protocol to obtain the original ROS message data. Although UDP itself does not guarantee the reliable transmission of data, in this system, by designing a cyclic redundancy check (CRC) based verification algorithm and a retransmission mechanism based on the sliding window protocol (when data errors or losses are detected, the sending end is requested to re-send according to the sequence number) at the receiving end (such as in the Unity engine), the accuracy and integrity of the data during transmission are ensured.
[0168] The Unity engine processes the parsed data information graphically, changes the virtual experiment scene and experimental equipment through matrix transformation, and simulates hand movement demonstrations. The formula is as follows:
[0169] ,
[0170] Among them, is the data information after matrix transformation, is the translation matrix, is the rotation matrix, is the scaling matrix.
[0171] Specifically, as the core operating platform of the virtual laboratory, the Unity engine uses its built-in network communication module, applies the network socket programming interface based on UDP, listens to the specified UDP port, and receives the interactive instruction data packets transmitted by UDP in real time. After receiving the data packets, the Unity engine converts the binary data into events and operation instructions recognizable inside the Unity engine according to the predefined parsing protocol. These instructions are used to update the virtual experiment scene and experimental equipment status in real time through the event-driven architecture and data-driven rendering update mechanism of the Unity engine. The Unity engine utilizes its rendering pipeline accelerated by the Graphics Processing Unit (GPU) and rich interactive component library. According to the received instructions, it changes the geometric properties such as the position, rotation, and scaling of the objects in the virtual experiment scene through matrix transformation in real time, realizing real-time and precise interaction between the virtual experiment scene, experimental equipment, and user operations, and providing users with a highly immersive experimental experience.
[0172] In addition, real-time interaction operations are carried out in the virtual experiment scene, such as using a physical simulation-based grasping algorithm to achieve virtual grasping of experimental instruments and adjusting experimental parameters through a parameterization modeling-based interaction method.
[0173] That is, according to the user's operations, real-time feedback is given through real-time physical engine simulation (such as rigid body dynamics simulation based on Newtonian mechanics and flexible body simulation based on finite element analysis) and data-driven feedback mechanisms, such as the response of experimental equipment according to physical laws and the dynamic changes of experimental data based on mathematical models. These feedback information is displayed to users in an intuitive and vivid manner through the rendering and interaction functions of the Unity engine, using skinning animation-based model-driven technology and data visualization-based chart rendering algorithms, such as the real movement of experimental instruments based on physical simulation, the lighting of indicator lights based on state machine control, and the real-time change of data charts based on data mapping.
[0174] In addition, the method further includes:
[0175] Based on the log recording of big data analysis and the behavior analysis algorithm of deep learning, comprehensively record and analyze the user's gesture demonstration, and determine whether the gesture demonstration is correct;
[0176] If so, complete the experimental simulation operation;
[0177] If not, obtain the experimental steps with incorrect gesture demonstrations for analysis.
[0178] It should be noted that comprehensively recording and deeply analyzing the user's operations is to provide multi-dimensional data support and decision-making basis for subsequent teaching evaluation and experimental optimization.
[0179] Compared with the prior art, adopting the experimental simulation method based on MediaPipe and virtual reality shown in this embodiment, by performing spatial component conversion on the gesture image and preprocessing the gesture image through the adaptive histogram equalization algorithm, it helps to more stably identify gesture features under different lighting conditions, improve the quality of the gesture image, lay a foundation for subsequent hand key point calculation, introduce a formula based on the relationship between adjacent key points to correct the hand key points, reduce the occurrence of finger occlusion, and improve the hand key points, thereby solving the technical problems in the prior art that traditional virtual experiments require special wearable devices or handles to implement, and are difficult to operate with high equipment costs.
[0180] Embodiment 2
[0181] Please refer to Figure 2 shown in the figure, which is an experimental simulation system based on MediaPipe and virtual reality provided by the second embodiment of the present invention. The system includes:
[0182] A virtual scene construction module 100, configured to render a 3D model through the Unity engine to construct a virtual experimental scene and experimental equipment;
[0183] A gesture image acquisition module 200, configured to acquire the user's gesture image and adjust the mapping of the gesture image through three-dimensional calibration;
[0184] An image preprocessing module 300, configured to perform spatial component conversion on the gesture image and preprocess the gesture image through the adaptive histogram equalization algorithm to obtain an initial image;
[0185] A coordinate correction module 400, configured to use the MediaPipe model to identify the hand key points in the initial image and correct the position coordinates of the hand key points according to the positional relationship of the hand key points;
[0186] The simulation experiment demonstration module 500 is used to deeply encapsulate the data information of the hand key points and transmit it to the Unity engine for parsing through the UDP communication protocol, and simulate the hand movement demonstration in the Unity engine.
[0187] Compared with the prior art, the experimental simulation system based on MediaPipe and virtual reality shown in this embodiment converts the spatial components of the gesture image through the image preprocessing module, and preprocesses the gesture image through the adaptive histogram equalization algorithm, which helps to more stably identify the gesture features under different lighting conditions, improve the quality of the gesture image, and lay a foundation for the subsequent calculation of hand key points. The coordinate correction module introduces a formula based on the relationship between adjacent key points to correct the hand key points, reduces the occurrence of finger occlusion, and improves the hand key points, thus solving the technical problems in the prior art that traditional virtual experiments require special wearable devices or handles to implement, and are difficult to operate and have high equipment costs.
[0188] The technical features of each of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0189] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable storage medium for an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can obtain instructions
[0190] In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0191] The above-described embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. An experimental simulation method based on MediaPipe and virtual reality, characterized in that The method includes: Rendering a 3D model through the Unity engine to construct a virtual experimental scenario and experimental equipment; Collecting the user's gesture images and adjusting the mapping of the gesture images through three-dimensional calibration; Performing spatial component conversion on the gesture images and preprocessing the gesture images through an adaptive histogram equalization algorithm to obtain initial images; Using the MediaPipe model to identify the hand key points in the initial images and correcting the position coordinates of the hand key points according to the positional relationship of the hand key points, including: Using the MediaPipe model to identify the hand key points in the initial images, calculating the distance and angle between adjacent hand key points, and the formula is as follows: , Among them, is the distance between adjacent hand key points, is the angle between adjacent hand key points, and are respectively the axis and axis coordinates of the hand key point , and are respectively the axis and axis coordinates of the hand key point . When it is detected that the distance and the angle exceed the preset range, it means that there is an occlusion situation of the gesture, and the position coordinates of the hand key points are corrected, including: Calculate hand key points To hand key points The key vectors between them are as follows: , Among them, is the key vector, Constructing a rotation matrix according to the angle, and the calculation formula is as follows: , wherein, is the angular difference, is the rotation matrix, is the minimum angle value, is the maximum angle value, Based on the rotation matrix and the key vector, calculating the rotation vector to be rotated, and the formula is as follows: , Among them, is a rotation vector, Based on the rotation vector, preliminarily correct the position coordinates of the hand key points, and the formula is as follows: , Among them, , are the position coordinates of the key points of the hand after preliminary correction, , , are respectively 's component, component; Deeply encapsulating the data information of the hand key points and transmitting it to the Unity engine for parsing through the UDP communication protocol, and simulating hand gesture demonstrations in the Unity engine.
2. The experimental simulation method based on MediaPipe and virtual reality according to claim 1, wherein The step of collecting the user's gesture images and adjusting the mapping of the gesture images through three-dimensional calibration specifically includes: Based on the spatial position of the camera and the change of the shooting angle, converting the three-dimensional points in the world coordinate system into three-dimensional points in the camera coordinate system, and the formula is as follows: , Among them, is the three-dimensional point data in the world coordinate system, are respectively the three-dimensional coordinates of is the three-dimensional point in the camera coordinate system, are respectively the three-dimensional coordinates of is the rotation matrix, is the translation vector; Then converting the camera coordinate system to the projection of the gesture image coordinate system, which is represented by the internal parameter matrix, and the formula is as follows: , Among them, is the internal parameter matrix, , are the focal lengths of the camera in the x-axis and y-axis directions respectively, is the tilt factor.
3. The experimental simulation method based on MediaPipe and virtual reality according to claim 2, characterized in that The step of performing spatial component conversion on the gesture images specifically includes: Extracting color components from the gesture images, and the color components include R component, G component, and B component, and their value ranges are between 0 and 1; Converting the color components into color space information, and the formula is as follows: , Among them, V is the brightness information, C is the color difference, S is the saturation, and H is the hue.
4. The experimental simulation method based on MediaPipe and virtual reality according to claim 3, characterized in that, The method further includes: Normalize the position coordinates ( of the hand key points after the preliminary correction to obtain the current coordinates; , ) In the simulated annealing process, the initial temperature is set, and iterative cooling is carried out according to the preset cooling strategy. In each round of iteration, random offsets are added in the and directions to perturb the current coordinates, and new coordinates are obtained. Calculating the fitness corresponding to the new coordinates, comparing it with the fitness of the current coordinates, calculating the difference, and obtaining the current coordinates in each round of iteration according to the difference; When the preset number of iterations is reached, obtain the current coordinates and the corresponding fitness values in each iteration, and calculate the hand key points by inverse operation using the current coordinates with the minimum fitness value. The final corrected position coordinates.
5. The experimental simulation method based on MediaPipe and virtual reality according to claim 4, characterized in that The calculation formula of the fitness function is: , Among them, F is the fitness of hand key points, , are weight coefficients, is the distance from the new coordinate of the hand key point to the position coordinate of the hand key point , is the angle from the new coordinate of the key point to the position coordinate of the hand key point , is the distance from the position coordinate of the hand key point in the correct gesture demonstration to the position coordinate of the hand key point , is the angle from the position coordinate of the hand key point in the correct gesture demonstration to the position coordinate of the hand key point .
6. The experimental simulation method based on MediaPipe and virtual reality according to claim 1, characterized in that, The step of deeply encapsulating the data information of the hand key points and transmitting it to the Unity engine for parsing through the UDP communication protocol, and simulating hand gesture demonstrations in the Unity engine specifically includes: Deeply encapsulating the data information of the hand key points through the ROS message type, and the data information includes position coordinates and color space information; Transmitting through the UDP communication protocol, and after cyclic redundancy check, transporting it to the Unity engine for parsing; The Unity engine performs graphic processing on the parsed data information, changes the virtual experimental scenario and experimental equipment through matrix transformation, and simulates hand gesture demonstrations, and the formula is as follows: , Among them, is the data information after matrix transformation, is the translation matrix, is the rotation matrix, is the scaling matrix.
7. The experimental simulation method based on MediaPipe and virtual reality according to claim 5, wherein The method further includes: Taking each hand key point as an independent variable and other hand key points as dependent variables, introducing non-linear features, and establishing a collaborative regression model; Solve the regression coefficients of the collaborative regression model by the least squares method, and iteratively optimize the regression coefficients to obtain the collaborative motion coefficients between different hand key points; Integrate a series of collaborative motion coefficients of different hand key points obtained, obtain a collaborative matrix, and dynamically optimize the collaborative matrix through a regularization algorithm; According to the optimized collaborative matrix, introduce a time delay model to obtain a collaborative angle; Assist in correcting the position coordinates of hand key points through the collaborative angle.
8. An experimental simulation system based on MediaPipe and virtual reality, characterized in that, The system is used to implement the experimental simulation method based on MediaPipe and virtual reality described in any one of claims 1 to 7. The system includes: A virtual scene construction module for rendering a 3D model through the Unity engine to construct a virtual experimental scene and experimental equipment; A gesture image acquisition module for acquiring a user's gesture image and adjusting the mapping of the gesture image through three-dimensional calibration; An image preprocessing module for performing spatial component conversion on the gesture image and preprocessing the gesture image through an adaptive histogram equalization algorithm to obtain an initial image; A coordinate correction module for using the MediaPipe model to identify hand key points in the initial image and correcting the position coordinates of the hand key points according to the positional relationship of the hand key points, including: Use the MediaPipe model to identify hand key points in the initial image, and calculate the distance and angle between adjacent hand key points. The formulas are as follows: , Among them, is the distance between adjacent hand key points, is the angle between adjacent hand key points, , are respectively the -axis and -axis coordinates of the hand key point , , are respectively the -axis and -axis coordinates of the hand key point . When it is detected that the distance and the angle exceed a preset range, it means that there is an occlusion situation in the gesture, and the position coordinates of the hand key points are corrected, including: Calculate hand key points To hand key points The key vectors between them are as follows: , Among them, is the key vector, Construct a rotation matrix according to the angle. The calculation formula is as follows: , Among them, is the angular difference, is the rotation matrix, is the minimum angle value, is the maximum angle value, Based on the rotation matrix and the key vector, calculate the rotation vector to be rotated. The formula is as follows: , Among them, is a rotation vector, Based on the rotation vector, preliminarily correct the position coordinates of the hand key points, and the formula is as follows: , Among them, , are the position coordinates of the hand key points after preliminary correction, , are respectively components of component, component; A simulation experiment demonstration module for deeply encapsulating the data information of the hand key points and transmitting it to the Unity engine for parsing through the UDP communication protocol, and simulating hand motion demonstrations in the Unity engine.
Citation Information
Patent Citations
Multi-camera ship height measurement method based on dynamic long baseline
CN115546280A
Experimental AR simulation system based on 3D hand posture estimation method
CN117648032A