MediaPipe and virtual reality-based experimental simulation method and system

Through MediaPipe and virtual reality technology, combined with the Unity engine and MediaPipe model, virtual experimental simulation without special equipment is realized, solving the problem of dependence and high cost of traditional virtual experimental equipment, and improving the stability and accuracy of gesture recognition.

CN120122831AActive Publication Date: 2025-06-10NANCHANG CAMPUS OF JIANGXI UNIV OF SCI & TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510599815.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-06-10
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

Traditional virtual experiments require special wearable devices or handles, which are difficult to operate and have high equipment costs.

Method used

Using an experimental simulation method based on MediaPipe and virtual reality, the 3D model is rendered through the Unity engine, virtual experimental scenes and experimental equipment are constructed, and the user's gesture images are collected and preprocessed. The MediaPipe model is used to identify the key points of the hand, and transmitted to the Unity engine through the UDP communication protocol for simulation.

Benefits of technology

Under different lighting conditions, the gesture features are more stable, the gesture image quality is improved, the fingers are blocked, and the hand key points recognition accuracy is improved, solving the equipment dependence and high cost problems of traditional virtual experiments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122831A_ABST
    Figure CN120122831A_ABST
Patent Text Reader

Abstract

The invention provides an experiment simulation method and system based on MediaPipe and virtual reality, and relates to the technical field of image recognition, and the method comprises the steps: carrying out the rendering of a 3D model through a Unity engine, and constructing a virtual experiment scene and experiment equipment; a gesture image of a user is collected, and mapping of the gesture image is adjusted through three-dimensional calibration; performing spatial component conversion on the gesture image, and preprocessing the gesture image to obtain an initial image; using a MediaPipe model to identify the hand key points, and correcting the position coordinates according to the position relation of the hand key points; the data information of the key points of the hand is deeply packaged and transmitted to a Unity engine through a UDP communication protocol to analyze and simulate hand action demonstration, and the technical problems that in the prior art, a traditional virtual experiment needs special wearable equipment or a handle to be achieved, operation is difficult, and the equipment cost is high can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly to an experimental simulation method and system based on MediaPipe and virtual reality. Background Art

[0002] Under the traditional experimental teaching mode, many thorny problems will emerge when students conduct actual operation experiments in person. For example, from the perspective of economic cost, during the operation process, due to unfamiliarity with the experimental procedures and inaccurate techniques, students are extremely likely to cause damage and additional losses to experimental equipment. From the safety aspect, students are at risk of potential safety accidents due to unskilled operation of equipment. Moreover, affected by the limitation of the site space size in traditional experimental teaching, it is difficult to accommodate a large number of students to conduct experiments simultaneously, and the number of experimental equipment also cannot meet the needs of large-scale teaching. Often, multiple people share a set of equipment, which compresses the actual operation time of students and greatly reduces the teaching efficiency. Therefore, virtual experiment simulation has begun to enter the teaching field.

[0003] Current virtual experiments usually require special wearable devices or hardware such as handles to be realized. Accurately simulating hand movements or touch operations during the experiment through sensing technology is difficult to operate, and the cost of these hardware devices is high. They also require maintenance and installation, which is cumbersome and complex, and will increase the workload of staff. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide an experimental simulation method and system based on MediaPipe and virtual reality, aiming to solve the technical problems in the prior art that traditional virtual experiments require special wearable devices or handles to be realized, and are difficult to operate and have high equipment costs.

[0005] On the one hand, the present invention provides an experimental simulation method based on MediaPipe and virtual reality, and the method includes: Rendering a 3D model through the Unity engine to construct a virtual experimental scene and experimental equipment; Collecting the gesture images of the user, and adjusting the mapping of the gesture images through three-dimensional calibration; Performing spatial component conversion on the gesture images, and preprocessing the gesture images through an adaptive histogram equalization algorithm to obtain initial images; Using a MediaPipe model to recognize the hand key points in the initial images, and correcting the position coordinates of the hand key points according to the positional relationship of the hand key points; Deeply encapsulate the data information of the hand key points and transmit it to the Unity engine for parsing through the UDP communication protocol, and simulate the hand action demonstration in the Unity engine.

[0006] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the experimental simulation method based on MediaPipe and virtual reality provided by the present invention, by performing spatial component conversion on the gesture image and preprocessing the gesture image through the adaptive histogram equalization algorithm, it helps to more stably identify gesture features under different lighting conditions, improve the quality of the gesture image, lay a foundation for subsequent hand key point calculation, introduce a formula based on the relationship between adjacent key points to correct the hand key points, reduce the occurrence of finger mutual occlusion, and improve the hand key points, thereby solving the technical problems in the prior art that traditional virtual experiments require special wearable devices or handles to implement, and are difficult to operate and have high equipment costs.

[0007] According to one aspect of the above technical solution, the steps of collecting the user's gesture image and adjusting the mapping of the gesture image through three-dimensional calibration specifically include: Based on the spatial position and shooting angle changes of the camera, convert the three-dimensional points in the world coordinate system into three-dimensional points in the camera coordinate system. The formula is as follows: , Among them, is the three-dimensional point data in the world coordinate system, are respectively 's three-dimensional coordinates, is the three-dimensional point in the camera coordinate system, are respectively 's three-dimensional coordinates, is the rotation matrix, is the translation vector; Then project the camera coordinate system to the gesture image coordinate system, which is represented by the internal parameter matrix. The formula is as follows: , Among them, is the internal parameter matrix, , are respectively the focal lengths of the camera in the x-axis and y-axis directions, is the tilt factor.

[0008] According to one aspect of the above technical solution, the steps of performing spatial component conversion on the gesture image specifically include: Extract the color components of the gesture image. The color components include the R component, G component, and B component, and their value ranges are between 0 and 1; Convert the color components into color space information, and the formula is as follows: , where V is the brightness information, C is the color difference, S is the saturation, and H is the hue.

[0009] According to one aspect of the above technical solution, the steps of using the MediaPipe model to identify the hand key points in the initial image and correcting the position coordinates of the hand key points according to the positional relationship of the hand key points specifically include: Use the MediaPipe model to identify the hand key points in the initial image, and calculate the distance and angle between adjacent hand key points. The formula is as follows: , where is the distance between adjacent hand key points, is the angle between adjacent hand key points, , are the axis and axis coordinates of the hand key point respectively, , are the axis and axis coordinates of the hand key point respectively; When it is detected that the distance and the angle exceed the preset range, there is an occlusion situation for the gesture, and the position coordinates of the hand key points are corrected.

[0010] According to one aspect of the above technical solution, the steps of calculating the key vector between the hand key point and the hand key point and correcting the position coordinates of the hand key point when it is detected that the distance and the angle exceed the preset range specifically include: Calculate the key vector between the hand key point , where is the key vector; Construct a rotation matrix according to the angle. The calculation formula is as follows: , where is the angle difference, is the rotation matrix, is the minimum angle value, is the maximum angle value; Based on the rotation matrix and the key vector, calculate the rotation vector of the required rotation, and the formula is as follows: , where, is the rotation vector; Based on the rotation vector, preliminarily correct the position coordinates of the hand key point , and the formula is as follows: , where, , are the position coordinates of the hand key point after preliminary correction, , are respectively of component, component.

[0011] According to one aspect of the above technical solution, the method further includes: Normalize the position coordinates ( , , ) of the hand key point after preliminary correction to obtain the current coordinates; Simulate the annealing process, set the initial temperature, and perform iterative cooling according to the preset cooling strategy. In each round of iteration, add random offsets in the directions of and to the current coordinates for perturbation to obtain new coordinates; Calculate the fitness corresponding to the new coordinates, compare it with the fitness of the current coordinates, calculate the difference, and obtain the current coordinates in each round of iteration according to the difference; When the preset number of iterations is reached, obtain the current coordinates and the corresponding fitness in each round of iteration, and calculate the final corrected position coordinates of the hand key point through inverse operation for the current coordinates with the minimum fitness.

[0012] According to one aspect of the above technical solution, the calculation formula of the fitness function is: , where, F is the fitness of the hand key point, , are the weight coefficients, is the distance from the new coordinates of the hand key point to the position coordinates of the hand key point , is the angle from the new coordinates of the key point to the position coordinates of the hand key point . The position coordinates of the hand key points for correct gesture demonstration to the position coordinates of the hand key points distance, The position coordinates of the hand key points for correct gesture demonstration to the position coordinates of the hand key points angle.

[0013] According to one aspect of the above technical solution, the data information of the hand key points is deeply encapsulated and transmitted to the Unity engine for parsing through the UDP communication protocol. The steps of simulating hand gesture demonstrations in the Unity engine specifically include: Deeply encapsulate the data information of the hand key points through the ROS message type. The data information includes position coordinates and color space information; Transmit through the UDP communication protocol, and after cyclic redundancy check, send it to the Unity engine for parsing; The Unity engine performs graphic processing on the parsed data information, changes the virtual experimental scene and experimental equipment through matrix transformation, and simulates hand gesture demonstrations. The formula is as follows: , where, is the data information after matrix transformation, is the translation matrix, is the rotation matrix, is the scaling matrix.

[0014] According to one aspect of the above technical solution, the method further includes: Taking each hand key point as an independent variable and other hand key points as dependent variables, introducing non-linear features, and establishing a collaborative regression model; Solving the regression coefficients of the collaborative regression model by the least squares method, and iteratively optimizing the regression coefficients to obtain the collaborative motion coefficients between different hand key points; Integrate a series of collaborative motion coefficients of different hand key points obtained to obtain a collaborative matrix, and dynamically optimize the collaborative matrix through a regularization algorithm; According to the optimized collaborative matrix, introduce a time delay model to obtain a collaborative angle; Assist in correcting the position coordinates of the hand key points through the collaborative angle.

[0015] Another aspect of the present invention is to provide an experimental simulation system based on MediaPipe and virtual reality, which is used to implement the above-mentioned experimental simulation method based on MediaPipe and virtual reality. The system includes: A virtual scene construction module, which is used to render 3D models through the Unity engine to construct a virtual experimental scene and experimental equipment; A gesture image acquisition module, which is used to acquire the gesture images of the user and adjust the mapping of the gesture images through three-dimensional calibration; An image preprocessing module, which is used to perform spatial component conversion on the gesture images and preprocess the gesture images through an adaptive histogram equalization algorithm to obtain initial images; A coordinate correction module, which is used to identify the hand key points in the initial image using the MediaPipe model and correct the position coordinates of the hand key points according to the positional relationship of the hand key points; A simulation experiment demonstration module, which is used to deeply encapsulate the data information of the hand key points and transmit it to the Unity engine for parsing through the UDP communication protocol, and simulate hand movement demonstrations in the Unity engine. Description of the Drawings

[0016] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where: Figure 1 It is a schematic flowchart of the experimental simulation method based on MediaPipe and virtual reality in Embodiment 1 of the present invention; Figure 2 It is a structural block diagram of the experimental simulation system based on MediaPipe and virtual reality in Embodiment 2 of the present invention; Explanation of the Reference Signs in the Drawings: Virtual scene construction module 100, gesture image acquisition module 200, image preprocessing module 300, coordinate correction module 400, simulation experiment demonstration module 500. Detailed Embodiments

[0017] To make the objectives, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0018] Embodiment 1 Please refer to Figure 1, an experimental simulation method based on MediaPipe and virtual reality provided by the first embodiment of the present invention, the method comprising steps S10 - S14: Step S10, rendering a 3D model through the Unity engine to construct a virtual experimental scene and experimental equipment; Among them, the virtual experimental scene and experimental equipment are presented in a highly refined three-dimensional model form within the user's field of vision. The projection device selects a professional product with ultra-high resolution (such as reaching 4K and above) and high brightness to ensure that the model interface can maintain excellent clarity and color restoration in different lighting environments, accurately replicating the visual effects of the real scene.

[0019] Furthermore, the displayed model interface is constructed relying on the powerful graphics rendering framework of the Unity engine. With the help of the Scriptable Render Pipeline (SRP) of the Unity engine, real-time and efficient rendering operations are performed on elements such as the complex geometric structure of the 3D model, high-precision material textures, and realistic lighting effects (including physically based rendering, PBR technology-based light and shadow simulation), generating an extremely realistic and highly immersive virtual scene. Therefore, the model interface integrates an advanced interactive logic architecture, and users can perform natural and smooth interactive operations with various devices in the scene through subsequent gesture recognition interaction technology.

[0020] Before the step of rendering a 3D model through the Unity engine to construct a virtual experimental scene and experimental equipment, it further includes: Based on deep learning speech recognition of the Transformer model, starting the required virtual experimental scene and experimental equipment through voice commands.

[0021] It should be noted that when the user starts the required virtual experimental scene and experimental equipment through voice commands, that is, using speech recognition technology to convert the user's voice into machine-recognizable commands. The speech recognition technology uses a speech recognition model based on deep learning, such as a model based on the Transformer architecture. Its core principle is to capture and analyze the features in the speech signal through the multi-head attention mechanism. This method can more accurately recognize speech commands under various accents and language habits compared to traditional speech recognition methods, greatly improving the convenience and accuracy of system startup. During the startup process, the system will automatically load relevant algorithm models and configuration files to ensure the normal operation of subsequent gesture recognition and interaction functions.

[0022] Step S11, collecting the user's gesture images and adjusting the mapping of the gesture images through three-dimensional calibration; For example, a professional motion capture marker set is worn by the subject to collect data of gesture images. A series of rich and diverse representative hand movements are designed.

[0023] By way of example and not limitation, in addition to common movements such as making a fist, stretching fingers, and bending specific fingers, it also includes grasping objects of different shapes (such as circles, squares, triangles) and sizes (such as circular objects with a diameter of 2 cm - 10 cm), and performing fine gesture operations (such as pinching small objects, rotating knobs, etc.). Each movement is repeated no less than 20 times to obtain sufficient and stable data samples.

[0024] Specifically, during the data collection process, a synchronization device is used to ensure the synchronization of the motion capture system and the video device that records the subject's movements. The shooting angle of the video device should cover all directions of the hand to facilitate subsequent multi-dimensional verification and analysis of the data.

[0025] The collected original joint position data is filtered to remove noise and outliers. A method combining Kalman filtering and median filtering is adopted. First, Kalman filtering is used to perform preliminary smoothing on the data, and then median filtering is used to remove possible remaining outliers. The filtering parameters are dynamically adjusted according to different movement types and data characteristics to achieve the best filtering effect.

[0026] Specifically, based on the spatial position of the camera and the change in the shooting angle, the three-dimensional points in the world coordinate system are converted into three-dimensional points in the camera coordinate system. The formula is as follows: , where, is the three-dimensional point data in the world coordinate system, are respectively the three-dimensional coordinates of, is the three-dimensional point in the camera coordinate system, are respectively the three-dimensional coordinates of, is the rotation matrix, is the translation vector; Then, the projection of the camera coordinate system onto the gesture image coordinate system is represented by the intrinsic matrix. The formula is as follows: , where, is the intrinsic matrix, , are respectively the focal lengths of the camera in the x-axis and y-axis directions, is the tilt factor.

[0027] It should be noted that the spatial position and shooting angle of the camera are precisely adjusted through a 3D calibration algorithm based on machine vision, and in combination with the triangulation principle of multi-view vision, the best shooting perspective is obtained, effectively avoiding problems such as image acquisition quality degradation caused by occlusion, environmental reflection and other factors.

[0028] By way of example and not limitation, for example, an industrial-grade high-definition camera with high frame rate (frame rate set at 120fps or above) and low latency characteristics is used to collect the user's gesture images in real time. The frame rate of the camera is optimized through strict timing control and synchronization algorithms to ensure that it can accurately capture the user's fast and subtle gesture movements.

[0029] Further, before the step of collecting the user's gesture images and adjusting the mapping of the gesture images through 3D calibration, the following steps are also included: Based on speech synthesis technology and text synthesis technology, voice or text reminder information is generated before each experimental step.

[0030] Among them, according to the status of the current experimental step and the semantic understanding analysis result of the virtual experimental scene, voice reminders are generated through cutting-edge speech synthesis technology, or the corresponding operation gesture guidance information is displayed in text form using a UI system based on vector graphics rendering on the display interface. Speech synthesis uses end-to-end text-to-speech (TTS) technology based on deep learning, such as the cutting-edge model FastSpeech2 based on the Transformer architecture. This model can generate natural and fluent speech close to real human voices through a multi-modal fusion attention mechanism and an improved variant of the long short-term memory network (LSTM). The text reminder is provided through the UI system of the Unity engine, using a graphic drawing algorithm based on Bezier curves and an adaptive layout algorithm to use simple and clear icons and text with visual guidance to illustrate the gesture operations that the user can perform currently, helping the user quickly understand and master the way to interact with the virtual experimental scene and experimental equipment. These operation guidance information is dynamically updated through real-time semantic analysis and decision tree algorithms according to the state transition of different experimental steps and the dynamic changes of the virtual experimental scene, ensuring that the user can always clearly know the operation process of the next experimental step.

[0031] Step S12: Perform spatial component conversion on the gesture image, and preprocess the gesture image through an adaptive histogram equalization algorithm to obtain an initial image; Specifically, color components are extracted from the gesture image, and the color components include R component, G component, and B component, and their value ranges are between 0 and 1; Convert the color components into color space information, and the formula is as follows: , Among them, V is the brightness information, C is the color difference, S is the saturation, and H is the hue.

[0032] It should be noted that after obtaining the gesture image collected by the camera, the conversion of the spatial components helps to more stably identify the gesture features under different lighting conditions.

[0033] Secondly, preprocessing operations are performed to improve the quality of the gesture image and lay a foundation for subsequent hand key point calculation, such as histogram equalization and Gaussian filtering.

[0034] Furthermore, the gesture image is preprocessed by an adaptive histogram equalization algorithm to obtain an initial image.

[0035] It should be noted that the adaptive histogram equalization algorithm divides the gesture image into multiple small blocks, performs histogram equalization on each small block separately, and at the same time limits the enhancement amplitude of the contrast to avoid noise amplification caused by over-enhancement. It can be implemented using the cv2.createCLAHE() function in OpenCV in Python.

[0036] Step S13, use the MediaPipe model to identify the hand key points in the initial image, and correct the position coordinates of the hand key points according to the positional relationship of the hand key points; Specifically, use the MediaPipe model to identify the hand key points in the initial image, calculate the distance and angle between adjacent hand key points, and the formulas are as follows: , Among them, is the distance between adjacent hand key points, is the angle between adjacent hand key points, , are the axis and axis coordinates of the hand key point respectively, , are the axis and axis coordinates of the hand key point respectively; When it is detected that the distance and the angle exceed the preset range, it means that there is an occlusion situation of the gesture, and the position coordinates of the hand key points are corrected.

[0037] Calculate the key vector between the hand key point and the hand key point , and the formula is as follows: , Among them, is the key vector; According to the said angle, construct a rotation matrix, and the calculation formula is as follows: , Among them, is the angle difference, is the rotation matrix, is the minimum angle value, is the maximum angle value; Based on the said rotation matrix and the said key vector, calculate the rotation vector required for rotation, and the formula is as follows: , Among them, is the rotation vector; Based on the said rotation vector, preliminarily correct the position coordinates of the said hand key points , and the formula is as follows: , Among them, , are the position coordinates of the hand key point after preliminary correction, , are respectively 's component, component.

[0038] It should be noted that although the MediaPipe model can detect hand key points well, in the case of finger occlusion, the detection results may have errors. To solve this problem, a formula based on the relationship between adjacent key points is introduced for preliminary correction. The movement and posture of fingers usually follow certain physiological structures and movement laws. The occlusion situation can be judged by analyzing the relationships such as the distance and angle between adjacent key points, and preliminary correction can be carried out. For example, the activity angle of the finger joints of the human body has certain physiological limitations. The minimum angle of the metacarpophalangeal joint of the index finger and the middle finger is -π / 6, and the maximum angle is π / 3. By preliminarily correcting the coordinates of the occluded fingers, the accuracy of the coordinates is improved.

[0039] The said method further includes: Normalize the position coordinates ( , ) of the hand key points after preliminary correction to obtain the current coordinates; Specifically, the numerical value of the coordinate normalization process is in the interval [0, 1].

[0040] The normalized position coordinates are respectively perturbed by adding random offsets in the and directions to obtain the current coordinates; For example but not limited to, and are respectively the random offsets in the and directions, is a uniform normal distribution, , is a small positive number set according to the actual situation to control the perturbation amplitude.

[0041] , Simulate the annealing process, set the initial temperature, and perform iterative cooling according to the preset cooling strategy. In each round of iteration, the current coordinates are respectively perturbed by adding random offsets in the and directions to obtain new coordinates; Among them, the preset cooling strategy can be expressed as , is a constant, slightly less than 1, is the temperature of the th round of iteration, is the temperature of the th round of iteration.

[0042] Calculate the fitness corresponding to the new coordinates and compare it with the fitness of the current coordinates, calculate the difference, and obtain the current coordinates in each round of iteration according to the difference; Specifically, the calculation formula of the fitness function is: , where F is the fitness of the hand key points, , are weight coefficients, is the distance from the new coordinate of the hand key point to the position coordinate of the hand key point , is the angle from the new coordinate of the key point to the position coordinate of the hand key point , is the distance from the position coordinate of the hand key point of the correct gesture demonstration to the position coordinate of the hand key point , is the distance from the position coordinate of the hand key point of the correct gesture demonstration to the position coordinate of the hand key point .

[0043] In addition, the method further includes: Determine whether the difference is less than a preset difference; If so, determine the new coordinate as the current coordinate; If not, calculate the coordinate probability of the new coordinate. When the coordinate probability exceeds the preset coordinate probability, determine the new coordinate as the current coordinate. When the coordinate probability does not exceed the preset coordinate probability, delete the new coordinate.

[0044] Among them, the coordinate probability can be expressed as: , is the coordinate probability, is the difference in the th round of iteration, is the th round of iteration temperature.

[0045] When the preset number of iterations is reached, obtain the current coordinate and the corresponding fitness in each round of iteration. Through inverse operation on the current coordinate with the minimum fitness, calculate the final corrected position coordinate of the hand key point

[0046] The method further includes: Record the change trend of the fitness in the past several rounds of iteration. If the fitness decreases slowly or no longer decreases continuously for multiple rounds, terminate the iteration in advance or adjust the cooling strategy; When the iteration is terminated in advance, obtain the current coordinate and the corresponding fitness in each round of iteration. Through inverse operation on the current coordinate with the minimum fitness, calculate the final corrected position coordinate of the hand key point

[0047] In order to further improve the accuracy of the preliminary correction, a simulated annealing process, a perturbation strategy, and a hand key point fitness detection are added to obtain the final corrected position coordinate, so as to obtain the most reasonable and highest-precision position coordinate.

[0048] In an actual scenario, the torsion of the metacarpophalangeal joint will produce a non-linear coupling effect. Therefore, only linear correction is not enough. To achieve non-linear correction, a collaborative regression model is introduced. Specifically: Taking each hand key point as an independent variable and other hand key points as dependent variables, introduce non-linear features and establish a collaborative regression model; For example, considering the non-linear features of hand joint movement, introduce polynomial terms and interaction terms to improve the fitting accuracy of the model.

[0049] ​​Among them, each piece of hand key point data collected is divided into a training set, a validation set, and a test set according to a certain ratio (such as 70%, 15%, 15%). During the division process, the stratified sampling method is adopted to ensure that the data distribution of different action types and subjects in each dataset is uniform. The data is normalized, and each piece of hand key point data is mapped to the range of [0, 1] or [-1, 1] to accelerate the training speed of the neural network and improve the stability of the model. The maximum-minimum normalization method is used to determine the normalization parameters according to the maximum and minimum values of the training set data, and then the validation set and test set data are normalized according to the same parameters. To increase the diversity of training data, data augmentation techniques are adopted.

[0050] In addition to performing operations such as random rotation, translation, and scaling on the hand key point data, noise interference can also be introduced, and different lighting conditions can be simulated. Through these operations, new training samples are generated to improve the robustness of the model. During the data augmentation process, the parameters of the augmentation operations are randomly controlled. For example, the rotation angle is randomly selected between -30° and 30°, and the translation distance is randomly selected between -5mm and 5mm to ensure that the generated training samples have sufficient diversity.

[0051] In addition, the number of neurons in the input layer is equal to the number of hand key points to ensure that all the angle information of the hand key points can be received. An input normalization layer is added to the input layer to further normalize the input data to improve the robustness of the model. The input normalization layer uses the batch normalization method to normalize the input data during each training, making the mean of the input data 0 and the variance 1. The middle layer contains multiple hidden layers, and the number of neurons in each hidden layer can be adjusted according to the experimental results.

[0052] Generally speaking, the number of neurons can be set in a decreasing manner. For example, the first hidden layer is set to 128 neurons, the second hidden layer is set to 64 neurons, etc. When setting the number of neurons, it is optimized by the grid search method to find the optimal combination of the number of neurons. A batch normalization layer and a ReLU activation function are added between each hidden layer to accelerate the convergence of the network and introduce non-linear features. The batch normalization layer can reduce the internal covariate shift and improve the training efficiency of the network. The ReLU activation function can introduce non-linearity, enabling the network to learn more complex features.

[0053] The number of neurons in the output layer is also equal to the number of hand key points. The output considers the hand key point data after the non - linear coupling effect. An output scaling layer is added to the output layer to restore the output data to the original hand key point data range. The output scaling layer performs the reverse operation according to the normalized parameters to restore the normalized output data to the actual hand key point data. The mean squared error (MSE) loss function is selected to measure the difference between the hand key point data output by the network and the actual hand key point data. The MSE loss function can intuitively reflect the error size of the network output. To improve the robustness of the model to outliers, the Huber loss function is adopted. The Huber loss function uses the MSE loss function when the error is small and the absolute value loss function when the error is large, which can effectively reduce the impact of outliers on model training.

[0054] Solve the regression coefficients of the collaborative regression model by the least - squares method, and iteratively optimize the regression coefficients to obtain the collaborative motion coefficients between different hand key points; Among them, in order to improve the accuracy of the solution, an iterative optimization method is adopted to iteratively update the regression coefficients multiple times.

[0055] For example, for the a - th hand key point and the b - th hand key point, the coefficient C is obtained through regression analysis. ab . Aggregate a series of C ab , to obtain the collaborative matrix C and optimize it.

[0056] Integrate a series of collaborative motion coefficients of different hand key points obtained, to obtain the collaborative matrix, and dynamically optimize the collaborative matrix through a regularization algorithm; For example, introduce a regularization algorithm, such as L1 or L2 regularization, to prevent overfitting and improve the generalization ability of the model. The parameters of the regularization algorithm are selected by the method of cross - validation to find the optimal regularization strength. Dynamically update the collaborative matrix. As new data is continuously collected, the coefficients in the collaborative matrix are adjusted in real - time to adapt to the changes in hand movement patterns. Adopt the method of incremental learning. When new data arrives, only update the affected coefficients to reduce the computational amount.

[0057] According to the optimized collaborative matrix, introduce a time - delay model to obtain the collaborative angle; Considering that the collaborative effect may have a delay, introduce a time - delay parameter to correct the change of the collaborative angle. Through experimental analysis, obtain the collaborative motion time - delay between different hand key points and establish a time - delay model. For example, when the angle change of the a - th hand key point occurs at time t, the change of the collaborative angle of the b - th hand key point may be reflected at time t + γ, where γ is the time - delay parameter.

[0058] The position coordinates of the hand key points are corrected by the collaborative angle assistance.

[0059] In step S14, the data information of the hand key points is deeply encapsulated and transmitted to the Unity engine for parsing through the UDP communication protocol, and the hand movement is simulated in the Unity engine.

[0060] Specifically, the data information of the hand key points is deeply encapsulated through the ROS message type, and the data information includes position coordinates and color space information. Among them, the data information is deeply encapsulated according to the strict message specifications of the Robot Operating System (ROS). The ROS message type uses the Interface Definition Language (IDL) based on metadata description to define the data structure and format, so as to ensure the efficient and reliable transmission and sharing capabilities of data between different system components. For example, for the data information, through a custom ROS message type, the IDL syntax is used to accurately describe the field information such as the serial number of the key point, the three-dimensional coordinate value (represented by the Cartesian coordinate system), and the coordinate confidence based on uncertainty estimation. Through the ROS message publishing and subscribing mechanism, using the distributed communication protocol based on the publish / subscribe mode, the encapsulated data is published to the message bus in the ROS network topology, and other nodes can obtain the data by subscribing to the corresponding message topic. This data encapsulation method based on the standardized data model and communication protocol greatly improves the scalability and compatibility of the system and is convenient for seamless integration with other heterogeneous software systems.

[0061] It is transmitted through the UDP communication protocol, and after cyclic redundancy check, it is sent to the Unity engine for parsing. Specifically, the User Datagram Protocol (UDP) is adopted as the underlying protocol for data transmission. As a connectionless transport layer protocol, UDP features a concise protocol stack design and an efficient data transmission mechanism, with extremely low transmission latency, making it particularly suitable for application scenarios with extremely high real-time requirements. In this system, the UDP protocol is responsible for transmitting the encapsulated ROS messages from the image acquisition and solution nodes to the node where the Unity engine is located. The transmission process of UDP mainly involves three key steps: data packaging, sending, and receiving. The sending end encapsulates the ROS messages into UDP data packets according to the data format specifications of the UDP protocol, adds the destination IP address (using the IPv6 protocol to meet the future network development needs) and port number, and then sends them out through the network interface based on socket programming; the receiving end listens at the specified UDP port using an event-driven asynchronous I / O model. After receiving the data packet, it unpacks the packet according to the checksum mechanism of the UDP protocol to obtain the original ROS message data. Although UDP itself does not guarantee reliable data transmission, in this system, by designing a cyclic redundancy check (CRC)-based verification algorithm and a retransmission mechanism based on the sliding window protocol at the receiving end (such as in the Unity engine), the accuracy and integrity of the data during transmission are ensured.

[0062] The Unity engine processes the parsed data information graphically, changes the virtual experiment scene and experimental equipment through matrix transformation, and simulates hand movement demonstrations. The formula is as follows: , Among them, is the data information after matrix transformation, is the translation matrix, is the rotation matrix, is the scaling matrix.

[0063] Specifically, the Unity engine serves as the core operating platform of the virtual laboratory. Through its built-in network communication module, it uses the UDP-based network socket programming interface to listen to a specified UDP port and receive interactive instruction data packets transmitted via UDP in real time. After receiving the data packets, the Unity engine converts the binary data into events and operation instructions recognizable within the Unity engine according to a predefined parsing protocol. These instructions are used to update the virtual experiment scene and the status of experimental equipment in real time through the event-driven architecture of the Unity engine and the data-driven rendering update mechanism. The Unity engine utilizes its rendering pipeline accelerated by the Graphics Processing Unit (GPU) and a rich library of interactive components. According to the received instructions, it can change the geometric properties such as the position, rotation, and scaling of objects in the virtual experiment scene in real time through matrix transformation, realizing real-time and precise interaction among the virtual experiment scene, experimental equipment, and user operations, and providing users with a highly immersive experimental experience.

[0064] In addition, real-time interaction operations are performed in the virtual experiment scene, such as using a physical simulation-based grasping algorithm to achieve virtual grasping of experimental instruments and adjusting experimental parameters through a parameterized modeling-based interaction method.

[0065] That is, according to the user's operations, real-time feedback is given through real-time physical engine simulation (such as rigid body dynamics simulation based on Newtonian mechanics and flexible body simulation based on finite element analysis) and a data-driven feedback mechanism, such as the response of experimental equipment according to physical laws and the dynamic changes of experimental data based on mathematical models. These feedback messages are presented to users in an intuitive and vivid manner through the rendering and interaction functions of the Unity engine, using skinning animation-based model-driven technology and data visualization-based chart rendering algorithms, such as the real movement of experimental instruments based on physical simulation, the lighting of indicator lights based on state machine control, and the changes in real-time data charts based on data mapping.

[0066] In addition, the method further includes: Based on big data analysis-based logging and deep learning-based behavior analysis algorithms, comprehensively record and analyze the user's gesture demonstrations and determine whether the gesture demonstrations are correct; If so, complete the experimental simulation operation; If not, obtain the experimental steps with incorrect gesture demonstrations for analysis.

[0067] It should be noted that comprehensively recording and deeply analyzing the user's operations is to provide multi-dimensional data support and decision-making basis for subsequent teaching evaluation and experimental optimization.

[0068] Compared with the prior art, by adopting the experimental simulation method based on MediaPipe and virtual reality shown in this embodiment, through performing spatial component conversion on the gesture image and preprocessing the gesture image through an adaptive histogram equalization algorithm, it helps to more stably identify gesture features under different lighting conditions, improve the quality of the gesture image, lay a foundation for subsequent hand key point calculation, introduce a formula based on the relationship between adjacent key points to correct the hand key points, reduce the occurrence of finger mutual occlusion, and improve the hand key points, thereby solving the technical problems in the prior art that traditional virtual experiments require special wearable devices or handles to implement, and are difficult to operate with high equipment costs.

[0069] Embodiment 2 Please refer to Figure 2 , which shows an experimental simulation system based on MediaPipe and virtual reality provided by the second embodiment of the present invention. The system includes: A virtual scene construction module 100, configured to render a 3D model through the Unity engine to construct a virtual experimental scene and experimental equipment; A gesture image acquisition module 200, configured to acquire a user's gesture image and adjust the mapping of the gesture image through three-dimensional calibration; An image preprocessing module 300, configured to perform spatial component conversion on the gesture image and preprocess the gesture image through an adaptive histogram equalization algorithm to obtain an initial image; A coordinate correction module 400, configured to use the MediaPipe model to identify hand key points in the initial image and correct the position coordinates of the hand key points according to the positional relationship of the hand key points; A simulation experiment demonstration module 500, configured to deeply encapsulate the data information of the hand key points and transmit it to the Unity engine for parsing through the UDP communication protocol to simulate hand movement demonstrations in the Unity engine.

[0070] Compared with the prior art, by adopting the experimental simulation system based on MediaPipe and virtual reality shown in this embodiment, through the image preprocessing module performing spatial component conversion on the gesture image and preprocessing the gesture image through an adaptive histogram equalization algorithm, it helps to more stably identify gesture features under different lighting conditions, improve the quality of the gesture image, lay a foundation for subsequent hand key point calculation, and through the coordinate correction module introducing a formula based on the relationship between adjacent key points to correct the hand key points, reduce the occurrence of finger mutual occlusion, and improve the hand key points, thereby solving the technical problems in the prior art that traditional virtual experiments require special wearable devices or handles to implement, and are difficult to operate with high equipment costs.

[0071] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0072] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable storage medium for an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can obtain instructions In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0073] The above-described embodiments merely represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.

Claims

1. An experimental simulation method based on MediaPipe and virtual reality, characterized in that: The method comprises: Render 3D models through the Unity engine to build virtual experimental scenes and experimental equipment; Collecting a gesture image of a user, and adjusting a mapping of the gesture image through three-dimensional calibration; Performing spatial component conversion on the gesture image, and preprocessing the gesture image by an adaptive histogram equalization algorithm to obtain an initial image; Using the MediaPipe model to identify the hand key points in the initial image, and correcting the position coordinates of the hand key points according to the position relationship of the hand key points; The data information of the key points of the hand is deeply encapsulated and transmitted to the Unity engine for parsing through the UDP communication protocol, and the hand movement demonstration is simulated in the Unity engine.

2. The experimental simulation method based on MediaPipe and virtual reality according to claim 1, characterized in that: The step of collecting the user's gesture image and adjusting the mapping of the gesture image through three-dimensional calibration specifically includes: Based on the spatial position of the camera and the change in shooting angle, the 3D point in the world coordinate system is converted into a 3D point in the camera coordinate system. The formula is as follows: , in, is the three-dimensional point data in the world coordinate system, They are The three-dimensional coordinates of is a 3D point in the camera coordinate system, They are The three-dimensional coordinates of is the rotation matrix, is the translation vector; Then the projection of the camera coordinate system is transformed to the gesture image coordinate system, which is represented by the intrinsic parameter matrix. The formula is as follows: , in, is the internal parameter matrix, , are the focal lengths of the camera in the x-axis and y-axis directions, is the tilt factor.

3. The experimental simulation method based on MediaPipe and virtual reality according to claim 2, characterized in that: The step of performing spatial component conversion on the gesture image specifically includes: Extracting color components from the gesture image, where the color components include R component, G component, and B component, and their value range is between 0 and 1; The color components are converted into color space information, and the formula is as follows: , Among them, V is the brightness information, C is the color difference, S is the saturation, and H is the hue.

4. The experimental simulation method based on MediaPipe and virtual reality according to claim 3, characterized in that: The steps of using the MediaPipe model to identify the hand key points in the initial image and correcting the position coordinates of the hand key points according to the position relationship of the hand key points specifically include: The MediaPipe model is used to identify the hand key points in the initial image, and the distance and angle between adjacent hand key points are calculated. The formula is as follows: , in, is the distance between adjacent hand key points, is the angle between adjacent hand key points, , The key points of the hand of axis, Axis coordinates, , The key points of the hand of axis, Axis coordinates; When it is detected that the distance and the angle are beyond the preset range, the gesture is blocked, and the position coordinates of the key points of the hand are corrected.

5. The experimental simulation method based on MediaPipe and virtual reality according to claim 4, characterized in that: When it is detected that the distance and the angle are beyond the preset range, the gesture is blocked, and the step of correcting the position coordinates of the key points of the hand specifically includes: Calculate hand key points To the hand key point The key vector between is as follows: , in, is the key vector; According to the angle, a rotation matrix is ​​constructed, and the calculation formula is as follows: , in, is the angle difference, is the rotation matrix, is the minimum angle value, is the maximum angle value; Based on the rotation matrix and the key vector, the rotation vector required for rotation is calculated, and the formula is as follows: , in, is the rotation vector; Based on the rotation vector, the hand key points are preliminarily corrected The position coordinates are as follows: , in, , Key points for the hand The position coordinates after preliminary correction, , They are of Quantity, Quantity.

6. The experimental simulation method based on MediaPipe and virtual reality according to claim 5, characterized in that: The method further comprises: The hand key points after the preliminary correction The position coordinates ( , ) Normalize the coordinate values ​​to obtain the current coordinates; Simulate the annealing process, set the initial temperature, and iterate the temperature according to the preset cooling strategy. In each round of iteration, the current coordinates are respectively and Add a random offset in the direction to perturb and get the new coordinates; Calculate the fitness corresponding to the new coordinates, compare it with the fitness of the current coordinates, calculate the difference, and obtain the current coordinates in each iteration based on the difference; When the preset number of iterations is reached, the current coordinates and corresponding fitness in each iteration are obtained, and the current coordinates with the minimum fitness are inversely calculated to calculate the key points of the hand. The final corrected position coordinates.

7. The experimental simulation method based on MediaPipe and virtual reality according to claim 6, characterized in that: The calculation formula of the fitness function is: , Among them, F is the fitness of the hand key points, , is the weight coefficient, Key points for the hand The new coordinates to the hand key points The distance of the position coordinates, For key points The new coordinates to the hand key points The angle of the position coordinates, Hand key points for correct gesture demonstration Position coordinates to hand key points The distance of the position coordinates, Hand key points for correct gesture demonstration Position coordinates to hand key points The angle of the position coordinates.

8. The experimental simulation method based on MediaPipe and virtual reality according to claim 1, characterized in that: The data information of the key points of the hand is deeply encapsulated and transmitted to the Unity engine for parsing through the UDP communication protocol. The steps of simulating the hand action demonstration in the Unity engine specifically include: Deeply encapsulate the data information of the key points of the hand through the ROS message type, wherein the data information includes position coordinates and color space information; Transmitted via UDP communication protocol, after cyclic redundancy check, sent to the Unity engine for analysis; The Unity engine processes the parsed data information into graphics, changes the virtual experimental scene and experimental equipment through matrix transformation, and simulates hand movement demonstration. The formula is as follows: , in, is the data information after the matrix changes, is the translation matrix, is the twist matrix, is the scaling matrix.

9. The experimental simulation method based on MediaPipe and virtual reality according to claim 7, characterized in that: The method further comprises: Each hand key point is used as an independent variable, other hand key points are used as dependent variables, nonlinear features are introduced, and a collaborative regression model is established; Solving the regression coefficient of the collaborative regression model by the least square method, iteratively optimizing the regression coefficient, and obtaining the collaborative motion coefficient between different hand key points; Integrate a series of cooperative motion coefficients of different hand key points to obtain a cooperative matrix, and dynamically optimize the cooperative matrix through a regularization algorithm; According to the optimized coordination matrix, the time delay model is introduced to obtain the coordination angle; The cooperative angle is used to assist in correcting the position coordinates of the key points of the hand.

10. An experimental simulation system based on MediaPipe and virtual reality, characterized in that: The system is used to implement the experimental simulation method based on MediaPipe and virtual reality according to any one of claims 1 to 9, and the system includes: The virtual scene construction module is used to render 3D models through the Unity engine and build virtual experimental scenes and experimental equipment; A gesture image acquisition module, used to acquire the user's gesture image and adjust the mapping of the gesture image through three-dimensional calibration; An image preprocessing module, used for performing spatial component conversion on the gesture image and preprocessing the gesture image through an adaptive histogram equalization algorithm to obtain an initial image; A coordinate correction module, used to identify the hand key points in the initial image using the MediaPipe model, and correct the position coordinates of the hand key points according to the position relationship of the hand key points; The simulation experiment demonstration module is used to deeply encapsulate the data information of the key points of the hand, and transmit it to the Unity engine for parsing through the UDP communication protocol, and simulate the hand movement demonstration in the Unity engine.

Citation Information

Patent Citations

  • High-speed rail simulated driving action virtual and real interaction method based on gesture recognition

    CN114840079A

  • Multi-camera ship height measurement method based on dynamic long baseline

    CN115546280A

  • Virtual reality control method, system and equipment based on hand skeleton and medium

    CN117420917A

  • Experimental AR simulation system based on 3D hand posture estimation method

    CN117648032A

  • Gesture recognition and man-machine interaction method capable of customizing gestures

    CN117765616A