Human-computer interaction method for naked eye 3D

Through the human-computer interaction method of naked-eye 3D, gesture video streaming and depth cameras are used to bind gesture prediction and model, solving the problem that existing gesture interaction methods rely on expensive hardware and poor stability, and achieving low latency, smooth multi-degree of freedom interaction.

CN120371135APending Publication Date: 2025-07-25ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510524606.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing gesture interaction methods rely on expensive hardware devices and have poor interaction stability, limiting their wide application.

Method used

The human-computer interaction method of naked-eye 3D is adopted to collect gesture video streams to make three-dimensional predictions of hand key points, use Kalman filtering and a long-term memory network model that integrates the spatiotemporal attention mechanism to make gesture prediction, and build a hand model in unity for interaction.

Benefits of technology

It realizes low latency, smooth and stable multi-degree-of-freedom human-computer interaction, reducing hardware costs and improving interaction stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371135A_ABST
    Figure CN120371135A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of human-computer interaction, and discloses a human-computer interaction method for naked eye 3D (three-dimensional), which comprises the following steps: acquiring a gesture video stream, and performing three-dimensional prediction on hand key points; performing digital filtering on the plurality of pieces of hand key point information to generate a plurality of pieces of filtered hand key point information, performing simple gesture prediction, and performing complex gesture prediction by using a long-short-term memory network model fused with a space-time attention mechanism; performing digital coding on the simple gestures and the complex gestures, transmitting the coded gestures to unity, and transmitting a plurality of pieces of filtered hand key point information to the unity; acquiring depth values of the hand key points, and fusing the depth values with the filtered information of the plurality of hand key points; constructing a hand model in unity, binding a plurality of pieces of hand key point information fused with depth information, and transmitting coded simple gestures and complex gestures to game control logic in the unity to realize interaction between real gestures and virtual gestures; according to the method, low-delay, smooth and stable multi-degree-of-freedom man-machine interaction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer interaction, and particularly relates to a human-computer interaction method for naked-eye 3D. Background Art

[0002] With the rapid popularization of intelligent terminals such as AR / VR devices, smart homes, in-vehicle systems, etc., the traditional touch and voice interaction modes have gradually shown limitations. As the most natural non-verbal communication method of humans, gestures have the characteristics of intuitiveness, low learning cost, and high freedom, and are regarded as one of the core directions of the next generation of human-computer interaction. However, most of the existing gesture interaction methods rely on expensive hardware sensors or the processors of computers, and have defects such as non-explanation, high cost, and poor use stability, which greatly limit the use of gesture interaction. Summary of the Invention

[0003] In view of the above deficiencies in the prior art, the present invention provides a human-computer interaction method for naked-eye 3D, which is used to solve the defects of high cost and poor interaction use stability caused by the dependence on expensive devices in the existing gesture interaction methods.

[0004] In order to achieve the above invention purpose, the technical solution adopted by the present invention is as follows:

[0005] A human-computer interaction method for naked-eye 3D includes the following steps:

[0006] S1. Collect a gesture video stream, perform three-dimensional prediction of hand key points, and generate a plurality of hand key point information;

[0007] S2. Based on the plurality of hand key point information, perform digital filtering to generate a plurality of filtered hand key point information;

[0008] S3. Based on the plurality of filtered hand key point information, perform simple gesture prediction, which includes prediction of the state of a single finger being erected or bent;

[0009] S4. Based on the plurality of filtered hand key point information, use a long short-term memory network model integrating a spatio-temporal attention mechanism to perform complex gesture prediction, which includes prediction of continuous gestures such as swiping left, swiping right, swiping up, swiping down, and shrinking;

[0010] S5. After digitally encoding the simple gestures and complex gestures, transmit them to the game control logic of unity, and transmit the plurality of filtered hand key point information to unity;

[0011] S6. Set the first sampling position and the second sampling position, use a depth camera to collect two sets of hand movements for depth calibration, generate the depth values of hand key points and transmit them to unity, where they are fused with a number of filtered hand key point information to generate a number of hand key point information with fused depth information;

[0012] S7. Build a hand model in unity. After binding a number of hand key point information with fused depth information, interact with the game control logic to achieve the interaction between the real and virtual gestures.

[0013] The present invention has the following beneficial effects:

[0014] A human-computer interaction method for naked-eye 3D proposed by the present invention, after processing the gesture video stream data, obtains encoded complex gestures and simple gestures, and transmits them to the game control logic of unity. At the same time, a hand model is established in unity and a number of hand key point information with fused depth information is bound. The hand model is controlled by the game control logic to interact, realizing the interaction with virtual gestures, and achieving low-latency, smooth, stable multi-degree-of-freedom human-computer interaction. Brief Description of the Drawings

[0015] Figure 1 It is a schematic flow chart of a human-computer interaction method for naked-eye 3D proposed by the present invention;

[0016] Figure 2 It is a coordinate and position schematic diagram of hand key point information in the embodiment;

[0017] Figure 3 It is a structural schematic diagram of a long short-term memory network model integrating a spatio-temporal attention mechanism in the embodiment;

[0018] Figure 4 It is a structural schematic diagram of the first SNRB block or the second SNRB block in the embodiment;

[0019] Figure 5 It is a structural schematic diagram of a spatio-temporal attention mechanism module in the embodiment;

[0020] Figure 6 It is a structural schematic diagram of a spatial attention module in the embodiment;

[0021] Figure 7 It is a structural schematic diagram of a temporal attention module in the embodiment;

[0022] Figure 8 It is a schematic diagram after the relative coordinates of the sample data of the hand close to the camera in the embodiment are transformed;

[0023] Figure 9Schematic diagram after conversion of relative coordinates of sample data with the hand far from the camera in the embodiment; Figure 10 Schematic diagrams of several different continuous gestures selected in the embodiment. Detailed implementation manners

[0024] The following describes the detailed implementation manners of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed implementation manners. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.

[0025] As Figure 1 shown, a human-computer interaction method for naked-eye 3D includes the following steps:

[0026] S1. Collect a gesture video stream, perform three-dimensional prediction of hand key points, and generate a number of hand key point information.

[0027] Specifically, step S1 specifically includes S11 - S12:

[0028] S11. Turn on the image acquisition device, perform hand simulation, and collect a gesture video stream.

[0029] In this embodiment, the hand simulation is various gesture states such as opening both hands or bending the thumb, etc.; then turn on the image acquisition device and place its position at the same position as the Realsense camera, that is, collect gesture video streams at 0.3 meters and 5 meters; among them, the image acquisition device can be the camera built in the laptop or an external device, such as a mobile phone camera, a video camera, etc. If it is an external device, the device needs to be connected to the laptop, and the purpose is to facilitate checking whether the gesture video stream is successfully collected in the subsequent steps.

[0030] S12. Determine whether the video stream capture function successfully accesses the gesture video stream. If so, execute step S13; otherwise, report an error message and execute step S11.

[0031] S13. Based on the gesture video stream, use the Mediapipe model to perform three-dimensional prediction of hand key points, and generate a number of hand key point information, which includes the three-dimensional coordinate positions of the hand key points.

[0032] In this embodiment, the video stream capture function in the Python OpenCV library is used to turn on the image acquisition device for video shooting. The video stream capture function is: cap = cv2.VideoCapture(index), where index is the number of the image acquisition device. During the process of capturing the gesture video stream, it will be determined whether the captured gesture video stream can be accessed. If it cannot be accessed, an error message will be printed to facilitate recapturing the gesture video stream.

[0033] S2. Based on a number of hand key point information, perform digital filtering to generate a number of filtered hand key point information.

[0034] In this embodiment, during the process of human-computer interaction, gesture recognition only analyzes 21 key point information using the video frames captured by the notebook to transfer and map the key point coordinate information to the virtual hand model. Due to noise interference and external interference, the predicted key point coordinate information has a large jump, which greatly interferes with the stability of the virtual model operation. Therefore, the Kalman filtering method is used to perform digital filtering on a number of hand key point information. By statistically predicting the model changes, the influence of interference is filtered out, making the data changes after model prediction smoother, so that the subsequent human-computer interaction process is more smooth and stable, and the problems of jump and lag are solved. The main features of the Kalman filter design are to set the calculation frequency to the millisecond level to meet the requirements of real-time interaction. The prediction confidence of this method is better than the information degree of the real collected data, enhancing the anti-interference ability, improving the stability and smoothness of the interaction operation, and setting a smaller initial covariance to reduce the number of iterations, thereby accelerating the data convergence speed, as follows:

[0035] Specifically, step S2 specifically includes S21 - S27:

[0036] S21. Input a number of hand key point information, initialize the state vector and covariance matrix of the Kalman filter, and at the same time create a 6-dimensional measurement matrix. After using the 6-dimensional measurement matrix to represent the coordinates and velocities of the x-axis, y-axis, and z-axis of a number of hand key point information respectively, set the state change equation using the uniform change model and create a state transition matrix;

[0037] Among them, the state change equation is:

[0038]

[0039] Among them, x(t+dt), y(t+dt), and z(t+dt) respectively represent the x-axis, y-axis, and z-axis coordinates of the hand key-point information at the (t+dt)-th moment, x(t), y(t), and z(t) respectively represent the x-axis, y-axis, and z-axis coordinates of the hand key-point information at the t-th moment, x, y, and z respectively represent the x, y, and z-axis coordinates of the hand key-point information, v represents the speed of the hand key-point information, and dt represents the differential of the moment t.

[0040] S22. Predict the state vector of a number of hand key-point information at the next moment by using the state transition matrix, that is:

[0041] x k|k-1 = Fx k-1

[0042] Among them, x k|k-1 represents the state vector of the k-th step based on the (k-1)-th step, F represents the state transition matrix, and x k-1 represents the state vector of the (k-1)-th step.

[0043] S23. Update the predicted covariance matrix, that is:

[0044] P k|k-1 = FP k-1 F T + Q

[0045] Among them, P k|k-1 represents the covariance matrix of the k-th step based on the (k-1)-th step, P k-1 represents the covariance matrix of the (k-1)-th step, T represents the transpose, and Q represents the process noise covariance matrix.

[0046] S24. Calculate the Kalman gain, that is:

[0047] K k = P k|k-1 H T (HP k|k-1 H T + R) -1

[0048] Among them, K k represents the Kalman gain of the k-th step, H represents the observation matrix, and R represents the noise covariance matrix.

[0049] S25. Update the state vector by using the Kalman gain, that is:

[0050] x k = x k|k-1 + K k (z k - Hx k|k-1 )

[0051] Among them, z k represents the observation value of the k-th step, and x k represents the state vector of the k-th step.

[0052] S26. Update the covariance matrix, that is:

[0053] P k =(I - K k H)P k|k-1

[0054] Among them, P k represents the covariance matrix of the k-th step, and I represents the identity matrix.

[0055] S27. Generate a number of hand key point information for filtering.

[0056] S3. Based on a number of hand key point information for filtering, perform simple gesture prediction, which includes predicting the state of a single finger being raised or bent.

[0057] In this embodiment, for simple gesture prediction, the simple gesture refers to the state of a single finger being raised or bent. The prediction process is specifically as follows: The coordinate information of 21 gesture key points after Kalman filtering is put into a list for storage. Its storage form is [id][x, y, z], where id represents the number corresponding to the gesture key point, and x, y, and z respectively represent the x, y, and z axis coordinate positions of the hand, specifically as Figure 2 shown. Based on Figure 2 the numbers and relationships of each finger shown, by comparing the x coordinates of the middle finger and the fingertip of each finger, the state of the finger, that is, whether it is bent or raised, can be obtained. The specific prediction process is as follows:

[0058] Specifically, step S3 specifically includes S31 - S32:

[0059] S31. Number and store a number of hand key point information for filtering, specifically:

[0060] Number the wrist of the right palm or left palm as 0. At the same time, divide the thumb, index finger, middle finger, ring finger, and little finger of the right palm or left palm into four parts from the root to the fingertip, which are the root of the finger, between the fingers, the middle of the finger, and the fingertip respectively. And in the order from the thumb to the little finger, take the root of the thumb as the initial node and the fingertip of the little finger as the end node, and number from 1 to obtain the numbers and coordinates of 21 hand key points.

[0061] S32. Select the middle and fingertip of a single finger for comparison, and judge whether the x coordinate of the middle of the finger is greater than the x coordinate of the fingertip of the finger. If so, the finger is raised; otherwise, the finger is bent, and finally generate the predicted simple gesture.

[0062] S4. Based on the information of several hand key points after filtering, a long short-term memory network model with a fused spatio-temporal attention mechanism is used for complex gesture prediction, including continuous gesture predictions of swiping left, swiping right, swiping up, swiping down, and zooming out.

[0063] In this embodiment, Figure 3 the structure and connection relationship of the long short-term memory network model with a fused spatio-temporal attention mechanism are shown, including a first branch, a second branch, a first feature fusion module, a global average pooling module, a bidirectional long short-term memory network, a spatio-temporal attention mechanism module, a second feature fusion module, a first fully connected layer, a dropout module, a second fully connected layer, and a normalization layer.

[0064] Among them, the first branch is the first SNRB block; the second branch block includes a first average pooling module and a second SNRB block; and both the first SNRB block and the second SNRB block include a depthwise separable convolutional layer, an RLUE non-linear activation function layer, and a batch normalization layer, as specifically shown in Figure 4 It is as follows: (1) The advantage of depthwise separable convolution is that it has fewer parameters and higher computational efficiency, which is suitable for mobile or resource-constrained environments; this is very important for gesture recognition, especially for gesture recognition involving real-time processing on mobile devices. Since the number of parameters of ordinary convolution is the number of input channels × the size of the convolution kernel × the number of output channels, while depthwise separable convolution decomposes it into depthwise convolution and pointwise convolution, thus greatly reducing the parameters. For example: if the input channels of the input image are 256, the convolution kernel size is 5, and the output is 64; the parameters of its ordinary convolution are 256 × 5 × 64 = 81,920, while the parameters of depthwise separable convolution are 256 × 5 (depthwise convolution) + 256 × 64 (pointwise convolution) = 1,280 + 16,384 = 17,664, and the parameters are reduced by about 78%, which helps with model compression and accelerating training, and at the same time can reduce the risk of overfitting; (2) The batch normalization module (BN) can accelerate training, reduce internal covariate shift, and make the input distribution of each layer more stable; especially in deep networks, the gradient propagation is more stable, allowing the use of a higher learning rate. Therefore, adding a BN layer after the convolutional layer can make the input distribution of the ReLU activation function more stable, avoiding the problem of gradient vanishing or exploding. In addition, BN also has a slight regularization effect, which helps to prevent overfitting. Therefore, the above three modules work together, using depthwise separable convolution to be responsible for efficient feature extraction, using the ReLU non-linear activation function to introduce non-linear dynamics, and using the batch normalization layer to ensure training stability, so that the model can reduce resource consumption such as MobileNet and EfficientNet while taking into account performance and inference speed, especially suitable for mobile or real-time scenario deployment, and balancing the lightweight and accuracy requirements.

[0065] Among them, the spatio-temporal attention mechanism module includes a spatial attention module and a temporal attention module, specifically as Figure 5 shown; the connection relationship and structure of the spatial attention module are as Figure 6 shown, including a second average pooling module, a max pooling module, a third fully connected layer, and a fourth fully connected layer. Its design principle is as follows: The spatial attention module adopts a spatial attention mechanism, introduces an average pooling layer and a max pooling layer, and captures channel information to more comprehensively reflect the importance of channels. The connection relationship and structure of the temporal attention module are as Figure 6 shown, including a fifth fully connected layer, a sixth fully connected layer, a tanh activation function layer, and a softmax activation function layer; its design principle is as follows: The temporal attention module adopts a temporal attention mechanism. Through dynamic weight allocation, the model can adaptively focus on key action stages, significantly improving the temporal modeling ability while ensuring computational efficiency; during actual deployment, the size of units can be adjusted according to specific hardware conditions to achieve the best balance between accuracy and speed; at the same time, two fully connected layers (the fifth fully connected layer and the sixth fully connected layer) are used to project the query vector into the same semantic space as the value vector and enhance the features of the original value vector; the tanh activation function restricts the score to the interval [-1, 1] to prevent softmax saturation and is more suitable for attention score calculation than the ReLU activation function to retain negative correlations, and the softmax activation function performs a 0-1 probability mapping change on the features; therefore, constructing a temporal attention mechanism can focus on the sequence relationship of key frames and prevent misjudgment of different sequences of the same gesture. In addition, the temporal attention module uses the features of the last time step as the query. The features of the last moment contain summary information, which is used to interact with all time steps and can automatically identify important sequence information and suppress environmental noise.

[0066] In summary, the long short-term memory network model with a fused spatio-temporal attention mechanism designed by the present invention uses the long short-term memory network as the base layer, increases the ability of the model to pay attention to information in long time series by adding multi-scale feature extraction (which is reflected in the multi-scale feature extraction of the first branch and the second branch) and the spatio-temporal attention mechanism, aggregates the key information of multiple sequences, and greatly increases the processing speed and accuracy of the model for sequence frame data.

[0067] In addition, the training process of the model is as follows: (1) Using a self-built dataset, only 10 samples are required for each gesture action to train a good effect. The images collected by the camera are predicted for key point information through the Mediapipe model using the python opencv library, and 30 frames of data of continuous gesture information are sequentially collected and written into an npy file as a sample set. The data volume of each sample is 3*21*2, where 3 represents three-dimensional coordinate information, 21 represents 21 key points, and 2 represents data of two hands; (2) Data preprocessing, including: 1) Relative coordinate transformation, that is: the coordinate information of the key points is transformed into the key point coordinates centered on the hand: X′ i 、Y′ i 、Z′ i represent the x-axis coordinate, y-axis coordinate, and z-axis coordinate of the i-th key point after data processing respectively, and X i 、Y i 、Z i represent the x-axis coordinate, y-axis coordinate, and z-axis coordinate of the i-th key point before data processing respectively, and X0, Y0, Z0 represent the x-axis coordinate, y-axis coordinate, and z-axis coordinate of key point 0 respectively; the purpose of relative coordinate transformation is: to eliminate the influence of the hand position; 2) After the key point coordinates are transformed, mean square error normalization is performed, that is: X″ i =(X′ i -X m ) / X s , X m represents the average value of the x-axis coordinates of 21 hand key points, X s represents the standard deviation of the sample data, and X″ i represents the x-axis coordinate of the i-th key point after mean square error standardization; the purpose of standardization is: to eliminate the influence of distance; the relative coordinate transformation results of the sample data with the hand close to the camera and the sample data with the hand far from the camera are shown in Figure 8 、 Figure 9 respectively. Therefore, after relative coordinate transformation, the coordinate of the first key point is reduced to 0 and then removed, so the input sample information is 3*20*2, which is expanded into a 120*1 matrix. In addition, a total of 5 different continuous gestures are used in this experiment, namely left swipe, right swipe, up swipe, down swipe, and shrink, as shown in Figure 10 ; (3) After the above sample data is processed, it is used as the training dataset and input into the long short-term memory network model integrating spatio-temporal attention mechanism for training, that is, the 30 frames of the above input data, the 21*3*2 hand key points predicted by Mediapipe analysis, the input is 30*126, and the input after data processing is 30*120; a better model is obtained through 1000 rounds of training and the weight file is saved, and then the pre-trained weight file is used for complex gesture recognition analysis in prediction, specifically as follows:

[0068] Specifically, step S4 specifically includes S401 - S410:

[0069] S401. Input the filtered information of several hand key points into the first SNRB block of the first branch for up - sampling feature extraction to generate the first feature. Specifically:

[0070] S4011. Input the filtered information of several hand key points into the depth - separable convolutional layer for depth convolution and point - wise convolution operations to generate depth features.

[0071] S4012. Input the depth features into the RLUE non - linear activation function layer for non - linear mapping to generate non - linear features.

[0072] S4013. Input the non - linear features into the batch normalization layer for batch standard normalization to generate the first feature.

[0073] S402. Input the filtered information of several hand key points into the first average pooling module and the second SNRB block of the second branch in sequence for down - sampling feature extraction to generate the second feature.

[0074] S403. Input the first feature and the second feature into the first feature fusion module for feature fusion to generate the third feature.

[0075] S404. Input the third feature into the bidirectional long - short - term memory network to generate the fourth feature by capturing the relationship between features before and after. At the same time, input the third feature into the global average pooling module for global feature pooling to generate the fifth feature.

[0076] S405. Input the fourth feature into the spatio - temporal attention mechanism module. After focusing on key information and adjusting important feature channels, generate the sixth feature. Specifically:

[0077] S4051. Input the fourth feature into the second average pooling module and the max - pooling module of the spatial attention module for average pooling and max - pooling operations respectively. Then input the average - pooled features into the third fully - connected layer and the max - pooled features into the fourth fully - connected layer for fully - connected operations. After that, perform element - wise addition on the output features of each layer to generate spatial attention features.

[0078] S4052. Input the spatial attention features into the fifth fully - connected layer and the sixth fully - connected layer of the temporal attention module for fully - connected operations respectively. Then perform element - wise addition on the output features of each layer, input them into the tanh activation function layer for non - linear mapping, then input them into the softmax function layer for 0 - 1 interval mapping. Finally, multiply the mapped features and the spatial attention features element - by - element to generate the fifth feature.

[0079] S406. Input the fifth feature and the sixth feature into the second feature fusion module respectively for feature fusion to generate the seventh feature.

[0080] S407. Input the seventh feature into the first fully connected layer for linear feature integration to generate the eighth feature.

[0081] S408. Input the eighth feature into the dropout module for random feature dropout to generate the ninth feature.

[0082] S409. Input the ninth feature into the second fully connected layer for linear feature integration to generate the tenth feature.

[0083] S410. Input the tenth feature into the normalization layer for normalization operation to obtain the classified complex gestures.

[0084] S5. After digitally encoding the simple gestures and the complex gestures, transmit them to the game control logic of unity, and transmit the filtered information of several hand key points to unity.

[0085] S6. Set the first sampling position and the second sampling position, use the depth camera to collect two groups of hand movements for depth calibration, generate the depth values of the hand key points and then transmit them to unity, and fuse them with the filtered information of several hand key points to generate the information of several hand key points with fused depth information.

[0086] Specifically, in step S6, the first sampling position and the second sampling position are 0.3 meters and 5 meters respectively.

[0087] Specifically, in step S6, the hand movement is the state of spreading out all five fingers.

[0088] In this embodiment, it is a virtual reality operation process for enhancing the immersion of human-computer interaction. The hardware used is the Realsense depth camera, whose function is to assist the Mediapipe model in predicting and calibrating the depth information of the hand, and send the spatial three-dimensional information of the predicted and calibrated key points to the unity game project (i.e., the hand model) through the network port, so as to map and assign the information collected in reality to the position information of the hand in the virtual game.

[0089] S7. Build a hand model in unity, bind the information of several hand key points with fused depth information, and then interact with the game control logic to achieve the interaction between real and virtual gestures.

[0090] In this embodiment, the interaction is as follows: by mapping real gesture actions to the game control logic in Unity, the game control logic is used to control the virtual hand model. That is, while recognizing real gestures, the virtual hand model can also move following the real gestures, realizing the mapping between the real world and the virtual world and enabling good immersive interaction.

[0091] In summary, a human-computer interaction method for naked-eye 3D proposed by the present invention can achieve low-latency, smooth, and stable multi-degree-of-freedom human-computer interaction.

[0092] In the present invention, specific embodiments are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0093] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principle of the present invention and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations without departing from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A human-computer interaction method for naked-eye 3D, characterized in that, It includes the following steps: S1. Collect a gesture video stream, perform three-dimensional prediction of hand key points, and generate a number of hand key point information; S2. Based on a number of hand key point information, perform digital filtering to generate a number of filtered hand key point information; S3. Based on a number of filtered hand key point information, perform simple gesture prediction, which includes prediction of the state of a single finger being erected or bent; S4. Based on a number of filtered hand key point information, use a long short-term memory network model with a fused spatio-temporal attention mechanism to perform complex gesture prediction, which includes prediction of continuous gestures such as swiping left, swiping right, swiping up, swiping down, and shrinking; S5. After digitally encoding the simple gestures and complex gestures, transmit them to the game control logic of unity, and transmit a number of filtered hand key point information to unity; S6. Set the first sampling position and the second sampling position, use a depth camera to collect two sets of hand actions for depth calibration, generate hand key point depth values and transmit them to unity, and fuse them with a number of filtered hand key point information to generate a number of hand key point information with fused depth information; S7. Build a hand model in unity, bind a number of hand key point information with fused depth information, and interact with the game control logic to achieve the interaction between real and virtual gestures.

2. The human-computer interaction method for naked-eye 3D according to claim 1, wherein, Step S1 specifically includes: S11. Turn on the image acquisition device, perform hand simulation, and collect a gesture video stream; S12. Determine whether the video stream capture function successfully accesses the gesture video stream. If so, execute step S13; otherwise, report an error message and execute step S11; S13. Based on the gesture video stream, use the Mediapipe model to perform three-dimensional prediction of hand key points, and generate a number of hand key point information, which includes the three-dimensional coordinate positions of the hand key points.

3. The human-computer interaction method for naked-eye 3D according to claim 2, wherein, Step S2 specifically includes: S21. Input a number of hand key point information, initialize the state vector and covariance matrix of the Kalman filter, and at the same time create a 6D measurement matrix. After using the 6D measurement matrix to represent the coordinates and velocities of the x-axis, y-axis, and z-axis of a number of hand key point information respectively, set the state change equation using a uniform change model and create a state transition matrix; Among them, the state change equation is: Among them, x(t+dt), y(t+dt), and z(t+dt) respectively represent the coordinates of the x-axis, y-axis, and z-axis of the hand key point information at the (t+dt)th moment, x(t), y(t), and z(t) respectively represent the coordinates of the x-axis, y-axis, and z-axis of the hand key point information at the tth moment, x, y, and z respectively represent the x, y, and z-axis coordinates of the hand key point information, v represents the velocity of the hand key point information, and dt represents the differential at time t; S22. Use the state transition matrix to predict the state vector of a number of hand key point information at the next moment, that is: x k|k-1 = Fx k-1 where x k|k-1 represents the state vector of the k-th step based on the (k - 1)-th step, F represents the state transition matrix, and x k-1 represents the state vector of the (k - 1)-th step; S23. Update the predicted covariance matrix, that is: P k|k-1 = FP k-1 F T + Q where, P k|k-1 represents the covariance matrix of the k-th step based on the (k-1)-th step, P k-1 represents the covariance matrix of the (k-1)-th step, T represents transpose, and Q represents the process noise covariance matrix; S24. Calculate the Kalman gain, that is: K k = P k|k-1 H T (HP k|k-1 H T + R) -1 Among them, K k represents the Kalman gain at the k-th step, H represents the observation matrix, and R represents the noise covariance matrix; S25. Update the state vector using the Kalman gain, that is: x k = x k|k-1 + K k (z k - Hx k|k-1 ) where z k represents the observation value at the k-th step, and x k represents the state vector at the k-th step; S26. Update the covariance matrix, that is: P k = (I - K k H)P k|k-1 Among them, P k represents the covariance matrix of the k-th step, and I represents the identity matrix; S27. Generate a number of filtered hand key point information.

4. The human-computer interaction method for naked-eye 3D according to claim 3, wherein Step S3 specifically includes: S31. Number and store the filtered information of several hand key points, specifically: Number the wrist of the right palm or left palm as 0. At the same time, divide the thumb, index finger, middle finger, ring finger, and little finger of the right palm or left palm from the finger root to the fingertip into four parts, namely the finger root, interphalangeal, middle finger, and fingertip. In the order from the thumb to the little finger, take the finger root of the thumb as the initial node and the fingertip of the little finger as the termination node, and number from 1 to obtain the numbers and coordinates of 21 hand key points; S32. Select the middle finger and fingertip of a single finger for comparison, and judge whether the x coordinate of the middle finger of the finger is greater than the x coordinate of the fingertip of the finger. If so, the finger is erected; otherwise, the finger is bent, and finally a predicted simple gesture is generated.

5. The human-computer interaction method for naked-eye 3D according to claim 4, characterized in that, The long short-term memory network model integrating the spatio-temporal attention mechanism in step S4 includes a first branch, a second branch, a first feature fusion module, a global average pooling module, a bidirectional long short-term memory network, a spatio-temporal attention mechanism module, a second feature fusion module, a first fully connected layer, a dropout module, a second fully connected layer, and a normalization layer; The first branch is the first SNRB block; the second branch block includes a first average pooling module and a second SNRB block; Both the first SNRB block and the second SNRB block include a depthwise separable convolutional layer, an RLUE non-linear activation function layer, and a batch normalization layer; The spatio-temporal attention mechanism module includes a spatial attention module and a temporal attention module; The spatial attention module includes a second average pooling module, a max pooling module, a third fully connected layer, and a fourth fully connected layer; The temporal attention module includes a fifth fully connected layer, a sixth fully connected layer, a tanh activation function layer, and a softmax activation function layer.

6. The human-computer interaction method for naked-eye 3D according to claim 5, characterized in that Step S4 specifically includes: S401. Input the filtered information of several hand key points into the first SNRB block of the first branch for upsampling feature extraction to generate a first feature; S402. Input the filtered information of several hand key points into the first average pooling module and the second SNRB block of the second branch in sequence for downsampling feature extraction to generate a second feature; S403. Input the first feature and the second feature into the first feature fusion module for feature fusion to generate a third feature; S404. Input the third feature into the bidirectional long short-term memory network to generate a fourth feature by capturing the relationship between features before and after. At the same time, input the third feature into the global average pooling module for global feature pooling to generate a fifth feature; S405. Input the fourth feature into the spatio-temporal attention mechanism module. After focusing on key information and adjusting important feature channels, generate a sixth feature; S406. Input the fifth feature and the sixth feature into the second feature fusion module respectively for feature fusion to generate a seventh feature; S407. Input the seventh feature into the first fully connected layer for linear feature integration to generate an eighth feature; S408. Input the eighth feature into the dropout module for random feature dropout to generate a ninth feature; S409. Input the ninth feature into the second fully connected layer for linear feature integration to generate a tenth feature; S410. Input the tenth feature into the normalization layer for normalization operation to obtain the classified complex gesture.

7. The human-computer interaction method for naked-eye 3D according to claim 6, wherein Step S401 specifically includes: S4011. Input the filtered information of several hand key points into the depthwise separable convolutional layer for depth convolution and pointwise convolution operations to generate depth features. S4012. Input the depth features into the RLUE non-linear activation function layer for non-linear mapping to generate non-linear features. S4013. Input the non-linear features into the batch normalization layer for batch standard normalization to generate the first feature.

8. The human-computer interaction method for naked-eye 3D according to claim 7, wherein, Step S405 specifically includes: S4051. After inputting the fourth feature into the second average pooling module and the max pooling module of the spatial attention module for average pooling and max pooling operations respectively, input the average pooling features into the third fully connected layer and the max pooling features into the fourth fully connected layer for fully connected operations, and then perform element-wise summation on the output features of each layer to generate spatial attention features. S4052. After inputting the spatial attention features into the fifth fully connected layer and the sixth fully connected layer of the temporal attention module for fully connected operations respectively, perform element-wise summation on the output features of each layer, input them into the tanh activation function layer for non-linear mapping, then input them into the softmax function layer for 0-1 interval mapping, and finally multiply the mapped features and the spatial attention features element-wise to generate the fifth feature.

9. The human-computer interaction method for naked-eye 3D according to claim 8, characterized in that In step S6, the first sampling position and the second sampling position are 0.3 meters and 5 meters respectively.

10. The human-computer interaction method for naked-eye 3D according to claim 9, wherein In step S6, the hand gesture is the state of spreading out all five fingers.