System and method for hand-pen interaction in virtual reality scene based on vision

By adopting hand three-dimensional reconstruction and hand-written interaction mapping technology in virtual reality scenarios, the problem of unnatural hand-written interaction in the existing VR painting system is solved, and efficient and accurate naked hand-written interaction is achieved, improving the immersion of the painting experience and the flexibility of operation.

CN119960600AActive Publication Date: 2025-05-09UNIV OF SCI & TECH BEIJING

Patent Information

Application Number
CN202510042969.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-09
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

The existing VR painting system has a disconnect between visual input and physical action, which increases the user's cognitive load and lacks a carefully designed interaction paradigm, making it difficult to provide natural, flexible and accurate hand interaction.

Method used

The hand-written interaction system in virtual reality scenarios is adopted, and the natural interaction between the naked hand and the virtual pen is realized through the hand three-dimensional reconstruction module, the hand-written interaction mapping module and the interactive presentation module. The system uses graph neural network to perform three-dimensional reconstruction of two-hands, recognizes pen-holding gestures, and maps gestures to basic functions in virtual reality scenes.

Benefits of technology

It effectively reduces cognitive load and operational threshold, provides an immersive painting experience, realizes natural, flexible and accurate hand interaction, and supports complex painting operations and fine-grained creative control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960600A_ABST
    Figure CN119960600A_ABST
Patent Text Reader

Abstract

The invention discloses a vision-based hand-pen interaction system in a virtual reality scene. The system comprises a hand three-dimensional reconstruction module, a hand-pen interaction mapping module and an interaction presentation module, the invention discloses a vision-based hand-pen interaction method in a virtual reality scene. The method comprises the following steps: S1, acquiring a hand image of a user visual angle through a visual sensor; s2, inputting the hand image into the constructed graph neural network to carry out two-hand three-dimensional reconstruction to obtain a model; s3, presetting a grasping posture, and distinguishing the change of each tiny gesture by identifying the difference between the actual posture and the preset posture; s4, designing several kinds of tiny gestures which are easy to distinguish, and mapping the tiny gestures into all functions of the drawing system to form a hand-pen interaction drawing system; according to the method, the user can intuitively create in the virtual environment, extra hardware is not needed, the cognitive load and the operation threshold are effectively reduced, and the immersive drawing experience is provided for the user in cooperation with an audio-visual feedback strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of virtual reality and human-computer interaction, and in particular to a handwriting interaction system and method in a vision-based virtual reality scene. Background Art

[0002] With the rapid development of virtual reality technology, VR painting and sketching applications are becoming more and more popular among artists and designers. These applications provide users with an immersive creative environment and intuitive spatial interaction methods, changing the traditional way of creation. In the early stages of design, hand-drawn sketches play a vital role in quickly exploring and communicating design concepts.

[0003] Existing VR painting systems mainly use hardware-based interaction methods, such as six-degree-of-freedom handheld controllers, styluses, VR drawing boards, etc. These controller-based interactions require users to judge the movement of devices outside their sight while watching the movement of the screen cursor. There is a disconnect between visual input and physical actions, which increases the user's cognitive load. Users need to concentrate to coordinate the two, which cannot provide the most natural and intuitive hand-drawing experience.

[0004] In response to the above problems, some researchers have begun to explore bare-hand interaction as an alternative to VR painting. Dudley et al. explored the possibility of three-dimensional drawing with bare hands in an AR environment. Vartak et al. proposed a virtual gesture painting system that uses hand gestures to draw lines on the screen. Rajail et al. developed a machine learning method to achieve virtual line drawing on the screen through hand gestures. These studies show that bare-hand interaction is expected to enhance the naturalness and expressiveness of VR painting.

[0005] However, existing bare-hand VR painting systems often lack a well-designed interaction paradigm to accurately identify user intent and support complex interactive operations. In professional-level VR painting applications, users need to use a variety of tools and controls, and existing systems have limited gesture recognition capabilities and are unable to provide fine painting control. The lack of natural, flexible, and precise hand interaction seriously affects the VR painting experience.

[0006] In addition, the natural hand-drawing process is often accompanied by frequent changes in brush strokes and adjustments to painting parameters, such as changing brush size and switching colors. Existing gesture interaction solutions are difficult to support such fine-grained operations smoothly and quickly, resulting in frequent interruptions in the interaction during the painting process. At the same time, the instability of the user's hand posture when gesturing in the air also poses a challenge to gesture recognition.

[0007] Therefore, in order to achieve natural bare-hand interaction in VR painting, an advanced gesture interaction paradigm is urgently needed that can accurately identify user intentions, provide rich painting function mappings, and combine intelligent drawing assistance and feedback mechanisms, ultimately bringing users a seamless, smooth, and fine-grained creative experience and unleashing the creativity of VR painting.

[0008] Recent studies have explored bare-hand interaction as an alternative to VR painting, but existing bare-hand VR painting systems often lack a well-designed interaction paradigm to accurately identify user intent and support complex interactive operations. The lack of natural, flexible, and precise hand interaction seriously affects the VR painting experience. Summary of the invention

[0009] The present invention provides a handwriting interaction system and method in a vision-based virtual reality scene, which can be widely used in scenes such as virtual reality painting and sketching, allowing users to create in a virtual environment intuitively without the need for additional hardware, effectively reducing cognitive load and operation threshold, and cooperating with audio-visual feedback strategies to provide users with an immersive painting experience.

[0010] The specific technical solutions are as follows:

[0011] A hand-pen interaction system in a vision-based virtual reality scene, the system comprising: a hand three-dimensional reconstruction module, a hand-pen interaction mapping module and an interaction presentation module;

[0012] The hand three-dimensional reconstruction module is connected to the hand-pen interaction mapping module, and the hand-pen interaction mapping module is connected to the interaction presentation module.

[0013] Preferably, the hand 3D reconstruction module is used to collect the RGB image sequence of the user's hands in real time through the RGB camera on the head display device, input it into the perturbation graph contrast learning network DGCLNet, extract the graph feature representation of the gesture through the graph U-Net, and optimize the graph features through the cosine similarity loss function to achieve accurate recognition of the gesture, and then map the gesture graph features to the vertex coordinates and joint rotation angles of the MANO hand model through the graph convolution upsampling network and the spatial transformation network, so as to reconstruct the 3D surface shape and joint posture of the hand.

[0014] Preferably, the hand-pen interaction mapping module is used to map the recognized gestures with the interactive operations of the virtual brush, and realize the virtual pen grabbing, menu calling and undo operations through pinching and multi-finger combination gestures, thereby realizing natural interaction between bare hands and the virtual pen.

[0015] Preferably, the interactive presentation module is used to control the virtual brush to paint in the three-dimensional virtual scene according to the result of the hand-pen interaction, and present the three-dimensional model of the hand, the state of the virtual pen and the painting result in real time in the virtual scene to provide visual feedback, and give audio prompts according to the interaction results to enhance the user's immersion and operation feedback.

[0016] Preferably, a handwriting interaction method in a vision-based virtual reality scene comprises:

[0017] S1, obtain the hand image from the user's perspective through the visual sensor;

[0018] S2. Input the hand image into the constructed graph neural network to perform three-dimensional reconstruction of both hands to obtain a model, wherein the model includes a backbone network, a neck network, and a prediction head;

[0019] S3, creating a virtual pen in a virtual reality scene, and classifying four pen-holding gestures by identifying the slight changes of each finger when holding the pen;

[0020] S4. Based on the four pen-holding gestures, the basic functions of the system in the virtual reality scene are mapped to form a hand-pen interaction paradigm.

[0021] Preferably, S2 includes the following sub-steps:

[0022] Sub-step S21, obtaining images of the user's hands;

[0023] Sub-step S22, inputting the images of both hands into the backbone network to extract image features, and inputting the image features into two independent multi-layer perceptrons MLP for left and right hands respectively for feature conversion;

[0024] Sub-step S23, using the Laplacian matrix to construct the features output by the multi-layer perceptron into graph features;

[0025] Sub-step S24: input the graph features into the perturbation graph contrast learning network DGCLNet, regenerate the graph features through the graph U-Net to form feature pairs,

[0026] Sub-step S25, the feature pairs are passed through a shared graph convolutional network GCN to maximize the similarity of the two output features;

[0027] Sub-step S26, inputting the graph features of the original branch into the upsampling layer, and obtaining the three-dimensional coordinates of 778 vertices of both hands in a MANO manner;

[0028] Sub-step S27, introducing a similarity distinction method to accurately calculate the similarity of feature pairs from the feature-pose pool to accelerate model convergence;

[0029] Sub-step S28, designing multiple loss functions to balance the relationship between hand posture, visual scale and viewpoint position to improve the quality of three-dimensional reconstruction;

[0030] Sub-step S29, iteratively optimize the model until convergence to obtain the final two-handed 3D reconstruction graph neural network model.

[0031] Preferably, S3 includes the following sub-steps:

[0032] Sub-step S31, creating a virtual body of a pen in a virtual reality environment, and grabbing the virtual pen by presetting a grabbing gesture on the virtual pen;

[0033] Sub-step S32, when the hand interacts with the virtual pen, measuring the error between the virtual hand and each finger joint of the preset gesture, and calculating the average position error of each joint;

[0034] Sub-step S33, by setting the error threshold of each finger joint position, determining whether the finger reaches the preset gesture position, and accurately identifying each finger gesture;

[0035] Sub-step S34, setting 4 finger gestures corresponding to four functions respectively, the 4 finger gestures include the pinching position of the thumb and index finger, the thumb, index finger and middle finger reaching the preset gesture finger position at the same time, the thumb, index finger and ring finger bending position at the same time, and all fingers reaching the preset position at the same time.

[0036] Preferably, S4 includes the following sub-steps:

[0037] Sub-step S41, the position of the thumb and index finger pinching corresponds to the function of moving the virtual pen, the virtual pen is selected by pinching the index finger and thumb, and the selection is cancelled by releasing, which is used to control the position of the virtual pen in the three-dimensional space;

[0038] Sub-step S42, the thumb, index finger and middle finger simultaneously reach the preset gesture finger position corresponding to calling out the menu function;

[0039] Sub-step S43, the thumb, index finger and ring finger are bent at the same time to the corresponding position to cancel the previous operation function;

[0040] Sub-step S44: When all fingers reach the preset position at the same time, it corresponds to the drawing operation.

[0041] The beneficial effects of the handwriting interaction system and method in a visual virtual reality scene of the present invention are as follows:

[0042] 1. The present invention integrates the main operations in the painting process into the single-hand compound hand-pen interaction paradigm, simplifying the drawing operations in three-dimensional space.

[0043] 2. The hand-pen interaction system can be widely used in virtual reality painting, sketching and other scenes, allowing users to create intuitively in a virtual environment without the need for additional hardware, effectively reducing cognitive load and operation thresholds, and cooperating with audio-visual feedback strategies to provide users with an immersive painting experience.

[0044] 3. The present invention has broad application prospects in the field of virtual reality interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is the overall flow chart of the present invention.

[0046] Figure 2 It is a schematic diagram of the overall architecture of the graph neural network of the present invention.

[0047] Figure 3 It is a schematic diagram of hand-pen interaction of the present invention.

[0048] Figure 4 It is a structural diagram of the handwriting interaction paradigm of the present invention.

[0049] Figure 5 It is a schematic diagram of preventing the pen tip from penetrating the canvas according to the present invention.

[0050] Figure 6 It is a schematic diagram of visual feedback of the present invention.

[0051] Figure 7 It is a schematic diagram of auditory feedback of the present invention.

[0052] Figure 8 It is a schematic diagram of the results created by the user of the present invention in this system. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] A handwriting interaction method in a virtual reality scene based on vision of the present invention comprises:

[0055] S1, obtain the hand image from the user's perspective through the visual sensor;

[0056] S2. Input the hand image into the constructed graph neural network to perform 3D reconstruction of both hands to obtain a model, wherein the model includes a backbone network, a neck network, and a prediction head;

[0057] S3, creating a virtual pen in a virtual reality scene, and classifying four pen-holding gestures by identifying the slight changes of each finger when holding the pen;

[0058] S4. Based on the four pen-holding gestures, the basic functions of the system in the virtual reality scene are mapped to form a hand-pen interaction paradigm.

[0059] S2 of this embodiment includes:

[0060] S21, obtaining images of both hands of the user;

[0061] S22, input the images of both hands into the backbone network to extract image features, and then input the features into two independent multi-layer perceptrons (MLP) for left and right hands respectively for feature conversion;

[0062] S23, using the Laplacian matrix to construct the features output by the MLP into graph features;

[0063] S24, input the graph features into the perturbation graph contrast learning network DGCLNet, regenerate the graph features through the graph U-Net to form feature pairs, and the feature pairs pass through the shared graph convolutional network GCN to maximize the similarity of the two output features;

[0064] S25, input the graph features of the original branch into the upsampling layer, and obtain the three-dimensional coordinates of 778 vertices of both hands in a MANO manner;

[0065] S26, introduce similarity distinction method to accurately calculate the similarity of feature pairs from the feature-pose pool to accelerate model convergence;

[0066] S27. Design multiple loss functions to balance the relationship between hand posture, visual scale and viewpoint position to improve the quality of 3D reconstruction. The loss functions include:

[0067] (1) The main training target losses include: 3D hand joint loss:

[0068]

[0069] 3D hand mesh vertex loss:

[0070]

[0071] 2D hand joint loss:

[0072]

[0073] (2) Auxiliary training target losses include:

[0074] Length loss:

[0075]

[0076] Consistency loss:

[0077]

[0078] Conversion loss:

[0079]

[0080] (3) Comparative training target losses include:

[0081] Cross-pose contrast loss:

[0082]

[0083] Self-pose contrast loss:

[0084]

[0085] in, Represent the true coordinates of the i-th 3D joint point, 3D vertex, and 2D joint point, respectively. is the corresponding predicted coordinate, ε is the edge set of the hand skeleton topology. R, T are the predicted hand rotation matrix and translation vector, R canonical , is the true value, f(·) is the graph feature mapping function learned by DGCLNet, H, H′, H + , H - They represent the characteristics of the original image, perturbation image, positive sample image, and negative sample image respectively. τ is the temperature hyperparameter.

[0086] S28. Iteratively optimize the model until convergence to obtain the final two-handed 3D reconstruction graph neural network model.

[0087] S22 of this embodiment includes:

[0088] S221. Use the pre-trained ResNet-50 as the backbone network, input the RGB images of both hands into the backbone network, and extract high-dimensional image features through convolution and pooling operations;

[0089] S222, the high-dimensional features output by the backbone network are respectively input into two independent multi-layer perceptrons MLP corresponding to the left hand and the right hand for nonlinear feature transformation;

[0090] S223, Multilayer Perceptron MLP consists of 3 fully connected layers. ReLU activation function and Dropout regularization are used after each fully connected layer to convert the image features extracted by the backbone network into compact feature representation. Let F be the image features extracted by the backbone network, and MLP can be expressed as:

[0091] H=σ(W3·σ(W2·σ(W1·F+b1)+b2)+b3) (9)

[0092] Among them, W i and b i are the weight matrix and bias vector of the i-th fully connected layer respectively, and σ is the ReLU activation function.

[0093] S224, the output dimension of the last fully connected layer is set to twice the number of target key points, and the left and right hand MLPs output the two-dimensional coordinates of the key points corresponding to the left and right hands respectively:

[0094]

[0095] in, and are the parameters of the last fully connected layer of the left and right hands, respectively, K left ∈R 2n and K right ∈R 2n are the two-dimensional coordinates of N key points of the left and right hands.

[0096] S23 of this embodiment includes:

[0097] S231, according to the skeleton topology structure of the key points of the hand, define an undirected graph G=(V, E), where V is a set of key points and E is a set of edges between the key points;

[0098] S232, use the two-dimensional coordinates of the key points as the initial features of the nodes, and construct the adjacency matrix A. If there is an edge connecting two key points, then A ij =1, otherwise A ij =0;

[0099] S233. Calculate the degree matrix D of the graph. D is a diagonal matrix with diagonal elements D ii is equal to the degree of node i, that is, the number of edges connected to node i;

[0100] S234, calculating the Laplacian matrix L=DA of the graph, where the Laplacian matrix is ​​a symmetric matrix that contains the topological structure information of the graph;

[0101] S235. Use the Laplacian matrix L to perform graph convolution on the node feature X output by the MLP to obtain a new node feature H:

[0102]

[0103] in, To add a self-loop adjacency matrix, for , W is the weight matrix of graph convolution, and H captures the node features and graph structure information.

[0104] S24 of this embodiment includes:

[0105] S241, input the graph feature H obtained in S23 into the perturbation graph contrastive learning network DGCLNet, which is composed of a graph U-Net and a graph convolutional network GCN;

[0106] S242, Graph U-Net adopts an encoder-decoder structure. The encoder consists of multiple graph convolutional layers and pooling layers, which downsamples the graph features into multi-scale graph representations; the decoder restores the graph feature resolution through upsampling and skip connections:

[0107]

[0108] Among them, H (l) is the output feature of the l-th layer decoder, is the operation of the l-th layer decoder, including upsampling and skip connections, H (l-1) It is the output of the Llth layer of the encoder, which is concatenated with the decoder features through a skip connection.

[0109] S243, Graph U-Net outputs a graph representation H′ with the same feature scale as the original graph, and (H, H′) is used as a positive sample pair; at the same time, the original graph H is perturbed by data enhancement methods such as node replacement to obtain a negative sample H-, and (H, H-) is used as a negative sample pair;

[0110] S244, the positive and negative sample pairs (H, H′) and (H, H-) are respectively input into two parameter-sharing GCNs. GCN aggregates the neighborhood information of the nodes in each graph convolution layer and extracts high-level semantic features:

[0111]

[0112] Among them, H (l) is the output feature of the l-th layer decoder, W (l) is the weight matrix of the lth layer of graph convolution, H (l-1) is the output of the Llth layer of the encoder, which is concatenated with the decoder features through a skip connection. To add a self-loop adjacency matrix, for The degree matrix of .

[0113] The node features output by the last layer of S245 and GCN are globally pooled to obtain graph-level representations f(H), f(H′), f(H - ), through the cosine similarity s(f(H), f(H′)) and s(f(H), f(H -))Calculate the similarity between positive and negative sample pairs:

[0114]

[0115] Among them, τ is the temperature hyperparameter, k is the number of negative samples, and by minimizing the contrast loss, DGCLNet can learn a robust graph representation.

[0116] S246. Maximize the similarity of positive sample pairs s(f(H), f(H′)), minimize the similarity of negative sample pairs s(f(H), f(H-)), and optimize the parameters of graph U-Net and GCN by comparing the learning objective function, so that DGCLNet can learn a robust graph representation.

[0117] S25 of this embodiment includes:

[0118] S251, the graph feature H learned by the original branch of DGCLNet mano Input into the multi-layer graph convolution upsampling network, and gradually restore the resolution of the graph features through the transposed convolution operation;

[0119] S252, the graph convolution upsampling network outputs high-resolution graph features that are consistent with the number of vertices of the MANO hand model template mesh, and propagates the features to each vertex of the MANO mesh through graph convolution:

[0120]

[0121] in, The operation of the lth upsampling layer includes graph transposed convolution and graph convolution.

[0122] S253,MANO template mesh contains 778 vertices for the left and right hands, and the three-dimensional coordinates of the vertices define the three-dimensional shape of the hand surface;

[0123] S254, vertex features are input into the spatial transformation network to predict the three-dimensional offset of each vertex relative to the hand root joint and the three-dimensional rotation angle of each joint:

[0124] ΔV=f V (H mano ), θ=f θ (H mano ) (16)

[0125] Among them, f V and f θ They are the prediction networks for vertex offset and joint rotation angle respectively.

[0126] S255. According to the predicted vertex offset and joint rotation angle, the MANO template mesh is deformed by the linear blend skinning LBS algorithm to obtain the three-dimensional surface shape and joint posture of the hand:

[0127] V=LBS(T,θ,ΔV) (17)

[0128] Where T is the template mesh of the MANO hand model, LBS is the linear mixed skinning function, θ is the parameter that controls the deformation of the hand model, and ΔV is the three-dimensional rotation angle of each joint.

[0129] S256, finally output the three-dimensional coordinates of 778 vertices on the left and right hands (V left , V right ), reconstructs a 3D hand model corresponding to the input image, and the parameterized representation of MANO makes it possible to control the hand model with low-dimensional posture and shape parameters to generate realistic 3D hand shapes.

[0130] The creation of virtual pen recognition of subtle gesture changes in S3 of this embodiment includes:

[0131] S31, creating a virtual pen in a virtual reality environment, and grabbing the virtual pen by presetting a grabbing gesture on the pen;

[0132] S32, when the hand interacts with the pen, calculating the error between each finger joint of the virtual hand and the preset gesture, and calculating the average position error of each joint;

[0133] S33, by setting an error threshold for each finger, judging whether each finger reaches a preset gesture position, dividing into four different pen-holding finger gestures, and introducing additional constraints to reduce misjudgment caused by occlusion;

[0134] S34, dividing the pen-holding finger gestures into four types corresponding to four functions respectively, including pinching the thumb and index finger together, the thumb, index finger and middle finger reaching the preset gesture finger positions at the same time, the thumb, index finger and ring finger bending at the same time, and all fingers reaching the preset positions at the same time.

[0135] S31 of this embodiment includes:

[0136] S311. Simply construct a three-dimensional virtual pen model in a virtual reality scene;

[0137] S312, recording the hand posture during grasping, and combining it with the virtual pen into a new three-dimensional object;

[0138] S313, pinching the index finger and thumb together is a necessary condition for grabbing the virtual pen. When the hand approaches the virtual pen and makes a pinching action, the recorded gesture will drive the virtual pen to be automatically adsorbed to the user's virtual hand.

[0139] S32 of this embodiment includes:

[0140] S321, when the user grabs the virtual pen, except for the index finger and thumb which are at the preset gesture positions, the remaining fingers can be bent / straightened at will;

[0141] S322, calculating the distance error between the middle finger, ring finger and little finger of the virtual hand and the knuckles of the preset gesture, the calculation formula is:

[0142] E ij =||p i -p j ||2, i, j∈{1, 2,...,N}, i≠j(18)

[0143] Assume the set of hand key points is P = {p i |i=1,2,...,N},where p i Represents the 3D coordinates of the i-th key point.

[0144] S33 of this embodiment includes:

[0145] S331, when the average joint error between the virtual hand fingers and the actual hand fingers is less than a certain error threshold (for example: 0.05cm), that is:

[0146]

[0147] If the finger reaches the preset gesture position, it is considered that the finger has reached the preset gesture position, otherwise it has not reached the preset position. By setting different threshold combination conditions, a variety of grasping micro-gestures can be defined, such as:

[0148] Thumb and index finger pinch:

[0149]

[0150] The thumb, index finger, and middle finger reach the preset position at the same time:

[0151]

[0152] The thumb, index finger, and ring finger reach the preset position at the same time:

[0153]

[0154] All fingers reach the preset posture:

[0155]

[0156] S332 of this embodiment includes:

[0157] First, the index finger and thumb pinch is defined as the hub state i (Hub), which serves as the default state of the entire interaction and the entrance to each subtask. In this state, the user can move the virtual pen naturally without triggering any special operations. When the system detects that the user switches to other gesture states, it enters the corresponding subtask mode.

[0158] When identifying each spoke, we introduced additional constraints to reduce misjudgment caused by occlusion:

[0159] C1 state iv: requires five fingers to reach or exceed the preset gesture position (can be a fist posture) at the same time, with the highest priority. This full-clenched gesture is easier to recognize and has a low misjudgment rate.

[0160] C2 state iii: requires the little finger to be in an extended state. In addition to the three required fingers being bent, the middle finger will contract together and maintain the original state, which can effectively distinguish iii from iv.

[0161] C3 state ii: The ring finger and little finger must be extended, while the middle finger is bent. This state only requires judging the bending state of one finger, and the recognition accuracy is relatively high.

[0162] C4 delay protection: A certain delay is set for switching between various states, which can effectively prevent the situation where the user goes from a radial state to a central state and then immediately returns to a radial state, giving the user enough time to determine the next interaction to be performed.

[0163] Through the above method, four basic pen-holding finger posture interaction paradigms are divided, which can be applied to most hand-object interaction systems.

[0164] S4 of this embodiment includes:

[0165] S41, the pinching position of the thumb and index finger corresponds to the function of moving the virtual pen. The virtual pen is selected by pinching the index finger and thumb, and the selection is cancelled by releasing the index finger and thumb, which is used to control the position of the virtual pen in the three-dimensional space;

[0166] S42, the thumb, index finger and middle finger simultaneously reach the preset gesture position corresponding to the call-out menu function;

[0167] S43, the position where the thumb, index finger and ring finger are bent at the same time corresponds to the function of undoing the last operation;

[0168] S44. When all fingers reach the preset gesture position at the same time, a corresponding drawing operation is performed.

[0169] S42 of this embodiment includes:

[0170] S421, a visual menu panel that can be clicked by the user's index finger;

[0171] S422, a color adjustment button, after clicking, expands a hexadecimal color palette, and the user can obtain the desired color by clicking the color palette;

[0172] S423. Copy, modify handwriting style, thickness, delete, combine and free shape operation buttons.

[0173] The present invention provides a handwriting interaction system in a virtual reality scene based on vision, the system comprising:

[0174] The hand image acquisition module is used to collect RGB image sequences of the user's hands in real time through the RGB camera on the head display device as input for subsequent gesture recognition and 3D reconstruction;

[0175] The gesture recognition module is used to input the acquired RGB images of both hands into the perturbation graph contrast learning network (DGCLNet), extract the graph feature representation of the gesture through the graph U-Net, and optimize the graph feature through the cosine similarity loss function to achieve accurate recognition of the gesture;

[0176] The gesture recognition module includes a 3D reconstruction module, which uses the MANO branch in DGCLNet to map the gesture graph features to the vertex coordinates and joint rotation angles of the MANO hand model through a graph convolution upsampling network and a spatial transformation network, and reconstruct the 3D surface shape and joint posture of the hand;

[0177] The hand-pen interaction mapping module is used to map the recognized gestures with the interactive operations of the virtual brush. Through gestures such as pinching and multi-finger combination, the virtual pen can be used to grab, call out menus, and cancel, thus achieving natural interaction between bare hands and the virtual pen.

[0178] The hand-pen interaction mapping module includes a painting control module, which is used to control the virtual brush to paint in the three-dimensional virtual scene according to the result of hand-pen interaction, including controlling the movement of the brush, drawing lines, controlling the size and posture of the drawn object, adjusting the color, canvas interaction, etc., to realize bare-hand painting control;

[0179] The presentation module is used to present the three-dimensional model of the hand, the status of the virtual pen, and the drawing results in real time in the virtual scene, provide visual feedback, and give audio prompts based on the interaction results to enhance the user's immersion and operation feedback.

[0180] Each module interacts with each other through data and control signals, and works together to realize a visual-based virtual reality bare-hand interactive painting system. Users can wear a VR headset, use their hands to interact naturally with the virtual pen, and use gestures to control the virtual pen to create paintings in a three-dimensional virtual space.

[0181] The hand three-dimensional reconstruction module is connected to the hand-pen interaction mapping module, and the hand-pen interaction mapping module is connected to the interaction presentation module.

[0182] like Figure 1 As shown, the embodiment of the present invention provides a hand-pen interaction method in a virtual reality scene based on vision. By creating two hands in real time, the average finger joint error between the preset grasping gesture and the created gesture is calculated, four different pen-holding finger postures are distinguished, and a hand-pen interaction paradigm is formed. The four pen-holding finger postures are respectively mapped to four basic painting system functions in a virtual reality three-dimensional scene and integrated into the hand-pen interaction system. The specific network details and hand-pen interaction paradigm are shown in FIG. Figure 2 , Figure 3 .

[0183] like Figure 2 The network structure diagram of virtual hand reconstruction is shown. Given an RGB image, our model uses ResNet-50 for feature extraction, and then divides the features into left and right hands through position embedding. The features of both hands are further processed by two independent MLP layers, and the knowledge of the Laplacian matrix is ​​used to construct graph features. Next, our DGCLNet regenerates the graph features through the graph u-net for perturbation. The graph feature pairs are passed through the same GCN and the similarity between them is maximized. Finally, the graph features of the original branch are fed into the upsampling layer to obtain the positions of 778 vertices in a standard MANO manner.

[0184] like Figure 3 The design concept for the pen-based interaction paradigm is to combine common operations of drawing in 3D space (moving, calling up menus, undoing, drawing) with the pen-based interaction paradigm. Using the virtual pen as a proxy for visually identifying changes in finger posture, users can create intuitively and freely in three-dimensional space without additional hardware input.

[0185] like Figure 4 Framework diagram of the hand-pen interaction paradigm: (1st from left) The gesture column represents the basic gestures of our interaction system; (2nd from left) Finger gestures are used to control the drawing process of the virtual pen; (2nd from right) Actions represent some interactive actions that users can complete through these basic gestures; (1st from right) Referring to the research of Li et al., we set up several types of pen holding for interaction with the pen: TFE, TRE, QRE, Pinch and Overhand, to meet the pen holding needs of different users.

[0186] like Figure 5 As shown in the figure, in order to prevent the pen tip from penetrating the canvas, we add ray detection at the pen tip. When the pen approaches the canvas, it triggers the detection of whether the pen tip is in front of or behind the canvas, and calculates the distance and difference vector from the pen tip to the plane.

[0187] First, we calculate the projection point of the pen tip on the canvas plane using the vector Indicates the direction vector from the pen tip to the pen handle It can be expressed as:

[0188]

[0189] in, is the position vector of any point on the plane, is the normal vector of the drawing board plane, P pen is the position coordinate of the middle of the pen. The signed distance d from the pen tip coordinate P to the canvas plane can be calculated using the point-to-plane distance formula:

[0190]

[0191] If the pen tip is on the back of the canvas and the distance is greater than the threshold, the pen is corrected to the front of the canvas and kept at a certain distance from the canvas plane to simulate the feeling of a pen hitting paper in real life. At this time, the pen position P′ pen Set to:

[0192]

[0193] Among them, ∈ is a small positive number, which is used to keep a certain distance between the pen tip and the canvas to avoid overlap. This can prevent the pen from penetrating the canvas and improve the realism of the interaction.

[0194] like Figure 6 : As shown in the figure, the visual feedback of the bare-hand pen interaction system in a vision-based virtual reality scene: a) the virtual pen will not penetrate the blackboard. When the real hand goes deep into the blackboard, the virtual pen can sense the user's strength and increase the width of the handwriting; b) when adjusting the color, the virtual pen will change with the color set on the palette; c)-e) are the color change process when the user grabs an object. The purpose is to show the user the currently hovering / selected object, c) is the original color of the object, c) is the original color of the object, d) is the hovering color, and e) is the color of the selected object.

[0195] like Figure 7 The figure shows the auditory feedback of the bare-hand pen interaction system in a vision-based virtual reality scene: when switching states, the virtual pen will emit different tones to prompt the user to switch states. (a) When calling out the menu, b) when releasing the virtual pen, c) when starting to draw.

[0196] like Figure 8 Shown is a collection of freely created works using the bare-hand pen interaction system in a vision-based virtual reality scene.

[0197] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A handwriting interaction system in a vision-based virtual reality scene, characterized in that: The system includes: a hand three-dimensional reconstruction module, a hand-pen interactive mapping module and an interactive presentation module; The hand three-dimensional reconstruction module is connected to the hand-stroke interaction mapping module, and the hand-stroke interaction mapping module is connected to the interaction presentation module.

2. The handwriting interaction system in a virtual reality scene based on vision according to claim 1 is characterized in that: The hand 3D reconstruction module is used to collect RGB image sequences of the user's hands in real time through the RGB camera on the head display device, input them into the perturbation graph contrast learning network DGCLNet, extract the graph feature representation of the gesture through the graph U-Net, and optimize the graph features through the cosine similarity loss function to achieve accurate recognition of the gesture, and then map the gesture graph features to the vertex coordinates and joint rotation angles of the MANO hand model through the graph convolution upsampling network and the spatial transformation network, so as to reconstruct the 3D surface shape and joint posture of the hand.

3. The handwriting interaction system in a virtual reality scene based on vision according to claim 1, characterized in that: The hand-pen interaction mapping module is used to map the recognized gestures with the interactive operations of the virtual brush, and realize the virtual pen grabbing, menu calling and undo operations through pinching and multi-finger combination gestures, thereby realizing natural interaction between bare hands and the virtual pen.

4. The handwriting interaction system in a virtual reality scene based on vision according to claim 1, characterized in that: The interactive presentation module is used to control the virtual brush to paint in the three-dimensional virtual scene according to the result of the hand-pen interaction, and at the same time present the three-dimensional model of the hand, the state of the virtual pen and the painting result in real time in the virtual scene to provide visual feedback, and give audio prompts according to the interaction results to enhance the user's immersion and operation feedback.

5. A handwriting interaction method in a vision-based virtual reality scene, characterized in that: The method comprises: S1, obtain the hand image from the user's perspective through the visual sensor; S2. Input the hand image into the constructed graph neural network to perform three-dimensional reconstruction of both hands to obtain a model, wherein the model includes a backbone network, a neck network, and a prediction head; S3, creating a virtual pen in a virtual reality scene, and classifying four pen-holding gestures by identifying the slight changes of each finger when holding the pen; S4. Based on the four pen-holding gestures, the basic functions of the system in the virtual reality scene are mapped to form a hand-pen interaction paradigm.

6. The handwriting interaction method in a virtual reality scene based on vision according to claim 1, characterized in that: The S2 comprises the following sub-steps: Sub-step S21, obtaining images of the user's hands; Sub-step S22, inputting the images of both hands into the backbone network to extract image features, and inputting the image features into two independent multi-layer perceptrons MLP for left and right hands respectively for feature conversion; Sub-step S23, using the Laplacian matrix to construct the features output by the multi-layer perceptron into graph features; Sub-step S24: input the graph features into the perturbation graph contrast learning network DGCLNet, regenerate the graph features through the graph U-Net to form feature pairs, Sub-step S25, the feature pairs are passed through a shared graph convolutional network GCN to maximize the similarity of the two output features; Sub-step S26, inputting the graph features of the original branch into the upsampling layer, and obtaining the three-dimensional coordinates of 778 vertices of both hands in a MANO manner; Sub-step S27, introducing a similarity distinction method to accurately calculate the similarity of feature pairs from the feature-pose pool to accelerate model convergence; Sub-step S28, designing multiple loss functions to balance the relationship between hand posture, visual scale and viewpoint position to improve the quality of three-dimensional reconstruction; Sub-step S29, iteratively optimize the model until convergence to obtain the final two-handed 3D reconstruction graph neural network model.

7. The method for handwriting interaction in a virtual reality scene based on vision according to claim 1, characterized in that: The S3 comprises the following sub-steps: Sub-step S31, creating a virtual body of a pen in a virtual reality environment, and grabbing the virtual pen by presetting a grabbing gesture on the virtual pen; Sub-step S32, when the hand interacts with the virtual pen, measuring the error between the virtual hand and each finger joint of the preset gesture, and calculating the average position error of each joint; Sub-step S33, by setting the error threshold of each finger joint position, determining whether the finger reaches the preset gesture position, and accurately identifying each finger gesture; Sub-step S34, setting 4 finger gestures corresponding to four functions respectively, the 4 finger gestures include the pinching position of the thumb and index finger, the thumb, index finger and middle finger reaching the preset gesture finger position at the same time, the thumb, index finger and ring finger bending position at the same time, and all fingers reaching the preset position at the same time.

8. The handwriting interaction method in a virtual reality scene based on vision according to claim 1, characterized in that: The S4 comprises the following sub-steps: Sub-step S41, the position of the thumb and index finger pinching corresponds to the function of moving the virtual pen, the virtual pen is selected by pinching the index finger and thumb, and the selection is cancelled by releasing, which is used to control the position of the virtual pen in the three-dimensional space; Sub-step S42, the thumb, index finger and middle finger simultaneously reach the preset gesture finger position corresponding to calling out the menu function; Sub-step S43, the thumb, index finger and ring finger are bent at the same time to the corresponding position to cancel the previous operation function; Sub-step S44: When all fingers reach the preset position at the same time, it corresponds to the drawing operation.

Citation Information

Patent Citations

  • Gesture recognition method and device, computer equipment and storage medium

    CN112904994A

  • Drawing method and device, computer equipment and storage medium

    CN113703577A

  • Gesture recognition drawing method and device, equipment and storage medium

    CN117237526A

  • Hand-tracked text selection and modification

    CN119213468A

  • Method and apparatus for virtual reality animation

    US20180165877A1

Cited By

  • Intelligent animation modeling method based on three-dimensional technology

    CN121053266A

  • An intelligent animation modeling method based on three-dimensional stereoscopic technology

    CN121053266B