Bare-hand-based interaction method and device in VR environment, electronic equipment and medium
By identifying and matching users' gestures in the VR environment, flexible switching of operation modes and grouping of virtual elements are achieved, and the lack of naturalness and flexibility of interaction methods in the prior art is solved, and interaction efficiency and user experience are improved.
Patent Information
- Application Number
- CN202510174922.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The interaction method in the existing VR environment limits the user's natural movements, making it difficult to adapt to complex gesture operations, and the interaction mode is single and cannot be flexibly switched.
By obtaining the user's gesture image information for the target virtual scene, identifying the user's gesture type, and matching with the preset set of control gesture types, the operation mode switching is achieved. In the target operation mode, the user's control gesture type is identified and the virtual elements are grouped.
It realizes natural, efficient and flexible virtual scene interaction, improves user experience and interaction efficiency, adapts to the diverse interaction needs in different virtual scenes, and enhances the universality and robustness of the system.
Smart Images

Figure CN120103975A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of virtual reality, and to methods, devices, electronic devices, and media for bare-hand interaction in a VR environment. Background Art
[0002] With the rapid development of virtual reality (VR) technology, the way users interact with virtual scenes is gradually changing from traditional controllers to more natural and intuitive bare-hand interactions. At present, the main interaction methods in VR environments are: interacting in VR environments through external input devices such as handles, gloves or dedicated interaction tools to process objects.
[0003] However, when the above method is used to process objects in a VR environment, the following technical problems often occur: the above method often restricts the user's natural movements, is difficult to adapt to complex gesture operations, and the interaction mode is single and cannot be flexibly switched.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention
[0005] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.
[0006] Some embodiments of the present disclosure propose bare-hand-based interaction methods, devices, electronic devices, and media in a VR environment to explain one or more of the technical issues mentioned in the above background technology section.
[0007] In the first aspect, some embodiments of the present disclosure propose a bare-hand-based interaction method in a VR environment, the method comprising: obtaining gesture image information of a user corresponding to a target virtual scene; identifying the user gesture type based on the gesture image information; matching the user gesture type with a preset set of manipulation gesture types to obtain a matching manipulation gesture type; in response to determining that the matching manipulation gesture type satisfies a mode switching condition, switching the operation mode according to the matching manipulation gesture type; in the target operation mode, identifying the user's manipulation gesture type; and performing group operation processing on at least one virtual element in the target virtual scene according to the user's manipulation gesture type.
[0008] In the second aspect, some embodiments of the present disclosure propose a bare-hand-based interaction device in a VR environment, the device comprising: an acquisition unit, configured to acquire gesture image information of a user corresponding to a target virtual scene; a first recognition unit, configured to identify the user gesture type based on the gesture image information; a matching unit, configured to match the user gesture type with a preset set of manipulation gesture types to obtain a matching manipulation gesture type; a switching unit, configured to switch the operation mode according to the matching manipulation gesture type in response to determining that the matching manipulation gesture type satisfies a mode switching condition; a second recognition unit, configured to identify the user's manipulation gesture type in the target operation mode; an operation unit, configured to perform group operation processing on at least one virtual element in the target virtual scene according to the user's manipulation gesture type.
[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the above-mentioned first aspect.
[0010] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the above-mentioned first aspect is implemented.
[0011] In a fifth aspect, some embodiments of the present disclosure provide a computer program product, including a computer program, which implements the method described in any implementation manner of the above-mentioned first aspect when executed by a processor.
[0012] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the bare-hand interaction method in the VR environment of some embodiments of the present disclosure, a natural, efficient and flexible virtual scene interaction can be achieved, and the user experience and interaction efficiency are improved. Specifically, the traditional VR interaction method often relies on external devices (such as handles or gloves), which restricts the natural movements of users and is difficult to adapt to complex gesture operations; at the same time, the interaction mode is single and cannot be flexibly switched, resulting in low operation efficiency. Based on this, the bare-hand interaction method in the VR environment of some embodiments of the present disclosure first obtains the gesture image information of the user corresponding to the target virtual scene. Thus, a real-time visual data basis is provided for subsequent recognition. Then, according to the above-mentioned gesture image information, the user gesture type is identified. Thus, the user gesture type can be identified according to the user's gesture through a customized model. Then, according to the above-mentioned user gesture type, it is matched with a preset control gesture type set to obtain a matching control gesture type. Thus, the identified user gesture type can be identified as a control gesture type. Secondly, in response to determining that the above-mentioned matching control gesture type meets the mode switching condition, the operation mode is switched according to the above-mentioned matching control gesture type. Thus, the operation mode is switched to the operation mode selected by the user. Then, in the target operation mode, the user's manipulation gesture type is identified. Thus, the user's gesture for group operation processing in the selected operation mode can be identified. Finally, according to the manipulation gesture type of the above-mentioned user, at least one virtual element in the above-mentioned target virtual scene is subjected to group operation processing. Thus, objects in the VR environment can be subjected to group operation processing, so that users can complete complex interactive tasks through natural gestures without relying on additional equipment, which significantly improves the naturalness and flexibility of the interaction. At the same time, multi-mode switching improves the interaction efficiency and reduces operation delays. In addition, since the method can accurately identify gestures and dynamically adjust according to the user's operation mode, it can adapt to the diverse interaction needs in different virtual scenes, and enhance the versatility and robustness of the system. Thus, by dynamically matching user gestures with preset operation modes and virtual elements, efficient and natural virtual scene interaction is achieved, which improves the overall user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0014] Figure 1 is a flowchart of some embodiments of the bare-hand-based interaction method in a VR environment according to the present disclosure;
[0015] Figures 2 to 6It is an application scenario diagram according to some embodiments of the bare-hand-based interaction method in a VR environment;
[0016] Figure 7 is a schematic structural diagram of some embodiments of a bare-hand-based interaction device in a VR environment according to the present disclosure;
[0017] Figure 8 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0019] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0020] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0021] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0024] Figure 1 The process 100 of some embodiments of the bare-hand-based interaction method in a VR environment according to the present disclosure is shown. The bare-hand-based interaction method in the VR environment includes the following steps:
[0025] Step 101, obtaining gesture image information of a user corresponding to a target virtual scene.
[0026] In some embodiments, the execution subject (e.g., head-mounted display device) of the bare-hand-based interaction method in a VR environment can obtain the gesture image information of the user in the target virtual scene through the equipped sensors. Among them, the above-mentioned execution subject can be VR glasses or VR helmets. The above-mentioned equipped sensors include cameras and depth sensors to provide two-dimensional and three-dimensional gesture image data and perform gesture recognition in different environments. The above-mentioned target virtual scene can be any virtual scene in which the user performs group operations. The above-mentioned target virtual scene usually includes at least one virtual element. The above-mentioned virtual elements can be the targets for users to interact and operate in the virtual scene, including objects in the three-dimensional model (such as virtual furniture, tools, building components, etc.). For example, in a virtual architectural design scene, the virtual elements can be virtual walls, doors and windows, furniture, etc. The above-mentioned gesture image information is usually image data containing gesture actions.
[0027] Step 102: Identify the user gesture type based on the gesture image information.
[0028] In some embodiments, the above-mentioned execution subject can identify the user gesture type based on the above-mentioned gesture image information. Among them, the above-mentioned gesture types include static gestures in which the hand maintains a fixed posture (such as open palm, fist, etc.) and dynamic gestures involving hand movements (such as finger movement trajectory, palm waving, etc.). The above-mentioned execution subject can use machine learning algorithms such as convolutional neural network (CNN) or support vector machine (SVM) to identify the gesture type. In practice, the above-mentioned execution subject can extract features from the gesture image information to obtain feature data of the gesture image. Then, the user gesture type can be identified by inputting the above-mentioned gesture image feature data into the above-mentioned machine learning algorithm.
[0029] In some optional implementations of some embodiments, the execution subject may identify the user gesture type through the following steps:
[0030] In the first step, the gesture image information is input into the input layer of the pre-trained gesture recognition model to obtain the preliminarily processed image information. The gesture recognition model may be a neural network model that takes the gesture image as input data and the recognized user gesture type as output. For example, the neural network model may be a CNN. The gesture recognition model may include the input layer, the initial convolution layer, the maximum pooling layer, a preset number of residual block groups, a global average pooling layer, a flattening layer, a fully connected layer and an output layer. In practice, the execution subject may receive the pixel value of the gesture image as input data through the input layer.
[0031] In the second step, the above-mentioned preliminarily processed image information is input into the above-mentioned initial convolution layer to obtain a low-level feature map. Among them, the above-mentioned low-level feature map includes low-level features such as hand shape features, hand texture features and gradient directions. The above-mentioned low-level feature map usually has a high spatial resolution, but relatively less feature semantic information. The convolution kernel size of the above-mentioned initial convolution layer is usually 3x3 or 5x5. The number of convolution kernels of the above-mentioned initial convolution layer is usually 32 or 64. The step size of the above-mentioned initial convolution layer is usually 1. The padding of the above-mentioned initial convolution layer is usually the same (keeping the input and output sizes the same). The activation function of the above-mentioned initial convolution layer is usually ReLU. In practice, the above-mentioned execution entity can perform preliminary feature extraction on the input image through the above-mentioned initial convolution layer to obtain the above-mentioned low-level feature map.
[0032] The third step is to input the low-level feature map into the maximum pooling layer to obtain a pooled feature map. The pooling window size of the maximum pooling layer is usually 2x2. The step size of the maximum pooling layer is usually the same as the pooling window size. In practice, the execution subject can reduce the spatial dimension of the feature map and the amount of calculation while maintaining important features through the maximum pooling layer.
[0033] In the fourth step, the above-mentioned pooled feature map is sequentially passed through the above-mentioned preset number of residual block groups for feature extraction to obtain an enhanced feature map. Each residual block may include 3 convolutional layers, and each convolutional layer may be followed by a batch normalization layer, an activation function, and a jump connection. The convolution kernel size in the convolution layer of each residual block is usually 3x3. The number of convolution kernels in the convolution layer of each residual block can be gradually increased, for example, from 32 to 64, and then to 129. The step size of the convolution layer of each residual block is usually 1, but may be 2 when the step size is downsampled. The padding of the convolution layer of each residual block is usually the same. The activation function of the convolution layer of each residual block usually uses ReLU. In practice, the above-mentioned execution subject can introduce an attention mechanism in the above-mentioned residual block to enhance the focus on important feature channels. The above-mentioned execution subject can introduce a feature pyramid between the above-mentioned preset number of residual block groups to enhance the recognition ability of gestures of different scales. The execution subject can solve the gradient vanishing problem in deep network training by performing residual connection on the preset number of residual block groups, thereby enhancing the feature extraction capability. The preset number of residual block groups can be a more expressive and discriminative feature representation obtained by performing multi-dimensional feature extraction and optimization on the pooled feature map through a series of network structures and mechanisms to obtain an enhanced feature map.
[0034] The fifth step is to input the enhanced feature map into the global average pooling layer to obtain a pooled feature vector. The size of the pooled feature vector is usually (1,1), that is, the enhanced feature map of each channel is compressed into a single value. In practice, the execution subject can calculate the average value of all pixel values for the enhanced feature map input to each channel through the global average pooling layer. Then the enhanced feature map of each channel is converted into a single value by compressing the height and width. This reduces the number of parameters while retaining global information. The length of the pooled feature vector finally generated is equal to the number of channels of the input enhanced feature map, and a pooled feature vector is obtained.
[0035] Step 6: Input the pooled feature vector into the flattening layer to obtain an input feature vector. The flattening layer usually converts a multi-dimensional vector into a one-dimensional vector by a flattening operation. In practice, the execution subject can flatten the multi-dimensional feature map into a one-dimensional vector through the flattening layer so as to input it into the fully connected layer.
[0036] In the seventh step, the input feature vector is input into the fully connected layer to obtain an output vector. The number of input features of the fully connected layer is usually the length of the flattened feature vector, such as 129. The number of output features of the fully connected layer is usually the number of categories, such as 6 (for gesture recognition tasks). In practice, the execution subject can perform a linear transformation on the input feature vector through the weight matrix and bias vector in the fully connected layer. The execution subject can integrate local features into global features by introducing nonlinearity through the activation function (usually ReLU) in the fully connected layer. Then, the execution subject can map the input feature vector to the target dimension through the learning weight matrix and bias vector in the fully connected layer. The target dimension refers to the dimension of the feature vector output by the model, which is usually related to the number of target categories of the task. For example, in the gesture recognition task, the target dimension is the number of gesture categories.
[0037] In the eighth step, the output vector is input into the output layer to obtain a probability distribution set to identify the user gesture type. Each probability distribution in the probability distribution set corresponds to a gesture type. The length of the probability distribution set is the number of preset control gesture types. The output layer usually uses Softmax for multi-classification tasks. In practice, the execution entity can map the output vector to the probability distribution set through the output layer. Then, the execution entity can use the output layer to use the gesture type corresponding to the probability distribution that meets the preset conditions (usually the highest probability value) in the probability distribution set as the recognized user gesture type.
[0038] The first to eighth steps mentioned above are an inventive point of an embodiment of the present disclosure, which solves the technical problem that "the existing gesture recognition technology has low recognition accuracy, high model complexity, and insufficient recognition ability of gestures of different scales, which makes it difficult to adapt to complex gesture operations". The reasons why the existing technology is difficult to adapt to complex gesture operations are as follows: when dealing with complex backgrounds, gestures of different scales, and the gradient vanishing problem in deep network training, the existing gesture recognition model is difficult to effectively extract key features, thereby affecting the recognition accuracy and model efficiency. If the above factors are solved, the effect of improving gesture recognition accuracy, reducing model complexity, and enhancing the recognition ability of gestures of different scales can be achieved. In order to achieve this effect, the present disclosure adopts a deep convolutional neural network (CNN) model, which receives image data through the input layer, extracts low-level features through the initial convolution layer, reduces the dimension through the maximum pooling layer, enhances features through the residual block group, retains global information through the global average pooling layer, converts the data format through the flattening layer, and integrates features through the fully connected layer. Finally, the output layer performs classification decisions to achieve efficient gesture recognition. In this way, the key features of the gesture image can be effectively extracted, the accuracy and efficiency of gesture recognition can be improved, and the model complexity can be reduced at the same time, and the recognition ability of gestures of different scales can be enhanced. This allows the model to adapt to complex gesture operations.
[0039] Step 103 : matching the user gesture type with a preset control gesture type set to obtain a matching control gesture type.
[0040] In some embodiments, the above-mentioned execution subject may match the above-mentioned recognized user gesture type with the manipulation gesture in the above-mentioned preset manipulation gesture type set to obtain the matching manipulation gesture type. Among them, the preset manipulation gesture type in the above-mentioned preset manipulation gesture type set may be a pre-set gesture that can be manipulated. The above-mentioned preset manipulation gesture type set includes a group of defined gesture actions. The above-mentioned preset manipulation gesture type set includes control gestures and switching gestures. In practice, the above-mentioned execution subject may search for the above-mentioned recognized user gesture type in the above-mentioned preset manipulation gesture type set to obtain the above-mentioned matching manipulation gesture type.
[0041] The gesture guide map of the above-mentioned preset control gesture type can be preset as follows: Figure 2 .in, Figure 2Including control gestures and switching gestures. Among them, the above-mentioned control gestures include: turning on or off gestures, confirmation gestures and cancel gestures. The above-mentioned confirmation gestures are usually used to confirm the user's intention to interact with virtual elements or to complete the switching of a certain operation mode. It plays a key triggering role in the interaction process, ensuring that the user's gesture operation is accurately recognized by the above-mentioned execution subject and executes the corresponding instructions. The above-mentioned switching gestures include: switching gestures for selection and deletion modes, switching gestures within selection and deletion modes, and switching gestures for finger and palm modes. The above-mentioned switching gestures within the selection and deletion mode include switching gestures within the selection mode and switching gestures within the deletion mode. The above-mentioned switching gestures include sub-mode switching gestures. The above-mentioned sub-mode switching gestures include: switching gestures for finger and palm modes and switching gestures within the deletion mode.
[0042] Step 104 : in response to determining that the matching manipulation gesture type satisfies a mode switching condition, switching the operation mode is performed according to the matching manipulation gesture type.
[0043] In some embodiments, the execution subject may switch the operation mode according to the matching control gesture type in response to determining that the matching control gesture type satisfies the mode switching condition. Among them, the mode switching condition is usually that the execution subject recognizes that the matching control gesture is the switching gesture. The operation mode may include a selection mode and a deletion mode. The switchable operation mode may include the selection mode and the deletion mode. In the selection mode, at least one virtual element in the virtual scene may be grouped. The selection mode includes a line selection mode and a block selection mode. In the line selection mode, the virtual elements in the virtual scene may be selected as the same group by connecting them by lines. In the block selection mode, the virtual elements in the virtual scene may be selected as the same group by surrounding them by blocks. In the deletion mode, the grouped virtual elements in the virtual scene may be changed to an ungrouped state. The initial unswitched operation mode is usually the line selection mode. In practice, the execution subject may switch the operation mode when recognizing that the matching control gesture type is a switching gesture in a preset control gesture type set, and after recognizing the confirmation gesture, complete the switching of the operation mode and use the current operation mode as the target operation mode. The target operation mode may be an operation mode determined by the execution subject after the execution subject recognizes the confirmation gesture by switching the operation mode. For example, when the execution subject recognizes that the user gesture type is a switching gesture between selection and deletion modes, the operation mode is selection mode at this time. After recognizing the confirmation gesture, the selection mode is switched to deletion mode. When the execution subject recognizes that the user gesture type is a switching gesture within the selection mode, the operation mode is line selection mode at this time. After recognizing the confirmation gesture, the line selection mode is switched to block selection mode.
[0044] Step 105: In the target operation mode, identify the user's manipulation gesture type.
[0045] In some embodiments, the execution subject can identify the type of user's manipulation gesture in the target operation mode. The type of user's manipulation gesture is usually a pre-set gesture for manipulation in the operation mode. The type of user's manipulation gesture includes: a gesture for controlling the movement of the selector and a sub-mode switching gesture. The selector can be a virtual tool or virtual object controlled by the user's gesture in the operation mode, and the selector is used to select, operate or process virtual elements in the virtual scene. The form and function of the selector will vary according to the operation mode. For example, in different selection modes or deletion modes, the selector can be the user's fingertip or palm; in the selection mode, the selector can also be used for grouping virtual elements or connecting and splitting virtual element groups. The gesture for controlling the movement of the selector is usually a gesture of sweeping across a path or area when the user uses the selector for group operation processing. The group operation processing is usually a method for the user to interact with the virtual elements in the target virtual scene through the selector in the operation mode. The above-mentioned group operation processing includes: the above-mentioned execution subject can change the virtual elements in the ungrouped state into the grouped state through the above-mentioned selection mode, the above-mentioned execution subject can change the virtual elements in the grouped state into the ungrouped state through the above-mentioned deletion mode, and the above-mentioned execution subject can connect and split the virtual element group through the connection and split functions of the above-mentioned selection mode. In practice, the above-mentioned execution subject can obtain the gesture image information of the user corresponding to the above-mentioned target virtual scene in the above-mentioned target operation mode. The user's manipulation gesture type is identified according to the above-mentioned gesture image information. For example, when the above-mentioned target operation mode is the block selection mode, the above-mentioned execution subject obtains the gesture image information of the user in the target virtual scene, and uses the deep learning framework (such as MediaPipe) to identify the sliding of the selector, and obtains that the manipulation gesture type of the above-mentioned user is a gesture for controlling the movement of the selector.
[0046] In some optional implementations of some embodiments, the execution subject may identify the user's manipulation gesture type in the target operation mode through the following steps:
[0047] The first step is to switch the sub-mode in response to the recognition that the manipulation gesture type of the above-mentioned user is a sub-mode switching gesture. Among them, the sub-modes of the above-mentioned line selection mode may include a finger selection mode and a palm selection mode. The above-mentioned finger selection mode may be a line selection mode using the user's fingertip as a selector. The above-mentioned palm selection mode may be a line selection mode using the user's palm as a selector. The sub-modes of the above-mentioned deletion mode include a single deletion mode and a group deletion mode. In the above-mentioned single deletion mode, the virtual elements swept by the selector can be changed from a grouped state to an ungrouped state. In the above-mentioned group deletion mode, the virtual elements in the virtual element group swept by the selector can be changed from a grouped state to an ungrouped state. For example, in the above-mentioned line selection mode, the above-mentioned execution subject recognizes the switching gesture of the above-mentioned finger-palm mode through the above-mentioned gesture recognition model, obtains that the manipulation gesture type of the above-mentioned user is a sub-mode switching gesture, and switches the above-mentioned finger selection mode and the above-mentioned palm selection mode.
[0048] The second step is to use the switched sub-mode as the target operation mode after recognizing the confirmation gesture. For example, in the above-mentioned deletion mode, after the above-mentioned execution subject recognizes the switching gesture in the deletion mode, it switches the above-mentioned single deletion mode and the above-mentioned group deletion mode, and after recognizing the above-mentioned confirmation gesture, the above-mentioned target operation mode is switched to one of the sub-modes of the above-mentioned deletion mode.
[0049] Step 106: performing group operation processing on at least one virtual element in the target virtual scene according to the user's manipulation gesture type.
[0050] In some embodiments, the execution subject may perform group operation processing on at least one virtual element in the target virtual scene according to the manipulation gesture type of the user in the target operation mode. In practice, the execution subject may identify the path or area swept by the selector of the target operation mode in the target operation mode. Then, the at least one virtual element corresponding to the path or area is subjected to group operation processing.
[0051] In some optional implementations of some embodiments, the execution subject may perform group operation processing on at least one virtual element in the target virtual scene according to the manipulation gesture type of the user through the following steps:
[0052] The first step, in response to determining that the above-mentioned target operation mode is the deletion mode, determines the sub-mode of the deletion mode according to the above-mentioned control gesture type, wherein the sub-mode control gesture type set corresponding to the deletion mode includes a single deletion mode control gesture type and a group deletion mode control gesture type. The above-mentioned single deletion mode control gesture type may be a gesture for controlling the movement of the selector recognized in the single deletion mode. The above-mentioned group deletion mode control gesture type may be a gesture for controlling the movement of the selector in the group deletion mode. In practice, the above-mentioned execution subject may determine the sub-mode of the above-mentioned deletion mode according to the above-mentioned user control gesture type.
[0053] The second step is to identify at least one virtual element that the deletion mode selector sweeps over following the user's manipulation gesture. The deletion mode selector may be the user's fingertip. In practice, the execution subject may identify the virtual element that the selector sweeps over following the user's manipulation gesture by following the following steps:
[0054] The first sub-step is to track the motion trajectory of the key points to obtain the user control gesture trajectory. The user control gesture trajectory is usually the movement trajectory of the user in the virtual scene identified by the execution subject, in which the selector is controlled by gesture movement. The user control gesture trajectory can be represented by a curve fitted by the key points. In practice, the execution subject can use a deep learning framework (such as MediaPipe) in combination with a camera or a depth sensor to capture the user control gesture trajectory. MediaPipe can usually detect the key points of the hand (such as fingertips) in real time through a pre-trained neural network model. In practice, the MediaPipe framework is used in combination with a camera to detect the three-dimensional coordinates of the key points of the hand (such as fingertips) in real time. For example, if a user draws a "C"-shaped gesture in the air, MediaPipe will track the movement trajectory of the fingertips to form a continuous curve.
[0055] The second sub-step is to define the position and range of the virtual element or virtual element group in three-dimensional space. In practice, the execution subject can define the range of the virtual element or virtual element group by boundary representation. The boundary representation is usually a representation method based on the boundary of a geometric entity, which represents the range of a geometric entity by defining faces, edges and vertices. For example, in a virtual environment, there is a virtual button, whose position and range can be defined as a three-dimensional rectangular area by boundary representation. The boundary of the area is determined by the coordinates of its minimum and maximum points.
[0056] The third sub-step is to convert the user's manipulation gesture trajectory into three-dimensional coordinates. In practice, the above-mentioned execution subject can calculate the specific position of the hand in the three-dimensional space based on the depth information of the key points of the hand obtained by the depth sensor (such as Kinect or IntelRealSense) and the two-dimensional image coordinates obtained by the camera. Therefore, the above-mentioned execution subject maps the two-dimensional pixel coordinates to three-dimensional space coordinates according to the depth map, RGB image and camera intrinsic parameter matrix provided by the depth sensor, so as to realize the positioning and tracking of the user's manipulation gesture trajectory in the three-dimensional space. For example, the coordinates of the user's finger in the two-dimensional image are (x, y), and the depth information is z. Combined with the camera intrinsic parameter matrix, the coordinates of the finger in the three-dimensional space (X, Y, Z) can be calculated.
[0057] The fourth sub-step is to determine whether the user's manipulation gesture trajectory intersects with a virtual element or a group of virtual elements based on a spatial collision detection algorithm (such as bounding box detection, ray casting, etc.). For example, the trajectory of the user's fingertip in three-dimensional space is a straight line from point A to point B. The above-mentioned execution subject uses a ray casting algorithm to take this straight line as a ray and detect whether it intersects with the bounding box of the virtual element. If the ray intersects with the bounding box, it is considered that the user's manipulation gesture trajectory sweeps over the virtual element.
[0058] The third step is to determine at least one virtual element that has been swept as the target object. In practice, in the above-mentioned single deletion mode, the above-mentioned execution subject determines the virtual element swept by the recognized user manipulation gesture trajectory as the target object. In the above-mentioned group deletion mode, the above-mentioned execution subject determines the group of virtual elements swept by the recognized user manipulation gesture trajectory as the target object.
[0059] The fourth step is to delete the target object in the target virtual scene, wherein the grouping state of the deleted target object including at least one virtual element is ungrouped. In practice, after recognizing the confirmation gesture, the execution subject deletes the connecting lines between the target objects or the block shapes surrounding the target objects, and changes the deleted target object to an ungrouped state.
[0060] The effect flow chart of the above deletion mode can be referred to Figure 3 ,in, Figure 3 Includes the effect flow chart of the single deletion mode and the effect flow chart of the group deletion mode.
[0061] In some optional implementations of some embodiments, the execution subject may perform group operation processing on at least one virtual element in the target virtual scene according to the manipulation gesture type of the user through the following steps:
[0062] In the first step, in response to determining that the target operation mode is the line selection mode, according to the manipulation gesture type, a line selection mode selector is determined, wherein the sub-mode manipulation gesture type set corresponding to the line selection mode includes a finger selection mode manipulation gesture type and a palm selection mode manipulation gesture type. The line selection mode selector may be a virtual tool or virtual object controlled by a user gesture in the line selection mode. The line selection mode selector is used to select, operate or process virtual elements in a virtual scene. The line selection mode selector includes a finger selection mode selector and a palm selection mode selector. The finger selection mode selector is a user's fingertip. The palm selection mode selector is a user's palm. The finger selection mode manipulation gesture type may be a gesture for controlling the movement of the selector identified in the finger selection mode. The palm selection mode manipulation gesture type may be a gesture for controlling the movement of the selector in the palm selection mode. In practice, the execution subject may determine the line selection mode selector according to the sub-mode selected by the user manipulation gesture type. For example, the execution subject recognizes that the target mode is the finger selection mode and uses the user's fingertip as the line selection mode selector.
[0063] The second step is to identify the area where the line selection mode selector follows the user's manipulation gesture. In practice, the execution subject can identify the area where the selector follows the user's manipulation gesture by following the following steps:
[0064] The first sub-step is to track the motion trajectory of the key points to obtain the trajectory of the user's manipulation gestures. The key points are usually specific parts of the hand, such as fingertips, palm center, etc. In practice, the execution subject can use a deep learning framework (such as MediaPipe or YOLO) combined with a camera or depth sensor to capture the trajectory of the user's manipulation gestures. MediaPipe can usually detect the key points of the hand (such as fingertips, palms, etc.) in real time through a pre-trained neural network model. For example, the user draws a circle with his fingertips in a virtual scene. The execution subject detects the key points of the fingertips in real time through a deep learning framework (such as MediaPipe) combined with a camera and a depth sensor, and tracks its motion trajectory to obtain a circular trajectory.
[0065] The second sub-step is to further extract the inner contour and outer contour of the user's manipulation gesture. The inner contour is usually the boundary line inside the gesture. The outer contour is usually the boundary line outside the gesture. In practice, the execution subject can make the inner contour overlap with the outer contour through a scaling operation, and determine the overlapping part as the movement trajectory of the gesture. The scaling operation is usually to adjust the scale of the image or data so that the inner contour overlaps with the outer contour. For example, the user draws a circle with his finger, and the execution subject extracts the inner contour and outer contour of the circle through an image processing algorithm (such as Canny edge detection or Sobel operator). Then, the inner contour and outer contour of the circle are overlapped through a scaling operation, and the overlapping part is determined to be the movement trajectory of the user's gesture. Moreover, if the circle drawn by the user is irregular, the irregular contour can be adjusted to a standard circular trajectory through a scaling operation.
[0066] The third sub-step is to define a three-dimensional space area according to the starting point, end point and intermediate path of the user's manipulation gesture trajectory. The three-dimensional space area is usually the area covered by the user's manipulation gesture trajectory in the three-dimensional space. In practice, the execution subject can use a minimum bounding box or other geometric methods to frame the area swept by the user's manipulation gesture. The minimum bounding box is usually a geometric method for framing the minimum cube or cuboid of a three-dimensional object or trajectory. For example, the user draws a straight line from one point to another with his finger in a virtual scene. The execution subject uses the minimum bounding box algorithm to frame the area covered by this straight line as a three-dimensional space area.
[0067] In the fourth sub-step, during the user operation, the execution subject updates the user's manipulation gesture movement trajectory in real time to dynamically adjust the boundary of the spatial area. The boundary of the spatial area is usually the boundary of the three-dimensional spatial area, which is used to define the range covered by the gesture. In practice, in response to detecting an interruption or abnormality in the user's manipulation gesture trajectory, the execution subject can resume tracking by re-detecting the key points of the gesture, and then dynamically adjust the boundary of the spatial area according to the direction and speed of the user's manipulation gesture.
[0068] The third step is to determine the identified area as the line selection mode selection area. For example, in the palm selection mode, the execution subject may determine the identified area swept by the user's palm as the line selection mode selection area.
[0069] The fourth step is to determine the virtual elements in the line selection mode selection area to obtain at least one virtual element. In practice, the execution subject will select the virtual elements in the line selection mode selection area to obtain at least one virtual element.
[0070] The fifth step is to perform line processing on the at least one virtual element in the target virtual scene. The line processing may be to connect different virtual elements with straight lines. The grouping state of the at least one virtual element after the line processing is a grouped state. The order of the line processing is the same as the order in which the line selection mode selection area is swept over the virtual elements. The order of the line processing shows the selection order of the virtual elements. In practice, after recognizing the confirmation gesture, the execution subject performs line processing on at least one virtual element in the line selection mode selection area, and groups the at least one virtual element connected with a straight line into a group.
[0071] In step 6, in response to detecting a group split operation, and the group split operation corresponds to a virtual element group in the online selection mode, in which the group state of at least two virtual elements is in the grouped state, identifying the path swept by the line selection mode selector following the user's manipulation gesture as the line selection mode split path. The group split operation includes: in the selection mode, the execution subject recognizes that the user uses the selection mode selector to slide the virtual element group obtained by the selection mode at the connection point of the virtual elements in the group. The virtual element connection point includes: the connection line between virtual elements or the connection area between virtual element groups. The selection mode selector includes selectors corresponding to the online selection mode and the block selection mode, respectively. The line selection mode split path can be the user manipulation gesture trajectory recognized by the execution subject when the group split operation is recognized in the online selection mode. In practice, the execution subject can obtain the path swept by the line selection mode selector by recognizing the user manipulation gesture trajectory, and then use the path swept by the line selection mode selector as the line selection mode split path.
[0072] The seventh step is to generate at least one intersection point between the virtual element group corresponding to the group segmentation operation in the line selection mode and the line selection mode segmentation path according to the line selection mode segmentation path. The virtual element group corresponding to the group segmentation operation in the line selection mode includes at least two virtual elements, and the at least two virtual elements are connected by a straight line. The at least one intersection point can be the intersection point between the straight line connecting the at least two virtual elements and the line selection mode segmentation path. In practice, the execution subject can capture at least one intersection point between the virtual element group corresponding to the group segmentation operation in the line selection mode and the line selection mode segmentation path through a depth camera.
[0073] In step 8, based on the at least one intersection, in the target virtual scene, the virtual element group corresponding to the group splitting operation in the line selection mode is split to obtain at least two virtual element groups after the splitting. In practice, after recognizing the confirmation gesture, the execution subject can delete the connection between the group members at the at least one intersection position to perform the splitting operation to obtain at least two virtual element groups after the splitting.
[0074] The effect flow chart of the above line selection mode can be referred to Figure 4 ,in, Figure 4 It includes a grouping effect flow chart of the finger selection mode, a grouping effect flow chart of the palm selection mode, and a group split operation effect flow chart of the line selection mode.
[0075] In some optional implementations of some embodiments, the execution subject may perform group operation processing on at least one virtual element in the target virtual scene according to the manipulation gesture type of the user through the following steps:
[0076] The first step is to identify the closed loop drawn by the block selection mode selector following the user's manipulation gesture in response to determining that the target operation mode is the block selection mode. The block selection mode selector can be the tip of the user's index finger. In practice, the execution subject can capture the movement trajectory of the key point (such as the tip of the index finger) through the deep learning framework (such as MediaPipe or YOLO) combined with a camera or a depth sensor to obtain the user's manipulation gesture trajectory. Then use the image processing algorithm (such as Canny edge detection) to extract the outer contour and inner contour of the user's manipulation gesture. Then simplify the contour by a polygonal approximation algorithm. Finally, check whether the start and end points of the contour coincide by calculating the Euclidean distance, and calculate the contour's area, perimeter and other features to identify whether the shape drawn by the block selection mode selector following the user's manipulation gesture is a closed loop.
[0077] The second step is to determine the area within the closed loop as the selection area of the block selection mode. In practice, the execution subject uses the area surrounded by the identified closed loop as the selection area of the block selection mode.
[0078] The third step is to determine the virtual elements in the block selection mode selection area to obtain at least one virtual element. In practice, the execution subject will select the virtual elements in the block selection mode selection area to obtain at least one virtual element.
[0079] The fourth step is to surround the at least one virtual element with a block shape matching the closed loop shape in the target virtual scene. The grouping state of the at least one virtual element surrounded by the block shape is a grouped state. In practice, after recognizing the confirmation gesture, the execution subject surrounds at least one virtual element in the block selection mode selection area with a block shape similar to the closed loop shape, and groups the at least one virtual element surrounded by the block shape into a group. For example, when the closed loop is a circle, the block shape may be a circle or an ellipse. When the closed loop is a rectangle, the block shape may be a rectangle. When the closed loop is an irregular polygon, the block shape may be a polygon with a similar shape.
[0080] Step 5: In response to detecting a group split operation, and the group split operation in the block selection mode corresponding to the virtual element group includes at least two virtual elements in a grouped state, identifying the path swept by the block selection mode selector following the user's manipulation gesture as the block selection mode split path. The block selection mode split path may be the user manipulation gesture trajectory identified by the execution subject when the group split operation is identified in the block selection mode. In practice, the execution subject may obtain the path swept by the block selection mode selector by identifying the user manipulation gesture trajectory, and then use the path swept by the block selection mode selector as the block selection mode split path.
[0081] The sixth step is to generate, based on the block selection mode segmentation path, at least one cross path between the group of virtual elements corresponding to the group segmentation operation in the block selection mode and the block selection mode segmentation path. The virtual element group corresponding to the group segmentation operation in the block selection mode includes at least two virtual elements, and the at least two virtual elements are surrounded by a block shape. The at least one cross path can be an intersection line between the block shape surrounding the at least two virtual elements and the block selection mode segmentation path. In practice, the execution subject can capture at least one cross path between the virtual element group corresponding to the group segmentation operation in the block selection mode and the block selection mode segmentation path through a depth camera.
[0082] In step 7, based on the at least one cross path, in the target virtual scene, the virtual element group corresponding to the group segmentation operation in the block selection mode is segmented to obtain at least two virtual element groups after segmentation. In practice, after recognizing the confirmation gesture, the execution subject may segment the block shape surrounding the at least two virtual elements according to the shape of the at least one cross path to obtain at least two virtual element groups after segmentation.
[0083] The effect flow chart of the above block selection mode can be referred to Figure 5 ,in, Figure 5 Includes a flow chart of the grouping effect of the block selection mode and a flow chart of the group splitting operation effect of the block selection mode.
[0084] Optionally, the above execution entity may further perform the following steps:
[0085] The first step, in response to detecting the above-mentioned group connection operation, and the grouping state of at least one virtual element included in the virtual element group corresponding to the above-mentioned group connection operation is a grouped state, identifying the connection starting point corresponding to the above-mentioned group connection operation, wherein the above-mentioned starting point corresponds to a virtual element group connected by lines or a virtual element group surrounded by block graphics. The above-mentioned group connection operation includes: in the selection mode, the above-mentioned execution subject recognizes that the user uses the selection mode selector to slide between at least two virtual element groups. In practice, the above-mentioned execution subject can obtain the virtual element group that the path swept by the above-mentioned selection mode selector passes through first by identifying the trajectory of the user's manipulation gesture through a depth camera, thereby identifying the connection starting point corresponding to the above-mentioned group connection operation.
[0086] The second step is to identify at least one intermediate connection starting point corresponding to the group connection operation, wherein each intermediate connection starting point corresponds to at least one virtual element connected by a line or at least one virtual element surrounded by a block graphic. In practice, the execution subject can obtain the group of virtual elements passed by the path swept by the selection mode selector by identifying the user's control gesture trajectory, and obtain the intermediate starting point corresponding to the group connection operation.
[0087] The third step is to identify the connection endpoint corresponding to the group connection operation, wherein the connection endpoint corresponds to at least one virtual element connected by a line or at least one virtual element surrounded by a block graphic. In practice, the execution subject can obtain the virtual element group that the path swept by the selection mode selector passes through last by identifying the user's control gesture trajectory, and obtain the connection endpoint corresponding to the group connection operation.
[0088] In the fourth step, in the target virtual scene, the virtual element group corresponding to the connection starting point, at least one virtual element in each group corresponding to the at least one intermediate connection starting point, and at least one virtual element corresponding to the connection end point are group-connected to obtain a virtual element group connected into one. In practice, after recognizing the confirmation gesture, the execution subject may connect at least two virtual element groups connected by the path swept by the selection mode selector to connect the at least two virtual element groups into one virtual element group.
[0089] The effect flow chart of the above group connection operation can be referred to Figure 6 ,in, Figure 6 Includes a flow chart of the effects of the above-mentioned group connection operation.
[0090] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the bare-hand interaction method in the VR environment of some embodiments of the present disclosure, a natural, efficient and flexible virtual scene interaction can be achieved, and the user experience and interaction efficiency are improved. Specifically, the traditional VR interaction method often relies on external devices (such as handles or gloves), which restricts the natural movements of users and is difficult to adapt to complex gesture operations; at the same time, the interaction mode is single and cannot be flexibly switched, resulting in low operation efficiency. Based on this, the bare-hand interaction method in the VR environment of some embodiments of the present disclosure first obtains the gesture image information of the user corresponding to the target virtual scene. Thus, a real-time visual data basis is provided for subsequent recognition. Then, according to the above-mentioned gesture image information, the user gesture type is identified. Thus, the user gesture type can be identified according to the user's gesture through a customized model. Then, according to the above-mentioned user gesture type, it is matched with a preset control gesture type set to obtain a matching control gesture type. Thus, the identified user gesture type can be identified as a control gesture type. Secondly, in response to determining that the above-mentioned matching control gesture type meets the mode switching condition, the operation mode is switched according to the above-mentioned matching control gesture type. Thus, the operation mode is switched to the operation mode selected by the user. Then, in the target operation mode, the user's manipulation gesture type is identified. Thus, the user's gesture for group operation processing in the selected operation mode can be identified. Finally, according to the manipulation gesture type of the above-mentioned user, at least one virtual element in the above-mentioned target virtual scene is subjected to group operation processing. Thus, objects in the VR environment can be subjected to group operation processing, so that users can complete complex interactive tasks through natural gestures without relying on additional equipment, which significantly improves the naturalness and flexibility of the interaction. At the same time, multi-mode switching improves the interaction efficiency and reduces operation delays. In addition, since the method can accurately identify gestures and dynamically adjust according to the user's operation mode, it can adapt to the diverse interaction needs in different virtual scenes, and enhance the versatility and robustness of the system. Thus, by dynamically matching user gestures with preset operation modes and virtual elements, efficient and natural virtual scene interaction is achieved, which improves the overall user experience.
[0091] Further references Figure 7 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a bare-hand-based interactive device in a VR environment. Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.
[0092] like Figure 7As shown, in some embodiments, the bare-hand-based interaction device 700 in a VR environment includes: an acquisition unit 701, a first recognition unit 702, a matching unit 703, a switching unit 704, a second recognition unit 705, and an operation unit 706. Among them, the acquisition unit 701 is configured to acquire the user's gesture image information corresponding to the target virtual scene; the first recognition unit 702 is configured to recognize the user's gesture type according to the gesture image information; the matching unit 703 is configured to match the user's gesture type with a preset control gesture type set according to the user's gesture type to obtain a matching control gesture type; the switching unit 704 is configured to switch the operation mode according to the matching control gesture type in response to determining that the matching control gesture type meets the mode switching condition; the second recognition unit 705 is configured to recognize the user's control gesture type in the target operation mode; and the operation unit 706 is configured to perform group operation processing on at least one virtual element in the target virtual scene according to the user's control gesture type.
[0093] It is understood that the units described in the device 700 are similar to those described in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 700 and the units included therein, and will not be described in detail here.
[0094] Reference below Figure 8 , which shows a structural schematic diagram of an electronic device 800 suitable for implementing some embodiments of the present disclosure. Figure 8 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0095] like Figure 8 As shown, the electronic device 800 may include a processing device 801 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 809 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0096] Typically, the following devices may be connected to the I / O interface 805: input devices 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; and communication devices 809. The communication devices 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Figure 8 The electronic device 800 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 8 Each block shown in the figure may represent one device, or may represent multiple devices as required.
[0097] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network through a communication device 809, or installed from a storage device 809, or installed from a ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.
[0098] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0099] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0100] The computer-readable medium may be included in the electronic device; or it may exist independently without being installed in the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains the user's gesture image information corresponding to the target virtual scene; identifies the user's gesture type according to the gesture image information; matches the user's gesture type with a preset set of manipulation gesture types to obtain a matching manipulation gesture type; in response to determining that the matching manipulation gesture type satisfies the mode switching condition, switches the operation mode according to the matching manipulation gesture type; in the target operation mode, identifies the user's manipulation gesture type; and performs group operation processing on at least one virtual element in the target virtual scene according to the user's manipulation gesture type.
[0101] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0102] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0103] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor, for example, may be described as: a processor including an acquisition unit, a first recognition unit, a matching unit, a switching unit, a second recognition unit, and an operation unit. The names of these units do not, in some cases, constitute limitations on the units themselves, for example, the first recognition unit may also be described as "a unit for recognizing the type of user gestures based on the above-mentioned gesture image information".
[0104] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0105] Some embodiments of the present disclosure further provide a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned bare-hand-based interaction methods in a VR environment.
[0106] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.
Claims
1. A bare-hand interaction method in a VR environment, comprising: Obtaining gesture image information of the user corresponding to the target virtual scene; Identifying the user's gesture type according to the gesture image information; According to the user gesture type, matching is performed with a preset control gesture type set to obtain a matching control gesture type; In response to determining that the matching manipulation gesture type satisfies a mode switching condition, switching the operation mode according to the matching manipulation gesture type; In the target operation mode, identifying the user's control gesture type; According to the manipulation gesture type of the user, a group operation process is performed on at least one virtual element in the target virtual scene.
2. The method according to claim 1, wherein: In the target operation mode, identifying the user's manipulation gesture type includes: In response to recognizing that the manipulation gesture type of the user is a sub-mode switching gesture, performing sub-mode switching; After the confirmation gesture is recognized, the sub-mode switched to is used as the target mode.
3. The method according to claim 2, wherein: The performing group operation processing on at least one virtual element in the target virtual scene according to the manipulation gesture type of the user includes: In response to determining that the target operation mode is a delete mode, determining a sub-mode of the delete mode according to the manipulation gesture type, wherein the sub-mode manipulation gesture type set corresponding to the delete mode includes a single delete mode manipulation gesture type and a group delete mode manipulation gesture type; Identify at least one virtual element that the deletion mode selector sweeps following the user's manipulation gesture; Determine the scanned at least one virtual element as a target object; In the target virtual scene, the target object is deleted, wherein the grouping state of at least one virtual element included in the deleted target object is an ungrouped state.
4. The method according to claim 2, wherein: The performing group operation processing on at least one virtual element in the target virtual scene according to the manipulation gesture type of the user includes: In response to determining that the target operation mode is a line selection mode, determining a line selection mode selector according to the manipulation gesture type, wherein a sub-mode manipulation gesture type set corresponding to the line selection mode includes a finger selection mode manipulation gesture type and a palm selection mode manipulation gesture type; Identify the spatial area swept by the line selection mode selector following the user's control gesture; determining the spatial region as a line selection mode selection region; Determine the virtual elements within the line selection mode selection area to obtain at least one virtual element; In the target virtual scene, connecting the at least one virtual element, wherein the grouping state of the at least one virtual element after the connecting process is a grouped state; In response to detecting a group split operation, and the virtual element group corresponding to the group split operation in the line selection mode includes at least two virtual elements in a grouped state, identifying a path swept by the line selection mode selector following the user's manipulation gesture as a line selection mode split path; generating, according to the line selection mode segmentation path, at least one intersection point between the virtual element group corresponding to the group segmentation operation in the line selection mode and the line selection mode segmentation path; Based on the at least one intersection, in the target virtual scene, a segmentation process is performed on the virtual element group corresponding to the group segmentation operation in the online selection mode to obtain at least two segmented virtual element groups.
5. The method according to claim 1, wherein: The performing group operation processing on at least one virtual element in the target virtual scene according to the manipulation gesture type of the user includes: In response to determining that the target operation mode is the block selection mode, identifying a closed loop drawn by the block selection mode selector following a manipulation gesture of the user; Determining an area in the closed loop as a block selection mode selection area; Determine the virtual elements within the block selection mode selection area to obtain at least one virtual element; In the target virtual scene, a block graphic matching the closed loop shape surrounds the at least one virtual element, wherein the grouping state of the at least one virtual element surrounded by the block graphic is a grouped state; In response to detecting a group split operation, and the group split operation corresponds to a virtual element group in a block selection mode in which the grouping state of at least two virtual elements is a grouped state, identifying a path swept by the block selection mode selector following a manipulation gesture of a user as a block selection mode split path; According to the block selection mode segmentation path, generating at least one intersection path between the group of virtual elements corresponding to the group segmentation operation in the block selection mode and the block selection mode segmentation path; Based on the at least one cross path, in the target virtual scene, a virtual element group corresponding to the group segmentation operation in the block selection mode is segmented to obtain at least two virtual element groups after segmentation.
6. The method according to claim 4 or 5, wherein: The method further comprises: In response to detecting a group connection operation, and the grouping state of at least one virtual element included in the virtual element group corresponding to the group connection operation is a grouped state, identifying a connection starting point corresponding to the group connection operation, wherein the starting point corresponds to the virtual element group connected by a line or the virtual element group surrounded by a block graphic; Identifying at least one intermediate connection starting point corresponding to the group connection operation, wherein each intermediate connection starting point corresponds to at least one virtual element connected by a line or at least one virtual element surrounded by a block graphic; Identifying a connection endpoint corresponding to the group connection operation, wherein the connection endpoint corresponds to at least one virtual element connected by a line or at least one virtual element surrounded by a block graphic; In the target virtual scene, group connection processing is performed on the virtual element group corresponding to the connection starting point, at least one virtual element in each group corresponding to the at least one intermediate connection starting point, and at least one virtual element corresponding to the connection end point to obtain a virtual element group connected into one.
7. An interactive device based on bare hands in a VR environment, comprising: An acquisition unit, configured to acquire gesture image information of a user corresponding to a target virtual scene; A first recognition unit is configured to recognize a user gesture type according to the gesture image information; A matching unit, configured to match the user gesture type with a preset manipulation gesture type set to obtain a matching manipulation gesture type; A switching unit, configured to, in response to determining that the matching manipulation gesture type satisfies a mode switching condition, switch the operation mode according to the matching manipulation gesture type; A second recognition unit is configured to recognize a user's manipulation gesture type in a target operation mode; The operation unit is configured to perform group operation processing on at least one virtual element in the target virtual scene according to the type of the user's manipulation gesture.
8. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Human-computer interaction method and device based on gesture recognition
CN107272890A
Interaction method and system based on virtual reality device
CN108874126A
Gesture recognition method and device and electronic equipment
CN113190106A
Interaction method and device for virtual object, equipment and storage medium
CN115328309A
VR interaction method, system and device based on gesture recognition and medium
CN118409654A