Method and apparatus for bare-hand interaction in VR environment, electronic device and medium

By recognizing user gestures and switching operation modes through deep convolutional neural networks, the limitations of existing VR interaction methods are overcome, enabling natural and flexible virtual scene interaction and improving user experience and efficiency.

CN120103975BActive Publication Date: 2026-04-24BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIHANG UNIV
Filing Date
2025-02-18
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing VR interaction methods restrict users' natural movements, making it difficult to adapt to complex gesture operations. Furthermore, the interaction modes are limited and cannot be switched flexibly, resulting in low operating efficiency.

Method used

By acquiring the user's gesture image information, a deep convolutional neural network is used to identify the gesture type and match it with a preset set of control gesture types to achieve operation mode switching, identify the user's control gesture type, and perform group operation processing in the virtual scene.

Benefits of technology

It enables natural, efficient, and flexible virtual scene interaction, improves user experience and interaction efficiency, adapts to diverse interaction needs in different virtual scenes, and enhances the system's versatility and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103975B_ABST
    Figure CN120103975B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a bare-hand-based interaction method and device in a VR environment, an electronic device, and a medium. A specific embodiment of the method includes: obtaining gesture image information of a user corresponding to a target virtual scene; identifying a user gesture type according to the gesture image information; matching the user gesture type with a preset set of control gesture types to obtain a matching control gesture type; in response to determining that the matching control gesture type meets a mode switching condition, performing operation mode switching according to the matching control gesture type; in a target operation mode, identifying a control gesture type of the user; and performing group operation processing on at least one virtual element in the target virtual scene according to the control gesture type of the user. This embodiment realizes efficient and natural virtual scene interaction by dynamically matching the user gesture with the preset operation mode and the virtual element, thereby improving the overall experience of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of virtual reality, specifically methods, devices, electronic devices, and media for bare-hand-based interaction in a VR environment. Background Technology

[0002] With the rapid development of virtual reality (VR) technology, the way users interact with virtual scenes is gradually shifting from traditional controllers to more natural and intuitive bare-handed interaction. Currently, the main interaction methods in VR environments are: interacting with objects in the VR environment through external input devices, such as controllers, gloves, or dedicated interaction tools.

[0003] However, when using the above methods to process objects in a VR environment, the following technical problems often arise: the above methods often restrict the user's natural movements, are difficult to adapt to complex gesture operations, and have a single interaction mode that cannot be flexibly switched.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the present disclosure concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion that follows. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure propose methods, devices, electronic devices, and media for bare-hand-based interaction in a VR environment to explain one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure propose a bare-hand-based interaction method in a VR environment. The method includes: acquiring gesture image information of a user in a target virtual scene; identifying the user's gesture type based on the gesture image information; matching the user's gesture type with a preset set of control gesture types to obtain a matched control gesture type; switching the operation mode based on the matched control gesture type in response to determining that the matched control gesture type satisfies a mode switching condition; identifying the user's control gesture type in the target operation mode; and performing group operation processing on at least one virtual element in the target virtual scene based on the user's control gesture type.

[0008] Secondly, some embodiments of this disclosure propose a bare-hand-based interactive device in a VR environment. The device includes: an acquisition unit configured to acquire gesture image information of a user corresponding to a target virtual scene; a first recognition unit configured to recognize the user's gesture type based on the gesture image information; a matching unit configured to match the user's gesture type with a preset set of control gesture types to obtain a matched control gesture type; a switching unit configured to switch the operation mode based on the matched control gesture type in response to determining that the matched control gesture type meets the mode switching condition; a second recognition unit configured to recognize the user's control gesture type in the target operation mode; and an operation unit configured to perform group operation processing on at least one virtual element in the target virtual scene based on the user's control gesture type.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0011] Fifthly, some embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0012] The above embodiments of this disclosure have the following beneficial effects: Through the bare-handed interaction method in a VR environment according to some embodiments of this disclosure, natural, efficient, and flexible virtual scene interaction can be achieved, improving user experience and interaction efficiency. Specifically, traditional VR interaction methods often rely on external devices (such as controllers or gloves), limiting the user's natural movements and making it difficult to adapt to complex gesture operations; at the same time, the interaction mode is singular and cannot be flexibly switched, resulting in low operation efficiency. Based on this, the bare-handed interaction method in a VR environment according to some embodiments of this disclosure first acquires the gesture image information of the user corresponding to the target virtual scene. This provides a real-time visual data foundation for subsequent recognition. Then, based on the gesture image information, the user's gesture type is identified. Thus, the user's gesture type can be identified through a custom model based on the user's gestures. Next, based on the user's gesture type, it is matched with a preset set of control gesture types to obtain a matched control gesture type. Thus, the identified user gesture type can be identified as a control gesture type. Secondly, in response to determining that the matched control gesture type meets the mode switching condition, the operation mode is switched according to the matched control gesture type. Thus, the operation mode is switched to the user-selected operation mode. Next, under the target operation mode, the user's control gesture type is identified. This allows for the identification of the user's gestures used for group operations within the selected operation mode. Finally, based on the user's control gesture type, at least one virtual element in the target virtual scene is subjected to group operation processing. This enables group operation processing of objects in the VR environment, allowing users to complete complex interactive tasks through natural gestures without relying on additional equipment, significantly improving the naturalness and flexibility of the interaction. Simultaneously, multi-mode switching improves interaction efficiency and reduces operation latency. Furthermore, because this method can accurately identify gestures and dynamically adjust according to the user's operation mode, it can adapt to diverse interaction needs in different virtual scenes, enhancing the system's versatility and robustness. Thus, by dynamically matching user gestures with preset operation modes and virtual elements, efficient and natural virtual scene interaction is achieved, improving the overall user experience. Attached Figure Description

[0013] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0014] Figure 1 This is a flowchart of some embodiments of the bare-hand-based interaction method in a VR environment according to the present disclosure;

[0015] Figures 2 to 6These are application scenario diagrams based on some embodiments of bare-hand-based interaction methods in a VR environment;

[0016] Figure 7 This is a schematic diagram of the structure of some embodiments of a bare-hand-based interactive device in a VR environment according to the present disclosure;

[0017] Figure 8 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0019] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0021] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0022] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0023] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0024] Figure 1 A flow 100 is shown illustrating some embodiments of a bare-hand-based interaction method in a VR environment according to this disclosure. This bare-hand-based interaction method in a VR environment includes the following steps:

[0025] Step 101: Obtain the gesture image information of the user's corresponding target virtual scene.

[0026] In some embodiments, the execution subject of a bare-hand-based interaction method in a VR environment (e.g., a head-mounted display device) can acquire user gesture image information in a target virtual scene through equipped sensors. The execution subject can be VR glasses or a VR headset. The equipped sensors include cameras and depth sensors to provide two-dimensional and three-dimensional gesture image data for gesture recognition in different environments. The target virtual scene can be any virtual scene where the user performs group operations. The target virtual scene typically includes at least one virtual element. The virtual element can be a target for user interaction and operation in the virtual scene, including objects in a three-dimensional model (such as virtual furniture, tools, building components, etc.). For example, in a virtual architectural design scene, virtual elements can be virtual walls, doors, windows, furniture, etc. The gesture image information is typically image data containing gesture actions.

[0027] Step 102: Identify the user's gesture type based on the gesture image information.

[0028] In some embodiments, the execution entity can identify the user's gesture type based on the gesture image information. The gesture type includes static gestures where the hand maintains a fixed posture (such as an open palm, a fist, etc.) and dynamic gestures involving hand movement (such as finger movement trajectories, hand waving, etc.). The execution entity can use machine learning algorithms such as convolutional neural networks (CNNs) or support vector machines (SVMs) to identify the gesture type. In practice, the execution entity can extract features from the gesture image information to obtain feature data of the gesture image. Then, the user's gesture type can be identified by inputting the gesture image feature data into the machine learning algorithm.

[0029] In some optional implementations of certain embodiments, the aforementioned execution entity may identify the user gesture type through the following steps:

[0030] The first step involves inputting the aforementioned gesture image information into the input layer of a pre-trained gesture recognition model to obtain pre-processed image information. This gesture recognition model can be a neural network model that takes gesture images as input data and outputs the recognized user gesture type. For example, the neural network model could be a CNN. The gesture recognition model may include the aforementioned input layer, initial convolutional layer, max-pooling layer, a predetermined number of residual block groups, global average pooling layer, flattening layer, fully connected layer, and output layer. In practice, the executing entity can receive the pixel values ​​of the gesture image as input data through the input layer.

[0031] The second step involves inputting the pre-processed image information into the initial convolutional layer to obtain a low-level feature map. This low-level feature map includes hand shape features, hand texture features, and gradient directions. These low-level feature maps typically have high spatial resolution but relatively little semantic information. The kernel size of the initial convolutional layer is usually 3x3 or 5x5. The number of kernels is usually 32 or 64. The stride is usually 1. The padding is usually the same (keeping the input and output sizes the same). The activation function is usually ReLU. In practice, the execution entity can use the initial convolutional layer to perform preliminary feature extraction on the input image to obtain the low-level feature map.

[0032] The third step involves inputting the low-level feature map into the max-pooling layer to obtain the pooled feature map. The pooling window size of the max-pooling layer is typically 2x2. The stride of the max-pooling layer is usually the same as the pooling window size. In practice, the execution entity can use the max-pooling layer to reduce the spatial dimensionality of the feature map, reduce computational cost, and retain important features.

[0033] The fourth step involves sequentially passing the pooled feature maps through a predetermined number of residual block groups for feature extraction, resulting in enhanced feature maps. Each residual block can include three convolutional layers, each followed by a batch normalization layer, an activation function, and skip connections. The kernel size in each residual block's convolutional layers is typically 3x3. The number of kernels in each residual block's convolutional layers can be progressively increased, for example, from 32 to 64, and then to 129. The stride in each residual block's convolutional layers is typically 1, but may be 2 during downsampling. The padding in each residual block's convolutional layers is typically the same. The activation function in each residual block's convolutional layers is typically ReLU. In practice, the execution entity can introduce an attention mechanism into the residual blocks to enhance attention to important feature channels. Furthermore, the execution entity can introduce a feature pyramid between the predetermined number of residual block groups to enhance the recognition capability of gestures at different scales. The aforementioned execution entity can address the vanishing gradient problem in deep network training by performing residual connections on the aforementioned preset number of residual block groups, thereby enhancing feature extraction capabilities. These preset number of residual block groups can be used to obtain more expressive and discriminative feature representations through a series of network structures and mechanisms that perform multi-dimensional feature extraction and optimization on the aforementioned pooled feature maps, resulting in enhanced feature maps.

[0034] The fifth step involves inputting the enhanced feature map into the global average pooling layer to obtain a pooled feature vector. The pooled feature vector typically has a size of (1,1), meaning the enhanced feature map for each channel is compressed into a single value. In practice, the execution entity can calculate the average of all pixel values ​​for each channel's input enhanced feature map using the global average pooling layer. Then, the enhanced feature map for each channel is compressed in height and width to become a single value, reducing the number of parameters while preserving global information. The final pooled feature vector has a length equal to the number of channels in the input enhanced feature map.

[0035] The sixth step involves inputting the pooled feature vector into the flattening layer to obtain the input feature vector. This flattening layer typically transforms a multidimensional vector into a one-dimensional vector through a flattening operation. In practice, the execution entity can use this flattening layer to flatten the multidimensional feature map into a one-dimensional vector for input into the fully connected layer.

[0036] Step 7: Input the aforementioned input feature vector into the fully connected layer to obtain the output vector. The number of input features in the fully connected layer is typically the length of the flattened feature vector, for example, 129. The number of output features in the fully connected layer is typically the number of categories, for example, 6 (for gesture recognition tasks). In practice, the executing agent can perform a linear transformation on the input feature vector using the weight matrix and bias vector in the fully connected layer. The executing agent can introduce non-linearity using the activation function (usually ReLU) in the fully connected layer to integrate local features into global features. Then, the executing agent can map the input feature vector to the target dimension using the learned weight matrix and bias vector in the fully connected layer. The target dimension refers to the dimension of the model's output feature vector, which is usually related to the number of target categories in the task. For example, in gesture recognition tasks, the target dimension is the number of gesture categories.

[0037] Step 8: Input the output vector into the output layer to obtain a set of probability distributions to identify the user gesture type. Each probability distribution in the set corresponds to a gesture type. The length of the probability distribution set is the preset number of gesture types. The output layer typically uses Softmax for multi-class classification tasks. In practice, the executing agent can map the output vector to the probability distribution set through the output layer. Then, the executing agent can use the output layer to identify the gesture type corresponding to the probability distribution that satisfies a preset condition (usually the highest probability value) as the recognized user gesture type.

[0038] Steps one through eight above constitute an inventive point of this disclosure, solving the technical problem of "low recognition accuracy, high model complexity, and insufficient ability to recognize gestures at different scales in existing gesture recognition technologies, making it difficult to adapt to complex gesture operations." The reasons why existing technologies are unable to adapt to complex gesture operations are as follows: existing gesture recognition models struggle to effectively extract key features when dealing with complex backgrounds, gestures at different scales, and the vanishing gradient problem during deep network training, thus affecting recognition accuracy and model efficiency. Solving these factors can improve gesture recognition accuracy, reduce model complexity, and enhance the ability to recognize gestures at different scales. To achieve this, this disclosure employs a deep convolutional neural network (CNN) model. Image data is received through an input layer, and features are extracted through an initial convolutional layer, reduced through a max-pooling layer, enhanced through residual block groups, preserved through a global average pooling layer, converted through a flattening layer, and integrated through a fully connected layer. Finally, classification decisions are made through the output layer, achieving efficient gesture recognition. This effectively extracts key features from gesture images, improves the accuracy and efficiency of gesture recognition, reduces model complexity, and enhances the ability to recognize gestures at different scales. This allows the model to adapt to complex gesture operations.

[0039] Step 103: Match the user's gesture type with a preset set of control gesture types to obtain the matched control gesture type.

[0040] In some embodiments, the executing entity can match the identified user gesture type with the control gestures in the preset control gesture type set to obtain a matched control gesture type. The preset control gesture type set can be pre-defined, operable gestures. The preset control gesture type set includes a set of predefined gesture actions. The preset control gesture type set includes control gestures and switching gestures. In practice, the executing entity can search for the identified user gesture type in the preset control gesture type set to obtain the matched control gesture type.

[0041] The gesture guide diagrams for the aforementioned preset control gesture types can be preset as follows: Figure 2 .in, Figure 2This includes control gestures and switching gestures. Control gestures include: on / off gestures, confirmation gestures, and cancel gestures. Confirmation gestures are typically used to confirm the user's interaction intent with virtual elements or to switch between operation modes. They play a crucial triggering role in the interaction process, ensuring that the user's gestures are accurately recognized and executed by the execution entity. Switching gestures include: switching between selection and deletion modes, switching within selection and deletion modes, and switching between finger and palm modes. Switching within selection and deletion modes includes switching within selection mode and switching within deletion mode. Sub-mode switching gestures include: switching between finger and palm modes and switching within deletion mode.

[0042] Step 104: In response to determining that the matching control gesture type meets the mode switching conditions, switch the operation mode according to the matching control gesture type.

[0043] In some embodiments, the execution entity may switch the operation mode according to the matching control gesture type in response to determining that the matching control gesture type meets the mode switching condition. The mode switching condition is typically that the execution entity recognizes the matching control gesture as the switching gesture. The operation mode may include a selection mode and a deletion mode. Switchable operation modes may include the selection mode and the deletion mode. In the selection mode, at least one virtual element in the virtual scene can be grouped. The selection mode includes a line selection mode and a block selection mode. In the line selection mode, virtual elements in the virtual scene can be selected as a group by connecting them with lines. In the block selection mode, virtual elements in the virtual scene can be selected as a group by surrounding them with blocks. In the deletion mode, grouped virtual elements in the virtual scene can be changed to an ungrouped state. The initial unswitched operation mode is typically the line selection mode. In practice, the execution entity may switch the operation mode when it recognizes that the matching control gesture type is a switching gesture from a preset set of control gesture types. After recognizing the confirmation gesture, the operation mode switch is completed, and the current operation mode is used as the target operation mode. The target operation mode can be the operation mode determined by the executing entity after switching operation modes and recognizing the confirmation gesture. For example, when the executing entity recognizes a user gesture that switches between selection and deletion modes, the operation mode is currently selection mode. After recognizing the confirmation gesture, it switches the selection mode to deletion mode. Similarly, when the executing entity recognizes a user gesture that switches within the selection mode, the operation mode is currently line selection mode. After recognizing the confirmation gesture, it switches the line selection mode to block selection mode.

[0044] Step 105: In the target operation mode, identify the user's control gesture type.

[0045] In some embodiments, the aforementioned execution entity can identify the user's control gesture type within the target operation mode. The user's control gesture type is typically a pre-set gesture used in the operation mode. This user control gesture type includes gestures controlling the movement of the selector and gestures for switching sub-modes. The selector can be a virtual tool or virtual object controlled by the user's gestures in the operation mode. The selector is used to select, manipulate, or process virtual elements in a virtual scene. The form and function of the selector vary depending on the operation mode. For example, in different selection or deletion modes, the selector can be the user's fingertip or palm; in selection mode, the selector can also be used for grouping virtual elements or connecting and separating groups of virtual elements. The gestures controlling the movement of the selector are typically gestures used by the user to scan a path or area when performing group operations with the selector. The group operations are typically methods by which the user interacts with virtual elements in the target virtual scene using the selector in the operation mode. The aforementioned group operation processing includes: the executing entity can change ungrouped virtual elements to grouped states using the aforementioned selection mode; the executing entity can change grouped virtual elements to ungrouped states using the aforementioned deletion mode; and the executing entity can connect and split virtual element groups using the connection and segmentation functions of the aforementioned selection mode. In practice, the executing entity can acquire the user's gesture image information corresponding to the aforementioned target virtual scene under the aforementioned target operation mode. Based on the aforementioned gesture image information, the user's control gesture type is identified. For example, when the aforementioned target operation mode is block selection mode, the executing entity acquires the user's gesture image information in the target virtual scene, uses a deep learning framework (such as MediaPipe) to identify the selector's swiping, and obtains that the user's control gesture type is a gesture controlling the selector's movement.

[0046] In some optional implementations of certain embodiments, the aforementioned execution entity can identify the user's control gesture type in the target operation mode through the following steps:

[0047] The first step involves switching the user's gesture type as a sub-mode switching gesture. The sub-modes of the line selection mode can include finger selection mode and palm selection mode. The finger selection mode can be a line selection mode using the user's fingertip as the selector. The palm selection mode can be a line selection mode using the user's palm as the selector. The sub-modes of the deletion mode include single-entity deletion mode and group deletion mode. In the single-entity deletion mode, the virtual elements scanned by the selector can be changed from a grouped state to an ungrouped state. In the group deletion mode, the virtual elements in a group of virtual elements scanned by the selector can be changed from a grouped state to an ungrouped state. For example, in the line selection mode, the executing entity, through the gesture recognition model, recognizes the finger / palm mode switching gesture, determines that the user's gesture type is a sub-mode switching gesture, and switches between the finger selection mode and the palm selection mode.

[0048] The second step is to switch to the target operation mode after recognizing the confirmation gesture. For example, in the above deletion mode, after the execution entity recognizes the switching gesture within the deletion mode, it switches between the single-entity deletion mode and the group deletion mode. After recognizing the confirmation gesture, the target operation mode is switched to one of the sub-modes of the deletion mode.

[0049] Step 106: Based on the user's control gesture type, perform group operation processing on at least one virtual element in the target virtual scene.

[0050] In some embodiments, the execution entity can perform group operation processing on at least one virtual element in the target virtual scene according to the user's control gesture type under the target operation mode. In practice, the execution entity can identify the path or area swept by the selector of the target operation mode under the target operation mode. Then, it can perform group operation processing on at least one virtual element corresponding to the path or area.

[0051] In some optional implementations of certain embodiments, the execution entity may perform group operation processing on at least one virtual element in the target virtual scene according to the user's control gesture type through the following steps:

[0052] The first step, in response to determining that the target operation mode is a deletion mode, is to determine a sub-mode of the deletion mode based on the aforementioned control gesture type. The set of control gesture types corresponding to the deletion mode includes single-unit deletion mode control gesture types and group deletion mode control gesture types. The single-unit deletion mode control gesture type can be a gesture recognizing the movement of the selector in the single-unit deletion mode. The group deletion mode control gesture type can be a gesture controlling the movement of the selector in the group deletion mode. In practice, the executing entity can determine the sub-mode of the deletion mode based on the aforementioned user control gesture types.

[0053] The second step is to identify at least one virtual element that the deletion mode selector has swept across following the user's gesture. The deletion mode selector can be the user's fingertip. In practice, the execution entity can identify the virtual element that the selector has swept across following the user's gesture using the following steps:

[0054] The first sub-step involves tracking the motion trajectories of key points to obtain the user's gesture trajectory. This user gesture trajectory is typically the movement trajectory of the user controlling the selector in the virtual scene, as identified by the executing entity. This user gesture trajectory can be represented by a curve fitted to the key points. In practice, the executing entity can use a deep learning framework (such as MediaPipe) combined with a camera or depth sensor to capture the user's gesture trajectory. MediaPipe can typically detect hand key points (such as fingertips) in real time using a pre-trained neural network model. In practice, the MediaPipe framework is used in conjunction with a camera to detect the 3D coordinates of hand key points (such as fingertips) in real time. For example, if a user draws a "C" shape in the air, MediaPipe will track the movement trajectory of the fingertips, forming a continuous curve.

[0055] The second sub-step defines the position and extent of the virtual element or group of virtual elements in three-dimensional space. In practice, the aforementioned execution entity can define the extent of the virtual element or group of virtual elements using boundary representation. This boundary representation is typically a method based on the boundaries of geometric entities, defining the extent of the geometry by defining faces, edges, and vertices. For example, in a virtual environment, the position and extent of a virtual button can be defined as a three-dimensional rectangular region using boundary representation. The boundary of this region is determined by the coordinates of its minimum and maximum points.

[0056] The third sub-step involves converting the user's gesture trajectory into three-dimensional coordinates. In practice, the aforementioned execution entity can calculate the specific position of the hand in three-dimensional space based on the depth information of key hand points obtained by a depth sensor (such as Kinect or Intel RealSense) and the two-dimensional image coordinates obtained by the camera. Thus, the execution entity can map two-dimensional pixel coordinates to three-dimensional spatial coordinates based on the depth map, RGB image, and camera intrinsic parameter matrix provided by the depth sensor, achieving the localization and tracking of the user's gesture trajectory in three-dimensional space. For example, if the user's finger has coordinates (x, y) in a two-dimensional image and depth information of z, combined with the camera intrinsic parameter matrix, the finger's coordinates (X, Y, Z) in three-dimensional space can be calculated.

[0057] The fourth sub-step involves using a spatial collision detection algorithm (such as bounding box detection or ray casting) to determine whether the user's gesture trajectory intersects with a virtual element or group of virtual elements. For example, the trajectory of the user's fingertip in 3D space is a straight line from point A to point B. The execution entity uses a ray casting algorithm to treat this straight line as a ray and detect whether it intersects with the bounding box of a virtual element. If the ray intersects with the bounding box, it is considered that the user's gesture trajectory has swept over the virtual element.

[0058] The third step is to identify at least one scanned virtual element as the target object. In practice, in the single-entity deletion mode, the executing entity identifies the virtual element scanned by the user's gesture trajectory as the target object. In the group deletion mode, the executing entity identifies the group of virtual elements scanned by the user's gesture trajectory as the target object.

[0059] Fourth, in the aforementioned target virtual scene, the target objects are deleted. The deleted target objects include at least one virtual element whose grouping state is now ungrouped. In practice, after recognizing the confirmation gesture, the executing entity deletes the lines connecting the target objects or the block shapes surrounding the target objects, and then returns the deleted target objects to an ungrouped state.

[0060] The flowchart illustrating the effects of the above deletion mode can be used as a reference. Figure 3 ,in, Figure 3 This includes flowcharts showing the effects of single-item deletion mode and group deletion mode.

[0061] In some optional implementations of certain embodiments, the execution entity may perform group operation processing on at least one virtual element in the target virtual scene according to the user's control gesture type through the following steps:

[0062] The first step involves determining a line selection mode in response to the determination that the target operation mode is a line selection mode. Based on the control gesture type, a line selection mode selector is determined. The set of sub-mode control gesture types corresponding to the line selection mode includes finger selection mode control gesture types and palm selection mode control gesture types. The line selection mode selector can be a virtual tool or virtual object controlled by user gestures in the line selection mode. The line selection mode selector is used to select, operate, or process virtual elements in a virtual scene. The line selection mode selector includes finger selection mode selectors and palm selection mode selectors. The finger selection mode selector is the user's fingertip. The palm selection mode selector is the user's palm. The finger selection mode control gesture type can be a gesture recognized in the finger selection mode that controls the movement of the selector. The palm selection mode control gesture type can be a gesture recognized in the palm selection mode that controls the movement of the selector. In practice, the executing entity can determine the line selection mode selector based on the sub-mode selected by the user's control gesture type. For example, if the executing entity recognizes the target mode as a finger selection mode, it uses the user's fingertip as the line selection mode selector.

[0063] The second step is to identify the area swept by the line selection mode selector following the user's gestures. In practice, the execution entity can identify the area swept by the selector following the user's gestures using the following steps:

[0064] The first sub-step involves tracking the motion trajectories of key points to obtain the user's gesture trajectory. These key points are typically specific parts of the hand, such as fingertips or the center of the palm. In practice, the executing entity can use a deep learning framework (such as MediaPipe or YOLO) combined with a camera or depth sensor to capture the user's gesture trajectory. MediaPipe, in particular, can typically detect hand key points (such as fingertips and the palm) in real time using a pre-trained neural network model. For example, if a user draws a circle with their fingertips in a virtual scene, the executing entity, using a deep learning framework (such as MediaPipe) combined with a camera and depth sensor, can detect the key points of the fingertips in real time and track their motion trajectory to obtain a circular trajectory.

[0065] The second sub-step further extracts the inner and outer contour lines of the user's gesture. The inner contour line is typically the boundary line inside the gesture, and the outer contour line is typically the boundary line outside the gesture. In practice, the executing entity can use a scaling operation to make the inner and outer contour lines coincide, and the overlapping portion is determined as the movement trajectory of the gesture. This scaling operation typically involves adjusting the scale of the image or data to make the inner and outer contour lines coincide. For example, if a user draws a circle with their finger, the executing entity uses an image processing algorithm (such as Canny edge detection or the Sobel operator) to extract the inner and outer contour lines of this circle. Then, a scaling operation is used to make the inner and outer contour lines of the circle coincide, and the overlapping portion is determined as the movement trajectory of the user's gesture. Moreover, if the circle drawn by the user is irregular, the scaling operation can adjust the irregular contour into a standard circular trajectory.

[0066] The third sub-step defines a three-dimensional spatial region based on the start point, end point, and intermediate path of the user's gesture trajectory. This three-dimensional spatial region is typically the area covered by the user's gesture trajectory in three-dimensional space. In practice, the executing entity can use minimum bounding boxes or other geometric methods to define the area swept by the user's gesture. The minimum bounding box is typically a geometric method used to define the smallest cube or cuboid that encloses a three-dimensional object or trajectory. For example, a user swipes their finger from one point to another in a virtual scene, forming a straight line. The executing entity uses a minimum bounding box algorithm to define the area covered by this straight line as a three-dimensional spatial region.

[0067] The fourth sub-step involves the aforementioned execution entity updating the user's gesture movement trajectory in real time during user operation to dynamically adjust the boundaries of the spatial region. These spatial regions typically refer to the boundaries of the aforementioned three-dimensional spatial region, used to define the range covered by the gesture. In practice, in response to the detection of an interruption or anomaly in the user's gesture trajectory, the execution entity can resume tracking by re-detecting the key points of the gesture, and then dynamically adjust the boundaries of the spatial region based on the direction and speed of the user's gesture.

[0068] The third step is to define the identified area as the selection area for the line selection mode. For example, in the palm selection mode, the executing entity can define the area scanned by the user's palm as the selection area for the line selection mode.

[0069] The fourth step is to identify the virtual elements within the selection area defined by the line selection mode, thus obtaining at least one virtual element. In practice, the executing entity will select the virtual elements within the selection area defined by the line selection mode, thereby obtaining at least one virtual element.

[0070] Fifth, in the aforementioned target virtual scene, at least one virtual element is connected by lines. This connection process can involve linking different virtual elements with straight lines. After the connection process, the at least one virtual element is in a grouped state. The order of the connection process is the same as the order in which the line selection mode scans the virtual elements. This order of connection process indicates the selection order of the virtual elements. In practice, after recognizing the confirmation gesture, the executing entity connects at least one virtual element within the line selection mode selection area and groups the at least one virtual element already connected by straight lines into a single group.

[0071] Step 6: In response to the detection of a group segmentation operation, and given that the virtual element group corresponding to the group segmentation operation in online selection mode includes at least two virtual elements in a grouped state, the path swept by the line selection mode selector following the user's control gesture is identified as the line selection mode segmentation path. The group segmentation operation includes: in selection mode, the executing entity detects that the user uses the selection mode selector to slide at the connection point of virtual elements within the virtual element group obtained through the selection mode. The virtual element connection point includes: a line connecting virtual elements or a connection area between virtual element groups. The selection mode selector includes selectors corresponding to online selection mode and block selection mode, respectively. The line selection mode segmentation path can be the user's control gesture trajectory detected by the executing entity when a group segmentation operation is detected in online selection mode. In practice, the executing entity can obtain the path swept by the line selection mode selector by recognizing the user's control gesture trajectory, and then use the path swept by the line selection mode selector as the line selection mode segmentation path.

[0072] Step 7: Based on the line selection mode segmentation path described above, generate at least one intersection point between the virtual element group corresponding to the group segmentation operation in the online selection mode and the line selection mode segmentation path described above. The virtual element group corresponding to the group segmentation operation in the online selection mode includes at least two virtual elements, and these at least two virtual elements are connected by a straight line. The at least one intersection point can be the intersection of the straight line connecting the at least two virtual elements and the line selection mode segmentation path. In practice, the executing entity can capture at least one intersection point between the virtual element group corresponding to the group segmentation operation in the online selection mode and the line selection mode segmentation path described above using a depth camera.

[0073] Step 8: Based on at least one of the aforementioned intersection points, in the aforementioned target virtual scene, the virtual element groups corresponding to the online selection mode of the aforementioned group segmentation operation are segmented to obtain at least two segmented virtual element groups. In practice, after recognizing the aforementioned confirmation gesture, the executing entity can delete the connections between group members at the aforementioned at least one intersection point to perform the segmentation operation and obtain at least two segmented virtual element groups.

[0074] The flowchart of the above line selection mode effect can be referenced. Figure 4 ,in, Figure 4 This includes flowcharts for the grouping effect in finger selection mode, grouping effect in palm selection mode, and group splitting operation effect in line selection mode.

[0075] In some optional implementations of certain embodiments, the execution entity may perform group operation processing on at least one virtual element in the target virtual scene according to the user's control gesture type through the following steps:

[0076] The first step, in response to determining that the target operation mode is a block selection mode, is to identify the closed loop drawn by the block selection mode selector following the user's gesture. The block selection mode selector can be the user's index fingertip. In practice, the executing entity can use a deep learning framework (such as MediaPipe or YOLO) combined with a camera or depth sensor to capture the movement trajectory of key points (such as the index fingertip) to obtain the user's gesture trajectory. Then, the image processing algorithm (such as Canny edge detection) is used to extract the outer and inner contour lines of the user's gesture. Next, a polygon approximation algorithm is used to simplify the contour. Finally, the Euclidean distance is calculated to check whether the start and end points of the contour coincide, and features such as the area and perimeter of the contour are calculated to identify whether the shape drawn by the block selection mode selector following the user's gesture is a closed loop.

[0077] The second step is to define the area within the closed loop as the selection area for the block selection mode. In practice, the executing entity uses the area enclosed by the identified closed loop as the selection area for the block selection mode.

[0078] The third step is to identify the virtual elements within the selected area of ​​the block selection mode, thus obtaining at least one virtual element. In practice, the executing entity will select the virtual elements within the selected area of ​​the block selection mode, thereby obtaining at least one virtual element.

[0079] Fourth, in the aforementioned target virtual scene, surround the at least one virtual element with a block shape matching the shape of the closed loop. The grouping state of the at least one virtual element surrounded by the block shape is defined as the grouped state. In practice, after recognizing the confirmation gesture, the executing entity surrounds at least one virtual element within the selected area of ​​the block selection mode with a block shape similar to the shape of the closed loop, and groups the at least one virtual element surrounded by the block shape into a group. For example, when the closed loop is a circle, the block shape can be a circle or an ellipse. When the closed loop is a rectangle, the block shape can be a rectangle. When the closed loop is an irregular polygon, the block shape can be a polygon with a similar shape.

[0080] Fifth, in response to the detection of a group segmentation operation, and given that the virtual element group corresponding to the group segmentation operation in block selection mode includes at least two virtual elements in a grouped state, the path swept by the block selection mode selector following the user's control gesture is identified as the block selection mode segmentation path. This block selection mode segmentation path can be the trajectory of the user's control gesture detected by the execution entity when a group segmentation operation is detected in block selection mode. In practice, the execution entity can obtain the path swept by the block selection mode selector by recognizing the user's control gesture trajectory, and then use this path as the block selection mode segmentation path.

[0081] Step 6: Based on the block selection mode segmentation path described above, generate at least one intersection path between the group of virtual elements corresponding to the block selection mode segmentation operation and the block selection mode segmentation path. The group of virtual elements corresponding to the block selection mode segmentation operation includes at least two virtual elements, and these at least two virtual elements are surrounded by a block shape. The at least one intersection path can be the line of intersection between the block shape surrounding the at least two virtual elements and the block selection mode segmentation path. In practice, the executing entity can use a depth camera to capture at least one intersection path between the group of virtual elements corresponding to the block selection mode segmentation operation and the block selection mode segmentation path.

[0082] Step 7: Based on the aforementioned at least one intersecting path, in the aforementioned target virtual scene, the virtual element group corresponding to the aforementioned group segmentation operation in block selection mode is segmented to obtain at least two segmented virtual element groups. In practice, after recognizing the aforementioned confirmation gesture, the executing entity can segment the block shape surrounding at least two virtual elements according to the shape of the aforementioned at least one intersecting path to obtain at least two segmented virtual element groups.

[0083] The flowchart of the above block selection mode can be used as a reference. Figure 5 ,in, Figure 5 This includes flowcharts showing the grouping effect of the block selection mode and the group splitting operation effect of the block selection mode.

[0084] Optionally, the aforementioned implementing entity may also perform the following steps:

[0085] First, in response to detecting the aforementioned group connection operation, and that at least one virtual element in the virtual element group corresponding to the group connection operation is in a grouped state, the connection starting point corresponding to the group connection operation is identified. This starting point corresponds to a line-connected virtual element group or a virtual element group surrounded by a block-shaped graphic. The group connection operation includes: in selection mode, the executing entity detects that the user is using a selection mode selector to slide between at least two virtual element groups. In practice, the executing entity can use a depth camera to identify the user's gesture trajectory to obtain the first virtual element group traversed by the path swept by the selection mode selector, thereby identifying the connection starting point corresponding to the group connection operation.

[0086] The second step is to identify at least one intermediate connection starting point corresponding to the above group connection operation, wherein each intermediate connection starting point corresponds to at least one virtual element connected by a line or at least one virtual element surrounded by a block-shaped graphic. In practice, the execution entity can obtain the group of virtual elements passed through the path swept by the selection mode selector by recognizing the user's control gesture trajectory, and thus obtain the intermediate starting point corresponding to the above group connection operation.

[0087] The third step is to identify the connection endpoint corresponding to the above group connection operation, wherein the connection endpoint corresponds to at least one virtual element connected by a line or at least one virtual element surrounded by a block graphic. In practice, the execution entity can obtain the last group of virtual elements traversed by the path swept by the selection mode selector by recognizing the user's control gesture trajectory, and thus obtain the connection endpoint corresponding to the above group connection operation.

[0088] Fourth, in the aforementioned target virtual scene, group connection processing is performed on the virtual element group corresponding to the connection starting point, at least one virtual element in each group corresponding to at least one intermediate connection starting point, and at least one virtual element corresponding to the connection ending point, to obtain a virtual element group connected as one. In practice, after recognizing the aforementioned confirmation gesture, the executing entity can connect at least two virtual element groups linked by the path scanned by the aforementioned selection mode selector, thus connecting the at least two virtual element groups into one virtual element group.

[0089] The flowchart of the above group connection operation can be used as a reference. Figure 6 ,in, Figure 6 The flowchart shows the effect of the above group connection operation.

[0090] The above embodiments of this disclosure have the following beneficial effects: Through the bare-handed interaction method in a VR environment according to some embodiments of this disclosure, natural, efficient, and flexible virtual scene interaction can be achieved, improving user experience and interaction efficiency. Specifically, traditional VR interaction methods often rely on external devices (such as controllers or gloves), limiting the user's natural movements and making it difficult to adapt to complex gesture operations; at the same time, the interaction mode is singular and cannot be flexibly switched, resulting in low operation efficiency. Based on this, the bare-handed interaction method in a VR environment according to some embodiments of this disclosure first acquires the gesture image information of the user corresponding to the target virtual scene. This provides a real-time visual data foundation for subsequent recognition. Then, based on the gesture image information, the user's gesture type is identified. Thus, the user's gesture type can be identified through a custom model based on the user's gestures. Next, based on the user's gesture type, it is matched with a preset set of control gesture types to obtain a matched control gesture type. Thus, the identified user gesture type can be identified as a control gesture type. Secondly, in response to determining that the matched control gesture type meets the mode switching condition, the operation mode is switched according to the matched control gesture type. Thus, the operation mode is switched to the user-selected operation mode. Next, under the target operation mode, the user's control gesture type is identified. This allows for the identification of the user's gestures used for group operations within the selected operation mode. Finally, based on the user's control gesture type, at least one virtual element in the target virtual scene is subjected to group operation processing. This enables group operation processing of objects in the VR environment, allowing users to complete complex interactive tasks through natural gestures without relying on additional equipment, significantly improving the naturalness and flexibility of the interaction. Simultaneously, multi-mode switching improves interaction efficiency and reduces operation latency. Furthermore, because this method can accurately identify gestures and dynamically adjust according to the user's operation mode, it can adapt to diverse interaction needs in different virtual scenes, enhancing the system's versatility and robustness. Thus, by dynamically matching user gestures with preset operation modes and virtual elements, efficient and natural virtual scene interaction is achieved, improving the overall user experience.

[0091] Further reference Figure 7 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a bare-hand-based interactive device in a VR environment. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0092] like Figure 7As shown, a VR environment-based bare-hand interactive device 700 in some embodiments includes: an acquisition unit 701, a first recognition unit 702, a matching unit 703, a switching unit 704, a second recognition unit 705, and an operation unit 706. The acquisition unit 701 is configured to acquire gesture image information of a user corresponding to a target virtual scene; the first recognition unit 702 is configured to recognize the user's gesture type based on the gesture image information; the matching unit 703 is configured to match the user's gesture type with a preset set of control gesture types to obtain a matched control gesture type; the switching unit 704 is configured to switch the operation mode according to the matched control gesture type in response to determining that the matched control gesture type meets the mode switching condition; the second recognition unit 705 is configured to recognize the user's control gesture type in the target operation mode; and the operation unit 706 is configured to perform group operation processing on at least one virtual element in the target virtual scene according to the user's control gesture type.

[0093] It is understandable that the units described in the device 700 are related to the reference. Figure 1 The steps in the described method correspond accordingly. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 700 and the units contained therein, and will not be repeated here.

[0094] The following is for reference. Figure 8 It shows a schematic diagram of the structure of an electronic device 800 suitable for implementing some embodiments of the present disclosure. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0095] like Figure 8 As shown, the electronic device 800 may include a processing unit 801 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 809 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0096] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 8 Each box shown can represent a device or multiple devices as needed.

[0097] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 809, or installed from a ROM 802. When the computer program is executed by the processing device 801, it performs the functions defined in the methods of some embodiments of this disclosure.

[0098] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0099] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0100] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire gesture image information of a user corresponding to a target virtual scene; identify the user's gesture type based on the gesture image information; match the user's gesture type with a preset set of control gesture types to obtain a matched control gesture type; in response to determining that the matched control gesture type satisfies a mode switching condition, switch the operation mode based on the matched control gesture type; in the target operation mode, identify the user's control gesture type; and perform group operation processing on at least one virtual element in the target virtual scene based on the user's control gesture type.

[0101] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0102] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0103] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an acquisition unit, a first recognition unit, a matching unit, a switching unit, a second recognition unit, and an operation unit. The names of these units do not necessarily limit the specific unit; for example, the first recognition unit may also be described as "a unit that recognizes the user's gesture type based on the aforementioned gesture image information."

[0104] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0105] Some embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any of the above-described bare-hand-based interaction methods in a VR environment.

[0106] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A bare-hand-based interaction method in a VR environment, comprising: Obtain the gesture image information of the user corresponding to the target virtual scene; Based on the gesture image information, the user's gesture type is identified; Based on the user's gesture type, it is matched with a preset set of control gesture types to obtain the matched control gesture type; In response to determining that the matched control gesture type meets the mode switching condition, the operation mode is switched according to the matched control gesture type; In the target operation mode, identify the type of user's control gesture; Based on the user's gesture type, group operation processing is performed on at least one virtual element in the target virtual scene, including: In response to determining that the target operation mode is a block selection mode, the closed loop drawn by the block selection mode selector following the user's control gesture is identified; The region within the closed loop is defined as the block selection mode selection region; Determine the virtual elements within the selected area of ​​the block selection mode to obtain at least one virtual element; In the target virtual scene, at least one virtual element is surrounded by block graphics that match the shape of the closed loop, wherein the grouping state of the at least one virtual element after being surrounded by block graphics is the grouped state; In response to the detection of a group splitting operation, and the group splitting operation in the block selection mode includes a virtual element group with at least two virtual elements in a grouped state, the path swept by the block selection mode selector following the user's control gesture is identified as the block selection mode splitting path. Based on the block selection mode segmentation path, generate a group of virtual elements corresponding to the block selection mode in the block selection mode and at least one intersection path with the block selection mode segmentation path; Based on the at least one intersection path, in the target virtual scene, the virtual element group corresponding to the group segmentation operation in block selection mode is segmented to obtain at least two segmented virtual element groups.

2. The method according to claim 1, wherein, The process of identifying the user's control gesture type in the target operation mode includes: In response to the recognition that the user's control gesture type is a sub-mode switching gesture, a sub-mode switch is performed; After recognizing the confirmation gesture, the sub-mode switched to will be used as the target mode.

3. The method according to claim 2, wherein, The step of performing group operation processing on at least one virtual element in the target virtual scene based on the user's control gesture type includes: In response to determining that the target operation mode is a deletion mode, a sub-mode of the deletion mode is determined according to the control gesture type, wherein the set of sub-mode control gesture types corresponding to the deletion mode includes single deletion mode control gesture types and group deletion mode control gesture types. The deletion mode selector identifies at least one virtual element that is swept across by the user's gesture. The scanned virtual element is identified as the target object; In the target virtual scene, the target object is deleted, wherein at least one virtual element included in the deleted target object is in an ungrouped state.

4. The method according to claim 2, wherein, The step of performing group operation processing on at least one virtual element in the target virtual scene based on the user's control gesture type includes: In response to determining that the target operation mode is a line selection mode, a line selection mode selector is determined according to the control gesture type, wherein the set of sub-mode control gesture types corresponding to the line selection mode includes finger selection mode control gesture type and palm selection mode control gesture type. Identify the spatial area swept by the line selection mode selector following the user's control gesture; The spatial region is defined as the selection region for the line selection mode; Determine the virtual elements within the selected area of ​​the line selection mode to obtain at least one virtual element; In the target virtual scene, at least one virtual element is connected by lines, wherein the grouping state of the at least one virtual element after the connection process is a grouped state. In response to the detection of a group segmentation operation, and the virtual element group corresponding to the group segmentation operation in the online selection mode includes at least two virtual elements whose grouping status is already grouped, the path swept by the line selection mode selector following the user's control gesture is identified as the line selection mode segmentation path. Based on the line selection mode segmentation path, generate at least one intersection point between the virtual element group corresponding to the line selection mode in the group segmentation operation and the line selection mode segmentation path; Based on the at least one intersection point, in the target virtual scene, the virtual element group corresponding to the group segmentation operation in the online selection mode is segmented to obtain at least two segmented virtual element groups.

5. The method according to claim 4, wherein, The method further includes: In response to the detection of a group connection operation, and the grouping status of at least one virtual element in the virtual element group corresponding to the group connection operation is a grouped state, the connection start point corresponding to the group connection operation is identified, wherein the connection start point corresponds to the virtual element group connected by a line or the virtual element group surrounded by a block graphic. Identify at least one intermediate connection start point corresponding to the group connection operation, wherein each intermediate connection start point corresponds to at least one virtual element connected by a line or at least one virtual element surrounded by a block graphic; Identify the connection endpoint corresponding to the group connection operation, wherein the connection endpoint corresponds to at least one virtual element connected by a line or at least one virtual element surrounded by a block graphic; In the target virtual scene, the virtual element group corresponding to the connection start point, at least one virtual element in each group corresponding to the at least one intermediate connection start point, and at least one virtual element corresponding to the connection end point are grouped together to obtain a virtual element group connected as one.

6. A bare-hand-based interactive device for a VR environment, comprising: The acquisition unit is configured to acquire gesture image information of the user's corresponding target virtual scene; The first recognition unit is configured to recognize the user's gesture type based on the gesture image information; The matching unit is configured to match the user's gesture type with a preset set of control gesture types to obtain a matched control gesture type. The switching unit is configured to switch the operation mode according to the matching control gesture type in response to determining that the matching control gesture type meets the mode switching condition. The second recognition unit is configured to recognize the user's control gesture type in the target operation mode; The operation unit is configured to perform group operation processing on at least one virtual element in the target virtual scene according to the user's control gesture type, including: In response to determining that the target operation mode is a block selection mode, the closed loop drawn by the block selection mode selector following the user's control gesture is identified; The region within the closed loop is defined as the block selection mode selection region; Determine the virtual elements within the selected area of ​​the block selection mode to obtain at least one virtual element; In the target virtual scene, at least one virtual element is surrounded by block graphics that match the shape of the closed loop, wherein the grouping state of the at least one virtual element after being surrounded by block graphics is the grouped state; In response to the detection of a group splitting operation, and the group splitting operation in the block selection mode includes a virtual element group with at least two virtual elements in a grouped state, the path swept by the block selection mode selector following the user's control gesture is identified as the block selection mode splitting path. Based on the block selection mode segmentation path, generate a group of virtual elements corresponding to the block selection mode in the block selection mode and at least one intersection path with the block selection mode segmentation path; Based on the at least one intersection path, in the target virtual scene, the virtual element group corresponding to the group segmentation operation in block selection mode is segmented to obtain at least two segmented virtual element groups.

7. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 5.

8. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Human-computer interaction method and device based on gesture recognition

    CN107272890A

  • Gesture recognition method and device and electronic equipment

    CN113190106A