Object moving method, device and equipment based on bare hand gestures and medium

By constructing manipulated cone models and auxiliary cone models in a virtual reality environment, responding to the user's naked hand gestures, the problems of unnatural operation and difficulty in movement in the occlusion environment caused by traditional controller ray methods are solved, and more natural interaction and efficient object movement are achieved.

CN120066275APending Publication Date: 2025-05-30BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510226478.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In virtual reality environments, traditional controller ray methods make users unnatural operations and it is difficult to accurately move target objects in high occlusion environments, increasing user operation and cognitive burden.

Method used

By building manipulated cone models and auxiliary cone models, in response to the user's naked hand switching and sliding gestures, the target object can be accurately moved and the occlusion will be removed.

Benefits of technology

It improves the naturalness and immersion of user interaction, reduces the user's operation burden, and enhances the movement accuracy and efficiency in high occlusion environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066275A_ABST
    Figure CN120066275A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an object moving method and device based on a bare hand gesture, equipment and a medium. A specific embodiment of the method comprises the following steps: in response to a bare hand switching gesture detected for the first time, carrying out model initialization on a cone model to obtain a controllable cone model and an auxiliary cone model; in response to determining that the controllable cone model completes initialization and the bare hand sliding gesture is detected for the first time, locking a bottom mapping region to generate a locked region; and in response to determining that the bottom mapping area is locked and the bare-hand sliding gesture is continuously detected, correspondingly moving the vertex of the controllable cone model and the vertex of the auxiliary cone model according to movement information corresponding to the bare-hand sliding gesture. According to the implementation mode, the dependence of man-machine interaction on physical equipment in the high-shielding virtual environment can be reduced, and the shielding object removal of the target object in the high-shielding virtual environment can be efficiently realized through various types of operation gestures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular, to a method, apparatus, device, and medium for moving an object based on bare - hand gestures. Background Art

[0002] In the application of virtual reality technology, the movement of objects in a virtual environment has become one of the basic operations for user interaction. In many application scenarios, such as intelligent manufacturing, medical simulation, and games, the object movement task is the key to further operations. The existing movement methods mainly use the ray of a controller to move the object.

[0003] However, in practice, it is found that when performing movement operations in a virtual reality environment using the above - mentioned method, the following technical problems often exist:

[0004] First, the traditional controller ray method usually needs to rely on physical devices, such as a handheld controller or a remote control, to move the object. Although this method can improve the movement accuracy to a certain extent, it sacrifices the naturalness and immersion of the interaction. Users must operate external devices to complete the interaction, and this dependence limits the freedom of interaction, making the operation of users in the virtual environment less intuitive and natural. Especially in virtual reality technology (VR), bare - hand interaction is considered to be a more human - like operation habit and more immersive way. Therefore, excessive dependence on physical devices obviously weakens the advantages of the natural interaction experience in the virtual environment.

[0005] Second, in a highly complex and severely occluded virtual environment, due to the occlusion between objects, the ray method often cannot accurately move the target object, and is prone to misselection, especially in a scene with dense or severely occluded objects. In addition, users may need to frequently adjust the ray position or use multiple operations to complete a simple movement task. Such frequent interaction operations not only increase the operation burden of users, but also increase the cognitive burden and reduce the overall user experience. Especially after long - term use, it is easy to cause fatigue and discomfort.

[0006] The above information disclosed in this background art section is only used to enhance the understanding of the background of the concept of the present disclosure. Therefore, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention

[0007] This summary of the present disclosure is used to briefly introduce concepts that will be described in detail in the following detailed implementation section. This summary of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0008] Some embodiments of the present disclosure propose a method, apparatus, device, and medium for moving an object based on bare - hand gestures to solve one or more of the technical problems mentioned in the above - mentioned background art section.

[0009] In a first aspect, some embodiments of the present disclosure provide a method for moving an object based on bare - hand gestures. The method includes: in response to first detecting a bare - hand switching gesture corresponding to a target user, initializing a cone model to obtain a manipulable cone model and an auxiliary cone model. Wherein, the manipulable cone model and the auxiliary cone model are constructed based on a double - bend camera model; in response to determining that the initialization of the manipulable cone model is completed and first detecting a bare - hand sliding gesture corresponding to the target user, locking the bottom mapping area corresponding to the manipulable cone model to generate a locked area. Wherein, the locked area is determined based on the position of the bare - hand sliding gesture corresponding to the target user at the current time; in response to determining that the locking of the bottom mapping area corresponding to the manipulable cone model is completed and continuously detecting the bare - hand sliding gesture corresponding to the target user, according to the movement information corresponding to the bare - hand sliding gesture, correspondingly moving the vertices of the manipulable cone model and the vertices of the auxiliary cone model, so that the positions of the projections of the target object in the locked area on the first projection plane and the second projection plane move correspondingly, realizing the removal of the occluder of the target object.

[0010] In a second aspect, some embodiments of the present disclosure provide a device for moving an object based on bare - hand gestures, including: a model initialization unit configured to, in response to first detecting a bare - hand switching gesture corresponding to a target user, initialize a cone model to obtain a manipulable cone model and an auxiliary cone model. Wherein, the manipulable cone model and the auxiliary cone model are constructed based on a double - bend camera model; a locking unit configured to, in response to determining that the initialization of the manipulable cone model is completed and first detecting a bare - hand sliding gesture corresponding to the target user, lock the bottom mapping area corresponding to the manipulable cone model to generate a locked area. Wherein, the locked area is determined based on the position of the bare - hand sliding gesture corresponding to the target user at the current time; a moving unit configured to, in response to determining that the locking of the bottom mapping area corresponding to the manipulable cone model is completed and continuously detecting the bare - hand sliding gesture corresponding to the target user, according to the movement information corresponding to the bare - hand sliding gesture, correspondingly move the vertices of the manipulable cone model and the vertices of the auxiliary cone model, so that the positions of the projections of the target object in the locked area on the first projection plane and the second projection plane move correspondingly, realizing the removal of the occluder of the target object.

[0011] In a third aspect, some embodiments of the present disclosure provide a device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner of the first aspect.

[0012] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through a method for moving an object based on bare - hand gestures in some embodiments of the present disclosure, by constructing a manipulable cone model and an auxiliary cone model, a user can complete the movement of a target object using bare - hand gestures, enhancing the user experience. Specifically, the reasons for increasing the user's manipulation burden are as follows: completely relying on physical devices restricts the degree of freedom of interaction and sacrifices the naturalness and immersion of interaction. Based on this, in a method for moving an object based on bare - hand gestures in some embodiments of the present disclosure, first, in response to the first detection of a bare - hand switching gesture corresponding to a target user, the cone model is initialized to obtain a manipulable cone model and an auxiliary cone model. Here, through the construction of the manipulable cone model and the auxiliary cone model, the user gets rid of the dependence on physical devices and controls the target object by controlling the manipulable cone model and the auxiliary cone model. Next, in response to determining that the above - mentioned manipulable cone model has been initialized and the first detection of a bare - hand sliding gesture corresponding to the above - mentioned target user, the bottom mapping area corresponding to the above - mentioned manipulable cone model is locked to generate a locked area. Among them, the above - mentioned locked area is determined based on the position of the bare - hand sliding gesture corresponding to the target user at the current time. Here, by locking the bottom mapping area corresponding to the manipulable cone model, the target object and the target occluder within the mapping area can be accurately locked to reduce the manipulation burden of the target user. Finally, in response to determining that the locking of the bottom mapping area corresponding to the above - mentioned manipulable cone model is completed and continuously detecting the bare - hand sliding gesture corresponding to the above - mentioned target user, according to the movement information corresponding to the bare - hand sliding gesture, the vertices of the above - mentioned manipulable cone model and the vertices of the above - mentioned auxiliary cone model are correspondingly moved so that the positions of the projections of the target object within the above - mentioned locked area on the first projection plane and the second projection plane are correspondingly moved, realizing the removal of the occluder of the above - mentioned target object. Here, through the corresponding movement of the target object and the target occluder within the above - mentioned locked area on the second projection plane, the naturalness and immersion of interaction can be enhanced, and the occluder of the target object can be removed while moving the target object. In summary, through the method of using bare - hand gestures to move an object in a high - occlusion virtual environment, the manipulation burden of the target user can be reduced and the user experience can be enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.

[0014] Figure 1 is a flowchart of some embodiments of a method for moving an object based on bare - hand gestures according to the present disclosure;

[0015] Figure 2 is a schematic diagram of a scenario suitable for implementing some embodiments of a method for moving an object based on bare - hand gestures according to the present disclosure;

[0016] Figure 3 is a schematic diagram of a scenario suitable for implementing other embodiments of a method for moving an object based on bare - hand gestures according to the present disclosure;

[0017] Figure 4 is a gesture state transition diagram of some embodiments of a method for moving an object based on bare - hand gestures according to the present disclosure;

[0018] Figure 5 is a schematic structural diagram of some embodiments of an apparatus for moving an object based on bare - hand gestures according to the present disclosure;

[0019] Figure 6 is a schematic structural diagram of a device suitable for implementing some embodiments of the present disclosure. Specific Embodiments

[0020] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0021] In addition, it should be noted that for the sake of convenience of description, only the parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0022] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependent relationships.

[0023] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0024] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0025] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0026] Reference Figure 1 , a flowchart 100 of some embodiments of an object moving method based on bare - hand gestures according to the present disclosure is shown. The above - mentioned object moving method based on bare - hand gestures includes the following steps:

[0027] Step 101, in response to first detecting a bare - hand switching gesture corresponding to a target user, initialize the cone model to obtain a manipulable cone model and an auxiliary cone model.

[0028] In some embodiments, in response to first detecting a bare - hand switching gesture corresponding to a target user, an execution entity (e.g., a head - mounted device) initializes a cone model to obtain a manipulable cone model and an auxiliary cone model. The above - mentioned execution entity can be an electronic device capable of receiving specific information in a virtual scene and performing specific tasks. The above - mentioned specific information can be the bare - hand gesture information of the above - mentioned target user. The above - mentioned target user can be a user performing bare - hand gesture operations. The above - mentioned specific task can be a purposeful activity of the above - mentioned target user (e.g., the activity is to move an object in the above - mentioned virtual scene). The above - mentioned virtual scene can be any virtual scene that interacts with the above - mentioned target user. The above - mentioned virtual scene generally includes at least one virtual element. The above - mentioned virtual element can be a target for the above - mentioned target user to interact with and operate in the above - mentioned virtual scene, including objects in a three - dimensional model (such as virtual furniture, tools, etc.) and the corresponding projections of the above - mentioned objects. For example, in a virtual building design, the above - mentioned virtual elements can be virtual tables, stools, sofas, models, etc. The above - mentioned cone model can be a virtual element of a cone in the above - mentioned virtual environment. The above - mentioned cone model can be used to provide a natural interaction method for the above - mentioned target user to operate conveniently. The above - mentioned interaction method can refer to a method of information exchange between the above - mentioned target user and the above - mentioned execution entity. The above - mentioned model initialization can be to construct the above - mentioned manipulable cone model and the above - mentioned auxiliary cone model based on a dual - bend camera model. The above - mentioned dual - bend camera model can be a virtual camera model. The above - mentioned dual - bend camera model has two frustums. The above - mentioned frustum can be a cone model used to simulate human eye vision. In practice, constructing the above - mentioned manipulable cone model and the above - mentioned auxiliary cone model based on the dual - bend camera model includes:

[0029] In the first step, construct two above - mentioned cone models corresponding to the above - mentioned manipulable cone model and the above - mentioned auxiliary cone model.

[0030] In the second step, configure the parameters of the above - mentioned manipulable cone model and the auxiliary cone model to complete the model initialization. The above - mentioned manipulable cone model and the above - mentioned auxiliary cone model can be configured as shown in Figure 2 The above - mentioned manipulable cone model m can include three parameters: p 1 、r 1 and v m . Among them, p 1 can be the position of the bottom center point corresponding to the above - mentioned manipulable cone model m, r 1 can be the bottom radius corresponding to the above - mentioned manipulable cone model m, and v m can be the position of the vertex of the above - mentioned manipulable cone model m. The above - mentioned auxiliary cone model a can also include three parameters: p 2 、r 2 and v a . Among them, p2 It can be the position of the center point at the bottom corresponding to the above-mentioned auxiliary cone model a, r 2 It can be the bottom radius corresponding to the above-mentioned auxiliary cone model a, v a It can be the position of the vertex of the above-mentioned auxiliary cone model a. In the initial construction state, v m and v a are the same as the viewpoint position of the above-mentioned target user. The above-mentioned viewpoint position can be the position of the eyes of the above-mentioned target user. V can be the above-mentioned viewpoint position. The above p 1 and p 2 can be constructed at positions 20 meters and 16 meters in front of the above-mentioned V respectively. r 1 can be 0.08 of the screen height. r 2 It can be r 1 1.2 times of.

[0031] The above-mentioned manipulable model may be the above-mentioned virtual element of a cone for projecting a target object onto a first projection plane. The above-mentioned auxiliary cone model may be the above-mentioned virtual element of a cone for projecting the above-mentioned target object onto a second projection plane. The above-mentioned manipulable cone model may allow the above-mentioned target user to change the position of the vertex of the above-mentioned manipulable cone model through a specific operation (for example, the operation may be a bare-handed sliding gesture). The above-mentioned target object may be any of the above-mentioned virtual elements in the above-mentioned virtual environment. The above-mentioned first projection plane may be a plane at a preset ratio (for example, the preset ratio is 0.8) of the straight line from the center point of the bottom corresponding to the above-mentioned manipulable cone model to the viewpoint position and perpendicular to the above-mentioned straight line. The above-mentioned second projection plane may be a plane at a preset ratio (for example, the preset ratio is 0.5) of the straight line from the center point of the bottom corresponding to the above-mentioned manipulable cone model to the viewpoint position and perpendicular to the above-mentioned straight line. The above-mentioned bare-handed gesture may be a preset hand posture, including: a bare-handed switching gesture, a bare-handed sliding gesture, and a bare-handed selection gesture. The above-mentioned bare-handed switching gesture may be a gesture for controlling the start and stop of the above-mentioned manipulable cone model in the form of bare hands. In practice, the hand posture corresponding to the above-mentioned bare-handed switching gesture may be preset by those skilled in the relevant art. For example, the hand posture corresponding to the above-mentioned bare-handed switching gesture is that the thumb and index finger of the above-mentioned target user are kept straight, while the other three fingers are curled inward. The above-mentioned bare-handed sliding gesture may be that the thumb and index finger of the above-mentioned target user are in a pinching state. The above-mentioned bare-handed selection gesture may be that the thumb and middle finger of the above-mentioned target user are pinched. In practice, the above-mentioned first detection may be that the above-mentioned execution entity first detects the above-mentioned bare-handed gesture in a preset state. The above-mentioned preset state may include an initial construction state, a cone control state, an intermediate state, a selection state, and a standby state. The above-mentioned initial construction state may be a state where the above-mentioned model initialization is completed. The above-mentioned cone control state may be a state that allows the target user to manipulate the vertex of the above-mentioned manipulable cone model. The above-mentioned selection state may be a state where the above-mentioned target user performs the selection of the above-mentioned target object. The above-mentioned intermediate state may be between the above-mentioned cone control state and the above-mentioned selection state. The above-mentioned standby state may be a state where the above-mentioned execution entity is in standby. The above-mentioned bare-handed switching gesture corresponding to the above-mentioned target user may be the above-mentioned bare-handed switching gesture being performed by the target user. The above-mentioned target object may be any object in the above-mentioned virtual environment.

[0032] In some alternative implementation manners of some embodiments, the above-mentioned bare-handed switching gesture corresponding to the above-mentioned target user may be detected through the following steps:

[0033] In the first step, input the bare - hand gesture image into the input layer of the first model to obtain a pre - processed image. Among them, the above - mentioned first model can include: the above - mentioned input layer, extraction layer, and first fully - connected layer. The above - mentioned bare - hand gesture image can be an image containing the above - mentioned bare - hand gesture acquired by the above - mentioned execution subject at the current time. The above - mentioned first model can be a neural network model for extracting the key points of the above - mentioned bare - hand gesture. The above - mentioned input layer can be a network layer that compresses the image to a preset size. The above - mentioned pre - processed image can be an image of the above - mentioned preset size (for example, the preset size is 368×368 pixels).

[0034] In the second step, input the above - mentioned pre - processed image into the above - mentioned extraction layer to obtain key - point information. Among them, the above - mentioned extraction layer can be a network layer for feature extraction. The convolution kernel size of the above - mentioned extraction layer is usually 3x3. The stride of the above - mentioned extraction layer is usually 1. The pooling window of the above - mentioned extraction layer is usually 2x2. The activation function of the above - mentioned extraction layer is usually the Rectified Linear Unit (ReLU). The key - point confidence of the above - mentioned extraction layer can be set to 0.5. The above - mentioned key - point information can be a one - dimensional vector of hand key points. The above - mentioned hand key points can be specific parts corresponding to the bare - hand gesture (for example, wrist, palm center, fingertip).

[0035] In the third step, input the above - mentioned key - point information into the above - mentioned first fully - connected layer to obtain key - point feature information. Among them, the above - mentioned first fully - connected layer can be a network layer for extracting local features of the above - mentioned key - point information. The above - mentioned local features can be detailed features of the above - mentioned hand key points (for example, the detailed features can be finger length, joint angle).

[0036] In the fourth step, input the above - mentioned bare - hand gesture image into the pre - processing layer of the second model to obtain standard image data. Among them, the above - mentioned second model includes: pre - processing layer, convolutional layer, pooling layer, and second fully - connected layer. The above - mentioned pre - processing layer can adjust the above - mentioned bare - hand gesture image to a preset size or perform normalization processing on the image data. The above - mentioned normalization processing can normalize the pixel values from [0, 255] to [0, 1]. The above - mentioned standard image data can be the data after normalizing the above - mentioned bare - hand switching gesture image.

[0037] Step 5: Input the above standard image data into the above convolutional layer to obtain convolutional feature information. Among them, the kernel size of the above convolutional layer is usually 3x3 or 5x5. The number of kernels in the above convolutional layer is usually 32 or 64. The stride of the above convolutional layer is usually 1 or 2. The activation function of the above convolutional layer is usually the ReLU function. In practice, the above convolutional layer may include a first convolutional layer and a second convolutional layer. The above first convolutional layer can be used to extract the edge feature information of the image. The above edge feature information may be the regional feature information where the brightness or color in the image changes significantly. The above second convolutional layer is used to extract the shape feature information of the image. The above shape feature information may be the semantic content of the shape features of the objects in the image.

[0038] Step 6: Input the above convolutional feature information into the above pooling layer to obtain pooling feature information. Among them, the pooling window size of the above pooling layer is usually 2x2 or 3x3. The stride of the above pooling layer is usually the same as the pooling window size. In practice, the above execution entity can reduce the dimension of the above feature vector through the above pooling layer. The above pooling feature information may be the feature vector after dimensionality reduction.

[0039] Step 7: Input the above pooling feature information into the above second fully connected layer to obtain global feature information. Among them, the above second fully connected layer can be used to integrate the above feature vector after dimensionality reduction into global feature information. The above global feature information may be a set of the above local feature information. In practice, the above execution entity can introduce non-linearity through the activation function in the above fully connected layer (for example, the activation function can be the ReLU function) to integrate the local features into global feature information.

[0040] Step 8: Fuse the above key point feature information and the above global feature information to obtain fused feature information. In practice, the above fusion can be a weighted sum of the above key point feature information and the above global feature information.

[0041] Step 9: Input the above fused feature information into the above classifier to determine the target bare - hand gesture. Among them, the above classifier usually uses the softmax function to obtain the probability of each bare - hand gesture. The above determination of the target bare - hand gesture may be to determine the bare - hand gesture corresponding to the maximum value of the above probability.

[0042] Step 10: In response to determining that the above target bare - hand gesture is the above bare - hand switching gesture, determine that the bare - hand switching gesture is detected.

[0043] The above first step to the tenth step are an inventive point of the embodiments of the present disclosure, which solve the technical problem of "low recognition accuracy and insufficient recognition ability for different bare - hand gestures in the existing gesture recognition technology, resulting in difficult user operation". Based on this technical problem, the detection method of the present disclosure extracts key - point information through the extraction layer, extracts edge features and shape features through the convolutional layer, reduces the dimension through the pooling layer, integrates global feature information and feature - fusion information through the fully - connected layer, and finally makes a classification decision through the classifier, realizing the efficient recognition of bare - hand gestures. Thus, it can efficiently detect the bare - hand gestures of the target user, facilitate user operation, and further improve the user experience.

[0044] In some optional implementation manners of some embodiments, the above - mentioned initializing the cone model in response to first detecting the bare - hand switching gesture corresponding to the target user to obtain a manipulable cone model and an auxiliary cone model may include:

[0045] In response to first detecting the bare - hand switching gesture corresponding to the target user, the execution subject switches the current state from the standby state to the initial construction state, and initializes the cone model in the initial construction state to obtain a manipulable cone model and an auxiliary cone model. The above - mentioned current state may be the above - mentioned preset state of the execution subject at the current time. The above - mentioned switching may be that the execution subject changes from one above - mentioned preset state to another above - mentioned preset state.

[0046] Step 102, in response to determining that the initialization of the manipulable cone model is completed and first detecting the bare - hand sliding gesture corresponding to the target user, lock the bottom mapping area corresponding to the manipulable cone model to generate a locked area.

[0047] In some embodiments, in response to determining that the initialization of the manipulable cone model is completed and first detecting the bare - hand sliding gesture corresponding to the target user, the execution subject may lock the bottom mapping area corresponding to the manipulable cone model. Among them, the above - mentioned bottom mapping area may be a circular area with the center point of the bottom of the manipulable cone model as the center and the bottom radius of the manipulable cone model as the radius. The above - mentioned bare - hand sliding gesture may be a gesture for locking the position of the center point of the bottom of the manipulable cone model and controlling the position of the vertex of the manipulable cone model. The above - mentioned completion of initialization may be the end of the initialization of the manipulable cone model. The above - mentioned locked area may be the bottom mapping area corresponding to the manipulable cone model. The above - mentioned locking may be that the execution subject makes the bottom mapping area corresponding to the manipulable cone model fixed.

[0048] In some alternative implementations of some embodiments, locking the bottom mapping area corresponding to the manipulable cone model in response to determining that the initialization of the manipulable cone model is completed and the bare - hand sliding gesture corresponding to the target user is detected for the first time to generate a locked area may include:

[0049] In response to determining that the initialization of the manipulable cone model is completed and the bare - hand sliding gesture corresponding to the target user is detected for the first time, the execution entity switches the current state from the initial construction state to the cone control state, and locks the bottom mapping area corresponding to the manipulable cone model in the cone control state to generate a locked area.

[0050] Step 103, in response to determining that the locking of the bottom mapping area corresponding to the manipulable cone model is completed and continuously detecting the bare - hand sliding gesture corresponding to the target user, according to the movement information corresponding to the bare - hand sliding gesture, correspondingly move the vertex of the manipulable cone model and the vertex of the auxiliary cone model, so that the positions of the projections of the target object within the locked area on the first projection plane and the second projection plane move correspondingly, realizing the removal of the occluder of the target object.

[0051] In some embodiments, in response to determining that the locking of the bottom mapping area corresponding to the manipulable cone model is completed and continuously detecting the bare - hand sliding gesture corresponding to the target user, the execution entity may correspondingly move the vertex of the manipulable cone model and the vertex of the auxiliary cone model according to the movement information corresponding to the bare - hand sliding gesture, so that the positions of the projections of the target object within the locked area on the first projection plane and the second projection plane move correspondingly, realizing the removal of the occluder of the target object. Wherein, the continuous detection may be continuously detecting the bare - hand sliding gesture corresponding to the target user. Further reference Figure 3 , V may be the viewpoint position. m may represent the manipulable cone model. v m may be the termination position of the vertex of m. a may represent the auxiliary cone model. v a may be the termination position of the vertex of a. S 1 may be the starting position of the target object. S 2 may be the position of the target occluder. The target occluder may be an item between the target object and the target user. p m may be the position of the center point at the bottom corresponding to m. Π 1 may be the first projection plane. Π 2 may be the second projection plane. Π 3 may be the motion surface of the dual - bend camera model. c may represent the first projection plane. p c’ can be the projection position of the above-mentioned target object on the above-mentioned c plane, p c can be the projection position of the above-mentioned target occluder on the above-mentioned c plane. p 3 can be the above-mentioned V and the above-mentioned p m where the straight line is located and the above-mentioned Π 2 the position of the intersection point of the plane. o can represent the target object, p o can be the above-mentioned target object on the above-mentioned Π 1 the projection of the plane on the above-mentioned Π 2 the projection position of the plane, S 3 can be the termination position of moving the above-mentioned target object, h can be the above-mentioned p m the vector from to the above-mentioned V. Among them, the first position can be the above-mentioned p c ′ corresponding position. The second position can be p o corresponding position. In practice, the above-mentioned target object is moved through the following steps:

[0052] First step, according to the movement information corresponding to the above-mentioned bare hand sliding gesture and the first formula, determine the movement speed of the vertex of the above-mentioned manipulable cone model. Among them, the above-mentioned movement information can be to determine the gesture start position and the gesture end position. Among them, the above-mentioned gesture start position can be the position where the above-mentioned execution subject first detects the above-mentioned bare hand sliding gesture in the above-mentioned initial construction state, and the gesture end position can be the position where the above-mentioned bare hand sliding gesture is released.

[0053] Second step, in response to obtaining the above-mentioned movement speed, determine the termination position v of the vertex of the above-mentioned manipulable cone model m .

[0054] Third step, according to the above-mentioned v m , determine the projection position p of the above-mentioned target occluder on the above-mentioned first projection plane c .

[0055] Fourth step, according to the above-mentioned p c and the second formula, determine the termination position v of the vertex of the above-mentioned auxiliary cone model a .

[0056] Fifth step, in response to the above-mentioned manipulable cone model projecting the above-mentioned target object onto the Π 1 plane, determine the projection of the above-mentioned target object on the above-mentioned Π 1 plane.

[0057] Sixth step, in response to the above-mentioned auxiliary cone model projecting the projection of the above-mentioned target object on the Π 1 plane onto the Π 2 plane, determine the projection position p of the above-mentioned target object on the Π 2 plane o .

[0058] Step 7, based on the vector between the above-mentioned V and the above-mentioned p o to determine the termination position S of the above-mentioned target object 3 .

[0059] In some alternative implementation manners of some embodiments, the above-mentioned bare - hand sliding gesture corresponding to the above-mentioned target user can be detected through the following steps:

[0060] Step 1, input the bare - hand gesture image into the input layer of the first model to obtain a pre - processed image. Among them, the above - mentioned first model includes: the above - mentioned input layer, an extraction layer, and a first fully - connected layer.

[0061] Step 2, input the above - mentioned pre - processed image into the above - mentioned extraction layer to obtain key - point information.

[0062] Step 3, input the above - mentioned key - point information into the above - mentioned first fully - connected layer to obtain key - point feature information.

[0063] Step 4, input the above - mentioned bare - hand gesture image into the pre - processing layer of the second model to obtain standard image data. Among them, the above - mentioned second model includes: a pre - processing layer, a convolutional layer, a pooling layer, and a second fully - connected layer.

[0064] Step 5, input the above - mentioned standard image data into the above - mentioned convolutional layer to obtain convolutional feature information.

[0065] Step 6, input the above - mentioned convolutional feature information into the above - mentioned pooling layer to obtain pooling feature information.

[0066] Step 7, input the above - mentioned pooling feature information into the above - mentioned second fully - connected layer to obtain global feature information.

[0067] Step 8, fuse the above - mentioned key - point feature information and the above - mentioned global feature information to obtain fused feature information.

[0068] Step 9, input the above - mentioned fused feature information into a classifier to determine the target bare - hand gesture.

[0069] Step 10, in response to determining that the above - mentioned target bare - hand gesture is the above - mentioned bare - hand sliding gesture, determine that the bare - hand sliding gesture is detected.

[0070] In some alternative implementation manners of some embodiments, the above - mentioned execution subject moves the vertices of the above - mentioned manipulable cone model and the vertices of the above - mentioned auxiliary cone model correspondingly according to the movement information corresponding to the above - mentioned bare - hand sliding gesture, which may include the following steps:

[0071] First step, determine the gesture start position and gesture end position in the above-mentioned movement information. Among them, the above-mentioned gesture start position can be the position where the above-mentioned execution entity first detects the above-mentioned bare - hand sliding gesture in the above-mentioned initial construction state, and the gesture end position can be the position where the above-mentioned bare - hand sliding gesture is released.

[0072] Second step, determine the first vector between the above - mentioned gesture start position and the above - mentioned gesture end position. Among them, the above - mentioned first vector can be the vector from the above - mentioned gesture start position to the above - mentioned gesture end position.

[0073] Third step, in response to the vector modulus corresponding to the above - mentioned first vector being less than or equal to the target value, set the movement speed corresponding to the vertex of the above - mentioned manipulable cone model to the value 0. Among them, the above - mentioned target value can be a preset value (for example, the target value is 0.1).

[0074] Fourth step, in response to the modulus of the above - mentioned first vector being greater than the above - mentioned target value, divide the above - mentioned first vector by the above - mentioned vector modulus to obtain a divided vector. Among them, the above - mentioned divided vector can be the above - mentioned first vector divided by the above - mentioned vector modulus.

[0075] Fifth step, multiply the above - mentioned divided vector by a predetermined constant to obtain a multiplied vector. Among them, the above - mentioned predetermined constant can be the value set by the above - mentioned target user (for example, the predetermined constant is 0.18).

[0076] Sixth step, use the above - mentioned multiplied vector as the velocity vector corresponding to the vertex of the above - mentioned manipulable cone model, and perform corresponding movement on the vertex of the above - mentioned manipulable cone model. Among them, the direction of the above - mentioned velocity vector is the same as the direction of the above - mentioned first vector.

[0077] As an example, the above - mentioned bare - hand sliding gesture can perform corresponding movement on the vertex of the above - mentioned manipulable cone model through the following first formula:

[0078] d = P e -P i

[0079]

[0080] Among them, i can represent the start. P i can be the above - mentioned gesture start position in the start i state. e can represent the end. P e can be the above - mentioned gesture end position in the end e state. d can be the vector from the above - mentioned gesture start position to the above - mentioned end position. r can be the above - mentioned target value. α can be the value set by the above - mentioned target user. can be the velocity vector corresponding to the vertex of the above - mentioned manipulable cone model.

[0081] In some alternative implementations of some embodiments, based on the movement information corresponding to the bare - hand sliding gesture, the execution entity correspondingly moves the vertex of the manipulable cone model and the vertex of the auxiliary cone model, and may further include the following steps:

[0082] First step, determine the first position and the second position.

[0083] Second step, determine the vector between the first position and the second position to obtain a second vector. Wherein, the second vector may be the vector from the first position to the second position.

[0084] Third step, determine the vector between the second position and the position of the center point of the bottom corresponding to the manipulable cone model to obtain a third vector. Wherein, the third vector may be the vector from the second position to the position of the center point of the bottom corresponding to the manipulable cone model.

[0085] Fourth step, determine the distance length from the position of the center point of the bottom corresponding to the manipulable cone model to the viewpoint position. Wherein the distance length may be the distance between the position of the center point of the bottom corresponding to the manipulable cone model and the viewpoint position.

[0086] Fifth step, perform a square operation on the second vector to obtain a first coefficient. Wherein, the first coefficient may be a non - negative number.

[0087] Sixth step, multiply the second vector by the third vector to obtain a second coefficient. Wherein, the second coefficient may be a real number.

[0088] Seventh step, subtract the square of the distance length from the square of the third vector to obtain a third coefficient. Wherein, the third coefficient may be a real number.

[0089] Eighth step, input the first coefficient, the second coefficient, and the third coefficient into a preset quadratic equation to obtain the solution of the quadratic equation. Wherein, the quadratic equation may be used to represent that the line where p 3 and p c intersects the sphere. The sphere may be a sphere with p 3 as the center of the sphere and the distance between V and p m as the radius.

[0090] Ninth step, multiply the solution of the quadratic equation by the second vector to obtain a fourth vector. Wherein, the fourth vector may be any vector from p 3 to v a

[0091] ​In the tenth step, add the above fourth vector to the above second position to obtain the moved termination position of the vertex of the above auxiliary cone model.

[0092] The above first step to tenth step are an inventive point of the embodiment of the present disclosure, which solves the technical problem of "the existing technology relies on physical devices for user operations, resulting in difficult operations for users in a virtual environment". Based on this technical problem, the moving method of the present disclosure calculates the vertex position of the above auxiliary cone model by determining the projection position of the above target occluder. Furthermore, it realizes the efficient movement of the occluded target object, facilitates user operations, and thus improves the user experience.

[0093] As an example, the corresponding movement position of the vertex of the above auxiliary cone model can be determined by the following second formula:

[0094] v = p c -p 3

[0095] A = v·v

[0096] B = 2v·(p 3 -p m )

[0097] C = (p 3 -p m )·(p 3 -p m ) - h 2

[0098] h = p m -V

[0099] t = QES(A, B, C)

[0100] v a = p 3 + tv

[0101] Wherein, V may be the above viewpoint position. m may represent the above manipulable cone model. p m may be the position of the bottom center point corresponding to the above m. c may represent the above first projection plane. p c may be the projection position of the above target occluder on the above c plane. p 3 may be the position of the intersection point of the straight line where the above V and the above p m are located and the above second projection plane. V may be the vector from the above p c to the above p 3 h may be the above p mVector to the above V. QES(A, B, C) can be a quadratic equation solver. The above solver can be a function of a software tool (e.g., the software tool is MATLAB). The above quadratic equation can be t 2 (v·v)+t(2v·(p 3 -p m ))+((p 3 -p m )·(p 3 -p m )-h 2 )=0. The above A, B, C can be the coefficients of the above quadratic equation. t can be the solution of the above quadratic equation. a can represent the above auxiliary cone model. v a can be the position of the vertex of the above a.

[0102] In some alternative implementations of some embodiments, the occluder of the above target object is removed through the following steps:

[0103] First step, determine the distance between the first position of the projection of the above target object on the above first projection plane and the starting position of the above target object, to obtain a first length. Among them, the above first projection plane can be the Π 1 plane. The above first position can be the p c ' position. The starting position of the above target object can be the S 1 position. The first length can be the distance between the above p c ' and the above S 1 .

[0104] Second step, determine the distance between the above first position and the second position of the projection of the above first position on the above second projection plane, to obtain a second length. Among them, the above second projection plane can be the Π 2 plane. The above second position can be the p o position. The second length can be the distance between the above p c ' and the above p o .

[0105] Third step, determine the termination position of the above target object according to the above first length, the above second length, and the auxiliary vector between the viewpoint position of the above target user and the above second position.

[0106] As an example, the termination position of the above target object can be determined through the following steps:

[0107] Sub-step 1: Determine the vector length from the viewpoint position to the termination position of the target object according to the above first length, the above second length, and the length corresponding to the above auxiliary vector. Wherein, the above auxiliary vector can be the vector from the above viewpoint position to the above second position. The above vector length can be the sum of the above first length, the above second length, and the length corresponding to the above auxiliary vector.

[0108] Sub-step 2: Determine the vector angle from the viewpoint position to the termination position of the target object according to the above auxiliary vector.

[0109] Sub-step 3: Determine the termination position of the target object according to the coordinates of the above viewpoint position, the above vector length, and the above vector angle, using the Pythagorean theorem.

[0110] Fourth step: In response to the above termination position satisfying the preset condition, determine that the occlusion removal of the above target object is completed. The above preset condition can be that the above execution entity does not detect the above bare hand sliding gesture of the above target user within a preset time period. The above preset time period can be a time period preset by relevant technical personnel (for example, the above preset time period is 1 second). The above occlusion removal can be that the above occlusion no longer occludes the above target object.

[0111] Figure 4 It is a gesture state transition diagram of some embodiments of an object movement method based on bare hand gestures according to the present disclosure.

[0112] In Figure 4In the application scenario, five states are preset for the above-mentioned maneuverable cone model: standby state (MCBOSOFF), initial construction state (Initial Status), cone control state (Cone Control), intermediate state (InterStatus), and selection state (Selection Process). Among them, the above-mentioned standby state can be that the above-mentioned execution entity is in the standby state, and at this time, the above-mentioned target user does not perform any of the above-mentioned barehand gesture operations. The above-mentioned initial construction state can be that the initialization of the cone model is completed, and at this time, the above-mentioned target user determines the position of the bottom center point corresponding to the above-mentioned maneuverable cone model through the above-mentioned barehand sliding gesture. The above-mentioned cone control state can be that the user is allowed to manipulate the vertex of the above-mentioned maneuverable cone model, and at this time, the target user can move the above-mentioned target object in the above-mentioned locked area through the above-mentioned barehand sliding gesture. The above-mentioned intermediate state can be between the above-mentioned cone control state and the above-mentioned selection state, and at this time, the target user can observe the effect after the above-mentioned target object moves. The above-mentioned selection state can be the stage where the above-mentioned target user performs the final selection of the above-mentioned target object. Through the above-mentioned barehand gestures, the conversion between the above five states can be achieved. Specifically, in the above-mentioned standby state, the above-mentioned target user can use the above-mentioned barehand switching gesture to switch between the above-mentioned standby state and the above-mentioned initial construction state; in the above-mentioned initial construction state, the above-mentioned target user can use the above-mentioned barehand switching gesture to return to the above-mentioned standby state, or use the above-mentioned sliding gesture to enter the above-mentioned cone control state, or directly jump to the above-mentioned selection state using the barehand selection gesture. In the above-mentioned cone control state, the above-mentioned target user can return to the above-mentioned standby state through the above-mentioned barehand switching gesture, transition to the above-mentioned intermediate state by releasing the above-mentioned barehand sliding gesture, or enter the above-mentioned selection state through the above-mentioned barehand selection gesture. In the above-mentioned intermediate state, the above-mentioned target user can continue to use the above-mentioned barehand sliding gesture to return to the above-mentioned cone control state or use the above-mentioned barehand selection gesture to enter the above-mentioned selection state. Finally, in the above-mentioned selection state, after the above-mentioned target user completes the selection, he exits by releasing the above-mentioned barehand selection gesture to return to the above-mentioned initial construction state. In some embodiments, the above-mentioned barehand selection gesture can be a gesture used to complete the final selection interaction process. In practice, the posture of the hand corresponding to the above-mentioned barehand selection gesture can be preset by those skilled in the relevant art. For example, the posture of the hand corresponding to the above-mentioned barehand selection gesture is the pinching of the thumb and middle finger of the above-mentioned target user.

[0113] Further referring to Figure 5 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an object moving device based on barehand gestures. These device embodiments correspond to Figure 1 the method embodiments shown, and the above devices can be specifically applied to various electronic devices.

[0114] AsFigure 5 As shown in the figure, an object moving device 500 based on bare - hand gestures includes: a model initialization unit 501, a locking unit 502, and a moving unit 503. Among them, the model initialization unit 501 is configured to, in response to first detecting a bare - hand switching gesture corresponding to a target user, perform model initialization on a cone model to obtain a manipulable cone model and an auxiliary cone model. Among them, the above - mentioned manipulable cone model and the above - mentioned auxiliary cone model are constructed based on a double - bend camera model; the locking unit 502 is configured to, in response to determining that the above - mentioned manipulable cone model is initialized and first detecting a bare - hand sliding gesture corresponding to the above - mentioned target user, lock the bottom mapping area corresponding to the above - mentioned manipulable cone model to generate a locked area. Among them, the above - mentioned locked area is determined based on the position of the bare - hand sliding gesture corresponding to the above - mentioned target user at the current time; the moving unit 503 is configured to, in response to determining that the locking of the bottom mapping area corresponding to the above - mentioned manipulable cone model is completed and continuously detecting a bare - hand sliding gesture corresponding to the above - mentioned target user, perform corresponding movement on the vertex of the above - mentioned manipulable cone model and the vertex of the above - mentioned auxiliary cone model according to the movement information corresponding to the above - mentioned bare - hand sliding gesture, so that the positions of the projections of the target object in the locked area on the first projection plane and the second projection plane move correspondingly, realizing the movement of the above - mentioned target object.

[0115] It can be understood that the various units described in the above - mentioned device 500 correspond to the respective steps in the method described with reference to Figure 1 Therefore, the operations, features, and beneficial effects described above for the method also apply to the above - mentioned device 500 and the units included therein, and will not be elaborated here.

[0116] Next, with reference to Figure 6 , which shows a schematic structural diagram of a device 600 suitable for implementing some embodiments of the present disclosure. Figure 6 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0117] As Figure 6 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in the read - only memory 602 or a program loaded from the storage device 608 into the random - access memory 603. In the random - access memory 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the read - only memory 602, and the random - access memory 603 are connected to each other through a bus 604. The input / output interface 605 is also connected to the bus 604.

[0118] Typically, the following devices can be connected to the input / output interface 605: input devices 606 including, for example, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, head-mounted displays (HMDs), speakers, vibrators, etc.; storage devices 608 including, for example, hard disks; and communication devices 609. The communication device 609 can allow the electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had. Figure 6 Each block shown in

[0119] can represent one device or, as needed, multiple devices.

[0120] It should be noted that, in some embodiments of the present disclosure, the above-mentioned computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0121] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0122] The above computer-readable medium may be included in the above electronic device; or may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: in response to first detecting a bare-hand switching gesture corresponding to a target user, perform model initialization on a cone model to obtain a manipulable cone model and an auxiliary cone model. Wherein, the above manipulable cone model and the above auxiliary cone model are constructed based on a dual-bend camera model; in response to determining that the above manipulable cone model has completed initialization and first detecting a bare-hand sliding gesture corresponding to the above target user, lock a bottom mapping area corresponding to the above manipulable cone model to generate a locked area. Wherein, the above locked area is determined based on the position of the bare-hand sliding gesture corresponding to the above target user at the current time; in response to determining that the locking of the bottom mapping area corresponding to the above manipulable cone model is completed and continuously detecting a bare-hand sliding gesture corresponding to the above target user, move the vertices of the above manipulable cone model and the vertices of the above auxiliary cone model correspondingly according to the movement information corresponding to the above bare-hand sliding gesture, so that the positions of the projections of the target object in the locked area on the first projection plane and the second projection plane move correspondingly, realizing the removal of the occluder of the above target object.

[0123] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0125] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a model initialization unit, a locking unit, and a moving unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the model initialization unit can also be described as "in response to first detecting a bare hand switching gesture corresponding to a target user, initializing a cone model to obtain a manipulable cone model and an auxiliary cone model".

[0126] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0127] Some embodiments of the present disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any of the above-mentioned object moving methods based on bare hand gestures.

[0128] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A method for moving an object based on bare-hand gestures, comprising: In response to detecting a bare-hand switching gesture corresponding to a target user for the first time, initializing a cone model to obtain a controllable cone model and an auxiliary cone model, wherein the controllable cone model and the auxiliary cone model are constructed based on a dual-bend camera model; In response to determining that the manipulable pyramid model has completed initialization and a bare-hand sliding gesture corresponding to the target user is detected for the first time, locking a bottom mapping area corresponding to the manipulable pyramid model to generate a locked area, wherein the locked area is determined based on a position of the bare-hand sliding gesture corresponding to the target user at a current time; In response to determining that the bottom mapping area corresponding to the manipulable cone model is locked and the bare-hand sliding gesture corresponding to the target user is continuously detected, the vertices of the manipulable cone model and the auxiliary cone model are moved accordingly based on the movement information corresponding to the bare-hand sliding gesture, so that the projection positions of the target object in the locked area on the first projection plane and the second projection plane are moved accordingly, thereby realizing the removal of obstructions to the target object.

2. The method according to claim 1, wherein: The step of locking the bottom mapping area corresponding to the manipulable cone model to generate a locking area includes: Determine the position information of the bottom center point corresponding to the controllable cone model; The bottom mapping area is locked according to the position information, wherein the bottom mapping area is a circular area with the bottom center point corresponding to the controllable cone model as the center and the bottom radius corresponding to the controllable cone model as the radius.

3. The method according to claim 1, wherein: The occluders of the target object are removed by the following steps: Determine a distance between a first position of a projection of the target object on the first projection plane and a starting position of the target object to obtain a first length; Determine the distance between the first position and a second position of a projection of the first position on the second projection plane to obtain a second length; Determine the end position of the target object according to the first length, the second length, and an auxiliary vector between the viewpoint position of the target user and the second position; In response to the termination position satisfying a preset condition, it is determined that the removal of the obstruction of the target object is completed.

4. The method according to claim 1, wherein: In response to detecting the bare-hand switching gesture corresponding to the target user for the first time, initializing the cone model to obtain a manipulable cone model and an auxiliary cone model, including: In response to detecting the bare-hand switching gesture corresponding to the target user for the first time, switching the current state from the standby state to the initial construction state, and initializing the cone model in the initial construction state to obtain a manipulable cone model and an auxiliary cone model; And in response to determining that the manipulable cone model has completed initialization and detected the bare-hand sliding gesture corresponding to the target user for the first time, locking the bottom mapping area corresponding to the manipulable cone model to generate a locked area, comprising: In response to determining that the manipulable cone model has completed initialization and detected the bare-hand sliding gesture corresponding to the target user for the first time, the current state is switched from the initial construction state to the cone control state, and the bottom mapping area corresponding to the manipulable cone model is locked in the cone control state to generate a locked area.

5. The method according to claim 4, further comprising: In response to determining that there is no blocking object in front of the target object and detecting the bare-hand switching gesture of the target user again, the current state is switched from the cone control state to the standby state.

6. The method according to claim 1, wherein: The method further comprises: Inputting the bare hand gesture image into the input layer of the first model to obtain a preprocessed image, wherein the first model includes: the input layer, the extraction layer and the first fully connected layer; Inputting the preprocessed image into the extraction layer to obtain key point information; Inputting the key point information into the first fully connected layer to obtain key point feature information; Inputting the bare-hand gesture image into a preprocessing layer of a second model to obtain standard image data, wherein the second model comprises: the preprocessing layer, a convolutional layer, a pooling layer, and a second fully connected layer; Inputting the standard image data into the convolution layer to obtain convolution feature information; Inputting the convolution feature information into the pooling layer to obtain pooling feature information; Inputting the pooled feature information into the second fully connected layer to obtain global feature information; Fusing the key point feature information with the global feature information to obtain fused feature information; The fused feature information is input into a classifier to determine the target bare hand gesture.

7. The method according to claim 1, wherein: The correspondingly moving the vertices of the manipulable cone model and the vertices of the auxiliary cone model according to the movement information corresponding to the bare-hand sliding gesture includes: Determine a gesture start position and a gesture end position in the movement information; Determine a first vector between the gesture start position and the gesture end position; In response to the vector modulus corresponding to the first vector being less than or equal to a target value, setting the movement speed corresponding to the vertex of the controllable cone model to a value of 0; In response to the first vector modulus being greater than the target value, dividing the first vector by the vector modulus to obtain a divided vector; Multiplying the divided vector by a predetermined constant to obtain a multiplied vector; The multiplied vector is used as a velocity vector corresponding to the vertex of the controllable cone model, and the vertex of the controllable cone model is moved accordingly.

8. An object moving device based on bare hand gestures, comprising: A model initialization unit is configured to initialize the cone model in response to detecting a bare-hand switching gesture corresponding to a target user for the first time, so as to obtain a manipulable cone model and an auxiliary cone model, wherein the manipulable cone model and the auxiliary cone model are constructed based on a dual-bend camera model; a locking unit, configured to, in response to determining that the manipulable pyramid model has completed initialization and a bare-hand sliding gesture corresponding to the target user is detected for the first time, lock a bottom mapping area corresponding to the manipulable pyramid model to generate a locked area, wherein the locked area is determined based on a bare-hand sliding gesture position corresponding to the target user at a current time; The moving unit is configured to, in response to determining that the bottom mapping area corresponding to the manipulable cone model is locked and the bare-hand sliding gesture corresponding to the target user is continuously detected, correspondingly move the vertices of the manipulable cone model and the vertices of the auxiliary cone model according to the movement information corresponding to the bare-hand sliding gesture, so that the projection positions of the target object in the locked area on the first projection plane and the second projection plane are correspondingly moved, thereby realizing the removal of the obstruction of the target object.

9. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.