Control method, model training method, apparatus, device, and storage medium
Through the lightweight gesture recognition network, the key point position and gesture category are directly output in the head-mounted display device, which solves the problem of poor interaction experience caused by insufficient computing power and achieves efficient human-computer interaction.
Patent Information
- Application Number
- PCT/CN2024/136567
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-12-03
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, in head-mounted display devices with low computing power, the gesture recognition computing process is slow, resulting in poor user interaction experience.
A lightweight gesture recognition network is used to detect and classify key points. By extracting hand feature vectors, the key point positions and gesture categories are directly output, reducing dependence on key point detection and reducing the amount of calculation.
It improves the user interaction experience of head-mounted display devices with low computing power, enriches human-computer interaction methods, and is suitable for devices such as MR and AR glasses.
Smart Images

Figure CN2024136567_03072025_PF_FP_ABST
Abstract
Description
Control method, model training method, device, equipment and storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 25, 2023, with application number 2023118057446, and invention name “Control method, model training method, device, equipment and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the technical field of head-mounted display devices, and in particular to a control method, a model training method, an apparatus, a device, and a storage medium. Background Art
[0003] Gesture recognition uses sensors, computer vision, and machine learning to capture and interpret human posture, gestures, and spatial position information, enabling natural interaction with computer systems. Applying gesture recognition technology to head-mounted display devices can provide users with a more intuitive, natural, and immersive interactive experience. Based on gesture recognition algorithms, head-mounted display devices can identify user gestures, present virtual content, and display relevant information based on received user actions.
[0004] The relevant technology requires gesture classification based on the calculation results of the network that detects key points, which reduces the speed of the entire calculation process and requires higher computing power. Therefore, it is not suitable for head-mounted display devices with lower computing power, and the user's interactive experience with the head-mounted display device is poor.
[0005] Application Contents
[0006] The main purpose of this application is to provide a control method, model training method, device, equipment and storage medium, which are suitable for head-mounted display devices with low computing power, and improve the user's interactive experience with the head-mounted display device.
[0007] In a first aspect, the present application provides a method for controlling a head-mounted display device, comprising:
[0008] Acquire an image to be recognized, where the image to be recognized includes at least a portion of a hand;
[0009] Extracting a hand feature vector of the image to be recognized;
[0010] Based on the gesture recognition network, key point detection is performed on the hand feature vector to obtain key point positions of a plurality of preset hand key points, and the gesture category corresponding to the image to be recognized is determined according to the hand feature vector;
[0011] According to the key point position of at least one of the preset hand key points and the gesture category, the head mounted display device is controlled to perform a preset task.
[0012] In a second aspect, the present application further provides a method for training a control model of a head-mounted display device, wherein the control model includes a gesture recognition network, and the training method includes:
[0013] Acquire training data, the training data including a plurality of images to be recognized and target recognition results corresponding to the images to be recognized, the target recognition results including target key point positions and target gesture categories;
[0014] Extracting a hand feature vector of the image to be recognized;
[0015] Based on the gesture recognition network, key point detection is performed on the hand feature vector to obtain current key point positions of a plurality of preset hand key points, and a current gesture category corresponding to the image to be recognized is determined based on the hand feature vector;
[0016] The model parameters of the control model are adjusted according to a deviation between the target recognition result and a current recognition result, wherein the current recognition result includes the current key point position and the current gesture category.
[0017] In a third aspect, the present application further provides a control device for a head-mounted display device, comprising:
[0018] an acquisition module, configured to acquire an image to be recognized, wherein the image to be recognized includes at least a portion of a hand;
[0019] An extraction module, configured to extract a hand feature vector of the image to be identified;
[0020] a recognition module configured to perform key point detection on the hand feature vector based on a gesture recognition network to obtain key point positions of a plurality of preset hand key points, and determine a gesture category corresponding to the image to be recognized based on the hand feature vector;
[0021] A control module is used to control the head-mounted display device to perform a preset task according to the key point position of at least one preset hand key point and the gesture category.
[0022] In a fourth aspect, the present application further provides a computer device, the computer device comprising a memory and a processor;
[0023] Memory for storing computer programs;
[0024] A processor is used to execute a computer program and implement the control method of the head-mounted display device as described above, or the training method of the control model of the head-mounted display device as described above when executing the computer program.
[0025] In a fifth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the control method of the head-mounted display device as described above, or the steps of the training method of the control model of the head-mounted display device as described above are implemented.
[0026] The present application provides a control method, a model training method, an apparatus, a device, and a storage medium, wherein the control method comprises: obtaining an image to be recognized, the image to be recognized including at least a portion of a hand; extracting a hand feature vector from the image to be recognized; performing key point detection on the hand feature vector based on a gesture recognition network to obtain key point positions of several preset hand key points, and determining a gesture category corresponding to the image to be recognized based on the hand feature vector; and controlling a head-mounted display device to perform a preset task based on the key point position and gesture category of at least one of the preset hand key points. The present application performs key point detection on the image to be recognized and classifies gestures in the image to be recognized based on a gesture recognition network. The key point positions and gesture categories are output simultaneously through the same network. Since gesture classification does not rely on the results of key point detection, the computational effort is reduced and the computing power requirement is low, making it suitable for head-mounted display devices with low computing power. The head-mounted display device is controlled to perform tasks based on the key point positions and gesture categories output by the network, so that the head-mounted display device can perform human-computer interaction tasks based on the recognized user gestures, enriching the human-computer interaction methods and improving the user's interactive experience with the head-mounted display device. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0028] FIG1 is a flow chart of a control method provided in an embodiment of the present application;
[0029] FIG2 is a schematic diagram showing the connection between the server and the smart glasses provided in an embodiment of the present application;
[0030] FIG3 is a schematic diagram of gesture classification provided in an embodiment of the present application;
[0031] FIG4 is a schematic diagram of key points of a hand provided in an embodiment of the present application;
[0032] FIG5 is a schematic diagram of a palm gesture and a pointing gesture provided in an embodiment of the present application;
[0033] FIG6 is a schematic diagram of preset content displayed on the palm of a hand according to an embodiment of the present application;
[0034] FIG7 is a flow chart of a control model training method provided in an embodiment of the present application;
[0035] FIG8 is a schematic block diagram of a control device provided in an embodiment of the present application;
[0036] FIG9 is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0037] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0038] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.
[0039] The related technology requires the use of two networks for gesture recognition, one network detects key points, and the other network needs to classify gestures based on the key point results detected by the previous network to obtain gesture recognition results. That is, in the related technology, the gesture classification network needs to classify gestures based on the calculation results of the network that detects key points. However, the two network loading operations will slow down the entire calculation process and require greater computing power. Therefore, it is not suitable for head-mounted display devices with lower computing power, and the user's interactive experience with the head-mounted display device is poor.
[0040] In view of this, the embodiments of the present application provide a control method, a model training method, an apparatus, a device, and a storage medium. Among them, the control method can be applied to a head-mounted display device. It should be noted that the head-mounted display device can be MR (Mixed Reality) glasses or AR (Augmented Reality) glasses. In addition, the control method can also be applied to a server, which can be a separate server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0041] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.
[0042] Please refer to Figure 1, which is a flow chart of a control method provided in an embodiment of the present application. It should be noted that the control method provided in an embodiment of the present application can be used in a head-mounted display device, and of course can also be used in a server.
[0043] As shown in FIG2 , the control method is applied to a server, the server and the terminal are in communication connection, and the server can send the preset task obtained according to the control method to the terminal.
[0044] In a specific implementation, the terminal includes a head-mounted display device, which can be MR glasses or AR glasses; the server can be a separate server, a server cluster, or a cloud server that provides cloud computing services.
[0045] As shown in FIG1 , the control method includes steps S101 to S105 .
[0046] Step S101: Acquire an image to be recognized, where the image to be recognized includes at least part of a hand.
[0047] The image to be identified may include hands with different gestures, or may include hands with different parts. For example, the image to be identified may include only the palm, the back of the hand, or fingers making different shapes. In addition, the embodiment of the present application may perform pre-processing operations such as image enhancement and image resizing on the image to be identified to meet the needs of image processing and analysis of head-mounted display devices with low computing power requirements. Specifically, the image size is generally 224*224. The embodiment of the present application may reduce the image size to effectively reduce the amount of calculation and running time. Of course, the specific reduction ratio may be determined according to the actual situation and is not specifically limited here.
[0048] Step S102: extracting the hand feature vector of the image to be recognized.
[0049] It should be noted that during the neural network processing, the image to be identified, which includes at least a portion of a hand, can first be input into a preset feature extractor to extract hand feature vectors from the image to be identified, for example, extracting palm feature vectors, feature vectors of multiple fingers, etc. The preset feature extractor in the embodiment of the present application can adopt a convolutional neural network.
[0050] In some embodiments, a convolutional neural network may include an input layer, a convolution layer, an activation layer, a pooling layer, and a fully connected layer. Specifically, the image to be identified is input into the input layer of the convolutional neural network; each neuron of the convolution layer is connected to a local area of the input layer. A convolution layer may have multiple different convolution kernels, each of which slides on the input image and processes only a small piece of the image at a time to extract the most basic features of the image to be identified. By processing the image to be identified using the convolution kernel, the feature information of the local area in the image to be identified can be extracted; the activation layer performs a nonlinear mapping on the output result of the convolution layer; the pooling layer compresses the input feature map to extract the main features, which can reduce the dimension of the data, thereby making the feature extractor more efficient; the fully connected layer is at the end of the convolutional neural network to connect all features to obtain the final output value, i.e., the hand feature vector.
[0051] In the preset feature extractor of the embodiment of the present application, the number of channels in each layer structure is reduced, that is, the number of feature maps output by each layer is reduced. For example, the number of channels in each layer structure can be reduced by a preset number to improve feature extraction efficiency.
[0052] Step S103: Based on the gesture recognition network, key point detection is performed on the hand feature vector to obtain the key point locations of several preset hand key points, and the gesture category corresponding to the image to be recognized is determined based on the hand feature vector. Furthermore, the gesture recognition network in the embodiments of the present application can employ a lightweight network suitable for head-mounted display devices with low computing power. For example, lightweight networks can include computer vision models such as MobileNets, ShuffleNet, and GhostNet. It should be noted that the MobileNets computer vision model can be used to handle a variety of training tasks, including face analysis, common object detection, and photo location. ShuffleNet is a neural network structure designed specifically for devices with limited computing resources, significantly reducing computational overhead while maintaining model accuracy. GhostNet is a neural network architecture whose core is a Ghost module that generates more feature maps, thereby ensuring accuracy while reducing computational effort and increasing computational speed. In practical applications, the network with the least computational time can be selected as the gesture recognition network in the embodiments of the present application based on the network's computational time on the head-mounted display device, thereby improving the efficiency of key point detection and gesture classification.
[0053] Specifically, the embodiment of the present application can perform coordinate mapping processing on the extracted hand feature vector through the gesture recognition network, thereby obtaining the key point positions of several preset hand key points. Specifically, the extracted hand feature vector is segmented to obtain several segment feature vectors, and the maximum predicted probability of each segment feature vector in the several segment feature vectors is counted. Then, based on the statistical maximum predicted probability, a heat map is drawn, and the X coordinate information and Y coordinate information corresponding to several preset hand key points are mapped according to the drawn heat map, thereby obtaining the key point positions of the preset hand key points.
[0054] In gesture recognition, key point detection can detect at least one key point of the hand. The number and location of these key points vary depending on the specific application.
[0055] Exemplarily, the preset hand key points may include at least one of the following: wrist key points, palm key points, and finger key points. The wrist key points can serve as reference points for hand position and orientation. Fingers are the most frequently used part of the hand for movement, so detecting finger key points can provide important information about gestures.
[0056] Optionally, the finger key points may include at least one of the following: fingertip key points, knuckle key points, and finger base key points. The palm key points may include at least one of the following: palm key points, metacarpophalangeal joint key points, and key points at the connection between the palm and wrist.
[0057] Key points for detecting fingers include the fingertips, knuckles, and bases. The palm is another important part of the hand, providing information such as its orientation and posture. Key points for detecting the palm include the center and edges. The back of the hand and wrist can also serve as key points for gesture detection, providing information such as hand orientation and tilt.
[0058] As shown in Figure 4, a general hand can include 21 key points. The preset hand key points in the embodiment of the present application can be key points on individual fingers or key points on the palm, that is, the gesture recognition network of the embodiment of the present application can include a key point detection head, based on which the key point detection head can output the key point positions of some of the 21 key points to reduce the amount of calculation.
[0059] It's important to note that people don't necessarily use every part of their hands in their daily work and daily lives. Generally, certain fingers are used more frequently. For example, using the thumb and index finger to touch a screen, editing text with the thumb or index finger, and clicking with the index finger can be helpful. Therefore, outputting a subset of the 21 key points can meet the needs of most scenarios. For example, a key point detection head, such as the one used in the YOLO object detection model, can be used to detect hand feature vectors in the image to be identified. This allows the palm and fingers to be located by detecting key points at specific locations on the palm and fingers. These key points can include fingertip key points, finger joint key points, and palm center key points. By detecting these key points, the position and orientation information of the palm and fingers corresponding to the key points can be obtained.
[0060] As shown in FIG3 , the gesture category may include at least one of the following: call, like, ok, palm, stop, etc.
[0061] In some embodiments, the gesture recognition network in the embodiments of the present application may include a classification head. After extracting the hand feature vector of the image to be recognized, the gesture category of the image to be recognized can be output through the fully connected layer in the classification head without relying on the key point detection results, thereby reducing the overall calculation process and improving classification efficiency.
[0062] In addition, in the prior art, the dimensions of the hand feature vector are generally relatively large. In the gesture recognition network of the embodiment of the present application, the dimension of each fully connected layer can be reduced while ensuring the accuracy of gesture classification to reduce the amount of calculation.
[0063] Step S104: Control the head-mounted display device to perform a preset task according to the key point position and gesture category of at least one preset hand key point.
[0064] Gesture recognition is a technology that identifies user intentions and instructions by analyzing and understanding human movements and gestures, enabling natural interaction between people and computer systems. Head-mounted display devices can include MR glasses and AR glasses. Among them, gesture recognition technology based on MR glasses is an application that combines augmented reality, virtual reality, and gesture recognition. MR glasses merge the real world with the virtual world to create a new environment. Users can see real objects in the real world and virtual objects in the virtual world. The virtual objects are placed in the real world, allowing users to interact with these virtual objects. Gesture recognition technology based on AR glasses is an application that combines augmented reality and gesture recognition. AR glasses can overlay virtual information on the real world, allowing users to interact with virtual information and the real world.
[0065] As shown in Figure 6, the embodiment of the present application can identify the user's intentions and instructions by obtaining gesture categories, such as palm gestures or pointing gestures, and then control the AR glasses to perform preset tasks, such as presenting virtual content, performing operations according to pointing gestures, etc.
[0066] Applying gesture recognition technology to AR glasses can provide users with a more intuitive, natural, and immersive interactive experience. By recognizing the user's gestures, AR glasses can understand the user's intentions and instructions and present virtual content, perform operations, or provide relevant information accordingly.
[0067] The control method provided in the above embodiment includes: obtaining an image to be recognized, the image to be recognized including at least a portion of a hand; extracting a hand feature vector from the image to be recognized; performing key point detection on the hand feature vector based on a gesture recognition network to obtain key point positions of several preset hand key points, and determining a gesture category corresponding to the image to be recognized based on the hand feature vector; and controlling a head-mounted display device to perform a preset task based on the key point position and gesture category of at least one preset hand key point. The present application performs key point detection on the image to be recognized and classifies gestures in the image to be recognized based on a gesture recognition network. The key point positions and gesture categories are simultaneously output through the network. Since gesture classification does not rely on the results of key point detection, the amount of computation is reduced and the computing power requirements are relatively low. Therefore, it is suitable for head-mounted display devices with lower computing power. Based on the key point positions and gesture categories output by the network, the head-mounted display device is controlled to perform tasks, so that the head-mounted display device can perform human-computer interaction tasks based on the recognized user gestures, enriching the human-computer interaction methods and improving the user's interactive experience with the head-mounted display device.
[0068] In an exemplary embodiment, the gesture recognition network includes a key point detection head, and step S103 may specifically include S1030.
[0069] Step S1030: Based on the key point detection head, key point detection is performed on the hand feature vector to obtain key point positions of several preset hand key points. The hand feature vector includes a finger feature vector, wherein the finger feature vector includes a thumb feature vector and / or an index finger feature vector.
[0070] The embodiment of the present application does not require detection of all 21 key points; instead, it can detect key points on the user's thumb and / or index finger, two fingers that are not easily obscured. Users can use their thumb or index finger alone to perform operations such as clicking and sliding. Using both thumb and index finger simultaneously allows for pointing, zooming in, zooming out, and inputting. Therefore, detecting key points on the thumb and / or index finger can identify the user's command.
[0071] Compared to the computational effort of calculating the 21 key points of all fingers, only calculating the two key points of the thumb and index finger can reduce the computational effort. Based on this, the gesture recognition network in the embodiment of the present application includes a key point detection head, which performs key point detection on the hand feature vector in the image to be recognized based on the key point detection head. Specifically, the hand feature vector includes the finger feature vector, and only the thumb feature vector or the index finger feature vector can be extracted, or both the thumb feature vector and the index finger feature vector can be extracted. In actual applications, people mainly rely on the thumb and index finger to complete various activities. Therefore, the embodiment of the present application can perform key point detection on the thumb feature vector and the index finger feature vector to reduce the computational effort.
[0072] In an exemplary embodiment, step S1030 may include step S1031 and / or step S1032.
[0073] Exemplarily, at least part of the hand includes a thumb, and the hand feature vector includes a thumb feature vector.
[0074] Step S1031: Based on the key point detection head, key point detection is performed on the thumb feature vector to obtain the thumb key point position corresponding to the thumb key point.
[0075] Exemplarily, at least part of the hand includes an index finger, and the hand feature vector includes an index finger feature vector.
[0076] Step S1032: Based on the key point detection head, perform key point detection on the index finger feature vector to obtain the index finger key point position corresponding to the index finger key point.
[0077] In some implementations, only key points of the thumb feature vector may be detected, only key points of the index finger feature vector may be detected, or key points of both the thumb and index finger feature vectors may be detected. Specifically, the thumb and index finger are less likely to be obstructed, which improves the stability of key point detection. Using key points of the thumb feature vector and / or the index finger feature vector, it is possible to determine whether the user's hand is performing an input, click, or swipe operation.
[0078] Exemplarily, the key point detection head in the embodiment of the present application can adopt the detection head of the YOLO object detection model, that is, two fully connected layers as the key point detection head, detect the thumb feature vector and the index finger feature vector, and output four numerical values representing the key point positions corresponding to the thumb key point and the index finger key point. Taking the YOLOv5 object detection model as an example, it uses a key point regressor module to predict the position of the key point. After the hand feature vectors such as the thumb feature vector and the index finger feature vector are input into the key point regressor, the coordinate information corresponding to the thumb key point, that is, the thumb key point position, or the coordinate information corresponding to the index finger key point, that is, the index finger key point position, can be obtained.
[0079] In an exemplary embodiment, step S104 includes step S1041A and step S1042A.
[0080] Step S1041A: When the gesture category is a palm gesture, determine the target display area according to the key point positions and the palm gesture, and the boundary of the target display area is determined according to the key point positions.
[0081] Step S1042A: Control the head-mounted display device to display preset content in the target display area.
[0082] In actual applications, text input of AR glasses is a particularly important part of human-computer interaction. The embodiment of the present application can display the virtual keyboard of AR glasses in the palm of one hand of the user. When the AR glasses receive various operations of the user's other hand's fingers on the virtual keyboard, the corresponding content is displayed. For example, the virtual keyboard does not require other media and can control the position of the displayed content, which is convenient for use in crowded and small spaces, thereby realizing human-computer interaction and improving the user's human-computer interaction experience.
[0083] Compared with the floating virtual keyboard, which requires the user to turn his head and use his fingers to tap the buttons floating in the air to select buttons on the keyboard, the embodiment of the present application does not require the user to turn his head to achieve the desired effect, making it easy to control the position of the keyboard, so that the display of virtual content is not restricted by space, and the human-computer interaction is efficient and convenient.
[0084] Compared with the solution of using a mobile phone for cross-screen input, the embodiment of the present application does not require the addition of other media, can control the position of the keyboard, and is more convenient.
[0085] In an exemplary embodiment, the preset hand key points include the index fingertip key point and the thumb fingertip key point, and step S1041A includes steps SA1 and SB1.
[0086] Step SA1: When the gesture category is a palm gesture, determine the boundary of the gesture area according to the palm gesture, the key point position corresponding to the key point of the index fingertip, and the key point position corresponding to the key point of the thumbtip.
[0087] Step SB1: Determine the target display area according to the boundary of the gesture area and the key point position corresponding to the key point of the thumb fingertip.
[0088] When the gesture category is identified as a palm gesture, the boundary of the gesture area that can just cover the palm gesture can be determined based on the detected position of the index fingertip key point, the position of the thumb fingertip key point and the direction of the palm gesture. Then, based on the key point position corresponding to the thumb fingertip key point, the intersection of the two adjacent boundaries of the gesture area closest to the thumb fingertip key point is determined. Then, using the intersection as the reference point, the boundary of the target display area is determined.
[0089] It is understood that the shape of the target display area and the shape of the gesture area can both be rectangular. The intersection of two adjacent boundaries of the target display area can be determined based on the intersection of two adjacent boundaries of the gesture area. Finally, the boundary of the target display area can be determined based on the size of the boundary of the gesture area, thereby obtaining the target display area. In addition, the size of the target display area in the embodiments of the present application can be adapted to the size of the palm of the user.
[0090] The target display area is determined by the gesture area, so that the target display area can be anchored exactly at the center of the user's palm, and the size of the target display area is determined by the size of the gesture area. Therefore, it can be suitable for hands of different sizes, thereby improving the user experience.
[0091] The following is a discussion based on actual application scenarios:
[0092] It is understandable that the target display area can be understood as the anchor position of the input box (virtual keyboard), and the preset content can be text, pictures, or keyboard keys in the input box (virtual keyboard). As shown in Figure 6, if the gesture category is a palm gesture, the boundary of the target display area can be determined based on the key point position to control the head-mounted display device to anchor the AR glasses virtual keyboard interface to the palm position of the palm gesture. As shown in Figure 5, the process of determining the boundary of the target display area is as follows:
[0093] 1. Detect the position of the hand, that is, the rectangular box ABCD in the figure;
[0094] 2. Identify the hand. When the palm gesture is recognized, detect the palm position in the palm gesture, that is, the rectangular frame EF in Figure 5 (target display area), and start the palm input method;
[0095] 3. Anchoring the palm input frame: Based on the identified key point position of the index fingertip, calculate the point closest to the index fingertip among the four corner points of the rectangular frame ABCD, which is point C in Figure 5. Then, select the edge closest to the thumb tip from the edges CD and BC passing through point C, which is CD in Figure 5. Using D as the reference point, the palm position can be estimated. The coordinates of E in the gesture frame ABCD are (1 / 10*CD, 1 / 3*AD), and the coordinates of F are (1 / 2*AB, 4 / 5*AD). The coefficients can be changed according to the actual product application. This is just an example.
[0096] 4. When the other hand appears, detection and recognition are performed. When the indication gesture is recognized, the click operation of the index finger tip can be responded to;
[0097] 5. When the tip of the index finger of the indicating gesture "indicate" touches the key corresponding to the palm position input box (target display area) of the palm gesture "palm", the corresponding information can be displayed.
[0098] In an exemplary embodiment, step S105 includes step S1051B, step S1052B, and step S1053B.
[0099] Step S1051B: When the image to be recognized includes the first hand and the second hand, and the gesture category of the first hand is a palm gesture and the gesture category of the second hand is a pointing gesture, the preset content is displayed at the palm position in the palm gesture.
[0100] Step S1052B: Determine the interactive action for the preset content based on the gesture category of the second hand and the key point position of the second hand.
[0101] Step S1053B: Control the head-mounted display device to perform the corresponding interactive task according to the interactive action.
[0102] In some embodiments, the image to be recognized may include a first hand and a second hand. If the first hand gesture is classified as a palm gesture, the head-mounted display device can be controlled to display preset content of the target display area at the palm position of the palm gesture, i.e., preset content such as text in an input box, an image, or keyboard keys. If the second hand gesture is classified as a pointing gesture, the second hand's interaction with the preset content can be determined based on the key point positions of the detected finger feature vectors of the second hand. For example, when AR glasses or MR glasses detect a user's palm gesture, a keyboard is displayed at the palm position of the palm gesture. At this time, the user's other hand can freely operate on the keyboard. If only the key points of the index fingertip and the key points near the connection between the index finger and thumb are detected, the hand gesture category can be determined as a pointing gesture. Then, based on the position of the detected key points of the index fingertip corresponding to the keyboard (such as the positions of the numbers 1, 0, 6 on the keyboard) or the movement range (continuous movement forming the unlock pattern S), the hand's interaction with the keyboard can be determined.
[0103] As shown in Figure 6, an input box is displayed in the palm of one hand, which includes multiple control buttons. The index finger of the other hand clicks on the input box, and the head-mounted display device performs corresponding interactive tasks under this interactive action, such as expanding, editing, deleting, sliding, and other interactive actions. The head-mounted display device performs corresponding interactive tasks based on interactive actions such as expanding, editing, deleting, sliding, etc., for example, expanding a multi-level directory, displaying an edit bar, deleting the content from the input box, or switching to the next page. The user displays the preset content such as the input box with the palm of one hand, and performs various interactive actions with the other hand. Therefore, the preset content of the target display area can be controlled to be displayed at any position, making the position and operating space of human-computer interaction more flexible, thereby improving the convenience and user experience of human-computer interaction.
[0104] Please refer to FIG. 7 in conjunction with the aforementioned embodiment. FIG. 7 is a schematic diagram of a control model training method provided in an embodiment of the present application.
[0105] This control model training method can be applied to a head-mounted display device, but is not limited thereto. For example, it can be applied to a computer or server. If the control model training method is applied to a server, the control model obtained by the training method can be sent to the head-mounted display device via the server. This is not a limitation and is not intended to be construed herein.
[0106] The control model includes a gesture recognition network, and the training method includes steps S201 to S206.
[0107] Step S201: Acquire training data, where the training data includes a plurality of images to be recognized and target recognition results corresponding to the images to be recognized, where the target recognition results include target key point positions and target gesture categories.
[0108] Step S202: extracting the hand feature vector of the image to be recognized.
[0109] Step S203: Based on the gesture recognition network, key point detection is performed on the hand feature vector to obtain the current key point positions of several preset hand key points, and the current gesture category corresponding to the image to be recognized is determined according to the hand feature vector.
[0110] Step S204: adjusting the model parameters of the control model according to the deviation between the target recognition result and the current recognition result, where the current recognition result includes the current key point position and the current gesture category.
[0111] Model parameters refer to the estimation ability of the control model, and the model parameters of the control model are very important to the accuracy of its estimation. In the process of training a control model based on a neural learning network, it is generally necessary to continuously adjust the model parameters so that the current recognition result estimated by the control model is the same as the target recognition result, that is, the current key point position is the same as the target key point position, and the current gesture category is the same as the target gesture category. Therefore, if the current recognition result obtained from training is very close to the target recognition result corresponding to the training data, it means that the estimation accuracy of the control model is relatively high; if the current recognition result obtained from training is significantly different from the target recognition result corresponding to the training data, it is necessary to adjust the model parameters of the control model and then repeat the training until the current recognition result is close to the target recognition result or the current recognition result is the same as the target recognition result.
[0112] In one embodiment, step S204 may be specifically as follows: based on a preset loss function, adjusting the model parameters of the control model according to the deviation between the target recognition result and the current recognition result, the preset loss function includes a classification loss sub-function and a key point loss sub-function.
[0113] The loss function of the embodiment of the present application can be composed of a classification loss sub-function and a key point loss sub-function. Specifically, the parameters of the preset loss function can be adjusted in combination with the loss value of the index finger key point and the loss value of the middle finger key point. Based on the preset loss function, the model parameters of the control model can be adjusted according to the deviation between the target recognition result and the current recognition result. The trained control model can output the key point position and gesture category at the same time. The gesture classification does not rely on the results of the key point detection, which reduces the amount of calculation and is suitable for head-mounted display devices with low computing power. The control model can be used to control the head-mounted display device to perform tasks, so that the head-mounted display device can perform human-computer interaction tasks according to the recognized user gestures, enriching the human-computer interaction mode and improving the user's interactive experience with the head-mounted display device.
[0114] Optionally, a weight ratio may be configured for the classification loss sub-function and a weight ratio may be configured for the key point loss sub-function. For example, the preset loss function may be a weighted sum of the classification loss sub-function and the key point loss sub-function.
[0115] Optionally, the weight ratio of the classification loss sub-function and the key point loss sub-function can be adjusted according to the training situation of the control model to obtain a preset loss function that can adapt to the training situation. Of course, the weight ratio can be set according to the actual situation and is not specifically limited here.
[0116] The training method of the control model provided in the above embodiment includes obtaining training data, the training data including multiple images to be recognized and target recognition results corresponding to the images to be recognized, the target recognition results including target key point positions and target gesture categories; extracting hand feature vectors of the images to be recognized; performing key point detection on the hand feature vectors based on a gesture recognition network to obtain current key point positions of several preset hand key points, and determining the current gesture category corresponding to the image to be recognized based on the hand feature vectors; and adjusting model parameters of the control model based on the deviation between the target recognition results and the current recognition results, the current recognition result including the current key point positions and the current gesture category. The control model trained in the embodiment of the present application can simultaneously output key point positions and gesture categories through the same gesture recognition network. Since gesture classification does not rely on the results of key point detection, the amount of calculation is reduced and the computing power requirement is relatively low. It is suitable for head-mounted display devices with low computing power. Based on the key point positions and gesture categories, the head-mounted display device is controlled to perform tasks, so that the head-mounted display device can perform human-computer interaction tasks based on the recognized user gestures, enriching the human-computer interaction mode and improving the user's interactive experience with the head-mounted display device.
[0117] Please refer to Figure 8, which is a schematic block diagram of a control device provided in an embodiment of the present application. The control device can be configured in a server or electronic device to execute the aforementioned control method.
[0118] As shown in FIG8 , the control device is applied to a head-mounted display device, and the device includes: an acquisition module 110 , an extraction module 120 , an identification module 130 and a control module 140 .
[0119] The acquisition module 110 is configured to acquire an image to be recognized, where the image to be recognized includes at least a portion of a hand.
[0120] The extraction module 120 is used to extract the hand feature vector of the image to be recognized.
[0121] The recognition module 130 is used to perform key point detection on the hand feature vector based on the gesture recognition network, obtain the key point positions of several preset hand key points, and determine the gesture category corresponding to the image to be recognized based on the hand feature vector.
[0122] The control module 140 is configured to control the head mounted display device to execute a preset task according to a key point position and a gesture category of at least one preset hand key point.
[0123] In an exemplary embodiment, the gesture recognition network includes a key point detection head, and the recognition module 130 is specifically used to perform key point detection on the hand feature vector based on the key point detection head to obtain key point positions of several preset hand key points. The hand feature vector includes a finger feature vector, wherein the finger feature vector includes a thumb feature vector and / or an index finger feature vector.
[0124] In an exemplary embodiment, the identification module 130 includes a first identification submodule and / or a second identification submodule.
[0125] At least part of the hand includes a thumb, the hand feature vector includes a thumb feature vector, and the first recognition submodule is used to perform key point detection on the thumb feature vector based on a key point detection head to obtain a thumb key point position corresponding to the thumb key point.
[0126] At least part of the hand includes an index finger, the hand feature vector includes an index finger feature vector, and the second recognition submodule is used to perform key point detection on the index finger feature vector based on a key point detection head to obtain an index finger key point position corresponding to the index finger key point.
[0127] In an exemplary embodiment, the control module 140 includes a region determination submodule and a first control submodule.
[0128] The area determination submodule is used to determine the target display area according to the key point positions and the palm gesture when the gesture category is a palm gesture. The boundary of the target display area is determined according to the key point positions.
[0129] The first control submodule is used to control the head mounted display device to display preset content in the target display area.
[0130] In an exemplary embodiment, the preset hand key points include an index fingertip key point and a thumb fingertip key point, and the region determination submodule may include a first determination subunit and a second determination subunit.
[0131] The first determining subunit is configured to determine a boundary of the gesture area according to the palm gesture, the key point position corresponding to the index fingertip key point, and the key point position corresponding to the thumb fingertip key point when the gesture category is a palm gesture.
[0132] The second determining subunit is configured to determine the target display area according to the boundary of the gesture area and the key point position corresponding to the key point of the thumb fingertip.
[0133] In an exemplary embodiment, the image to be recognized includes a first hand and a second hand, and the gesture category of the first hand is a palm gesture. The control module 140 includes a display submodule, an action determination submodule and a second control submodule.
[0134] The display submodule is used to display preset content at the palm position in the palm gesture when the gesture category of the second hand is an indicating gesture.
[0135] The action determination submodule is used to determine the interactive action for the preset content according to the gesture category of the second hand and the key point position of the second hand.
[0136] The second control submodule is used to control the head mounted display device to perform the corresponding interactive task according to the interactive action.
[0137] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0138] The method of the present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0139] Illustratively, the above-mentioned method and apparatus may be implemented in the form of a computer program, which may be run on a computer device.
[0140] Please refer to Figure 9, which is a schematic block diagram of the structure of a computer device provided in an embodiment of the present application. The computer device can be a server or an electronic device.
[0141] As shown in FIG9 , the computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a storage medium and an internal memory.
[0142] The storage medium may store an operating system and a computer program. The computer program includes program instructions that, when executed, enable the processor to perform any of the steps of a method for controlling a head-mounted display device or a method for training a control model.
[0143] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0144] The internal memory provides an environment for the operation of the computer program in the storage medium. When the computer program is executed by the processor, the processor can execute the steps of any method for controlling a head-mounted display device or a method for training a control model.
[0145] This network interface is used for network communication, such as sending assigned tasks.
[0146] Those skilled in the art will understand that the structure shown in Figure 9 is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0147] It should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0148] In one embodiment, the processor is configured to execute a computer program and implement the following steps when executing the computer program:
[0149] Acquire an image to be recognized, where the image to be recognized includes at least a portion of a hand;
[0150] Extract the hand feature vector of the image to be recognized;
[0151] Based on the gesture recognition network, key point detection is performed on the hand feature vector to obtain the key point positions of several preset hand key points, and the gesture category corresponding to the image to be recognized is determined based on the hand feature vector;
[0152] Controlling a head-mounted display device to perform a preset task according to a key point position and a gesture category of at least one preset hand key point.
[0153] In one embodiment, the gesture recognition network includes a key point detection head. Based on the gesture recognition network, key point detection is performed on the hand feature vector to obtain key point positions of several preset hand key points, including:
[0154] Based on the key point detection head, key point detection is performed on the hand feature vector to obtain key point positions of several preset hand key points, where the hand feature vector includes a finger feature vector, wherein the finger feature vector includes a thumb feature vector and / or an index finger feature vector.
[0155] In one embodiment, at least part of the hand includes a thumb, and the hand feature vector includes a thumb feature vector. Key point detection is performed on the hand feature vector based on a key point detection head to obtain key point positions of several preset hand key points, including:
[0156] Based on the key point detection head, key point detection is performed on the thumb feature vector to obtain the thumb key point position corresponding to the thumb key point; and / or
[0157] At least part of the hand includes an index finger, and the hand feature vector includes an index finger feature vector. Based on the key point detection head, key point detection is performed on the hand feature vector to obtain key point positions of several preset hand key points, including:
[0158] Based on the key point detection head, key point detection is performed on the index finger feature vector to obtain the index finger key point position corresponding to the index finger key point.
[0159] In one embodiment, controlling a head mounted display device to perform a preset task based on a key point position and a gesture category of at least one preset hand key point includes:
[0160] When the gesture category is a palm gesture, the target display area is determined according to the key point positions and the palm gesture, and the boundary of the target display area is determined according to the key point positions;
[0161] Control the head-mounted display device to display preset content in the target display area.
[0162] In one embodiment, the preset hand key points include the index fingertip key point and the thumb fingertip key point. When the gesture category is a palm gesture, the target display area is determined based on the key point positions and the palm gesture, including:
[0163] When the gesture category is a palm gesture, the boundary of the gesture area is determined based on the palm gesture, the key point position corresponding to the index fingertip key point, and the key point position corresponding to the thumb fingertip key point;
[0164] The target display area is determined based on the boundary of the gesture area and the key point position corresponding to the thumb fingertip key point.
[0165] In one embodiment, controlling a head mounted display device to perform a preset task based on a key point position and a gesture category of at least one preset hand key point includes:
[0166] When the image to be recognized includes a first hand and a second hand, and the gesture category of the first hand is a palm gesture and the gesture category of the second hand is a pointing gesture, the preset content is displayed at the palm position of the palm gesture;
[0167] Determine the interactive action for the preset content based on the gesture category of the second hand and the key point position of the second hand;
[0168] According to the interactive action, the head-mounted display device is controlled to perform the corresponding interactive task.
[0169] Accordingly, in one embodiment, the processor is configured to execute the computer program and implement the following steps when executing the computer program:
[0170] Acquire training data, where the training data includes multiple images to be recognized and target recognition results corresponding to the images to be recognized, where the target recognition results include target key point locations and target gesture categories;
[0171] Extract the hand feature vector of the image to be recognized;
[0172] Based on the gesture recognition network, key point detection is performed on the hand feature vector to obtain the current key point positions of several preset hand key points, and the current gesture category corresponding to the image to be recognized is determined based on the hand feature vector;
[0173] The model parameters of the control model are adjusted according to the deviation between the target recognition result and the current recognition result, where the current recognition result includes the current key point position and the current gesture category.
[0174] It should be noted that, technical personnel in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process of the control method of the head-mounted display device described above can refer to the corresponding process in the embodiment of the control method of the head-mounted display device, and the specific working process of the training method of the control model can refer to the corresponding process in the embodiment of the training method of the control model, which will not be repeated here.
[0175] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. The method implemented when the computer program is executed by the processor can refer to the various embodiments of the control method of the head-mounted display device of the present application, or the training method of the control model of the head-mounted display device.
[0176] The computer-readable storage medium may be an internal storage unit of the computer device in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash memory card, etc. equipped on the computer device.
[0177] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0178] It should also be understood that the term "and / or" used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system that includes a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system that includes the element.
[0179] The serial numbers of the embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments. The above are only specific implementation methods of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A control method for a head-mounted display device, characterized in that Including: Obtain an image to be recognized, where the image to be recognized includes at least a part of a hand; Extract the hand feature vector of the image to be recognized; Based on a gesture recognition network, perform key point detection on the hand feature vector to obtain the key point positions of several preset hand key points, and determine the gesture category corresponding to the image to be recognized according to the hand feature vector; According to the key point positions of at least one of the preset hand key points and the gesture category, control a head-mounted display device to execute a preset task.
2. The control method of the head-mounted display device according to claim 1, characterized in that The gesture recognition network includes a key point detection head. Based on the gesture recognition network, performing key point detection on the hand feature vector to obtain the key point positions of several preset hand key points includes: Based on the key point detection head, perform key point detection on the hand feature vector to obtain the key point positions of several preset hand key points. The hand feature vector includes finger feature vectors, where the finger feature vectors include a thumb feature vector and / or an index finger feature vector.
3. The control method of the head-mounted display device according to claim 2, characterized in that, The at least part of the hand includes a thumb, and the hand feature vector includes a thumb feature vector. Based on the key point detection head, performing key point detection on the hand feature vector to obtain the key point positions of several preset hand key points includes: Based on the key point detection head, perform key point detection on the thumb feature vector to obtain the thumb key point position corresponding to the thumb key point; and / or The at least part of the hand includes an index finger, and the hand feature vector includes an index finger feature vector. Based on the key point detection head, performing key point detection on the hand feature vector to obtain the key point positions of several preset hand key points includes: Based on the key point detection head, perform key point detection on the index finger feature vector to obtain the index finger key point position corresponding to the index finger key point.
4. The control method of the head-mounted display device according to claim 2 or 3, characterized in that The controlling the head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category includes: When the gesture category is a palm gesture, determine a target display area according to the key point positions and the palm gesture, and the boundary of the target display area is determined according to the key point positions; Control the head-mounted display device to display preset content in the target display area.
5. The control method of the head-mounted display device according to claim 4, characterized in that, The preset hand key points include an index finger tip key point and a thumb tip key point. When the gesture category is a palm gesture, determining the target display area according to the key point positions and the palm gesture includes: When the gesture category is a palm gesture, determine the boundary of the gesture area according to the palm gesture, the key point position corresponding to the index finger tip key point, and the key point position corresponding to the thumb tip key point; Determine the target display area according to the boundary of the gesture area and the key point position corresponding to the thumb tip key point.
6. The control method of the head-mounted display device according to claim 2 or 3, characterized in that, The controlling the head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category includes: When the image to be recognized includes the hands of a first hand and a second hand, and the gesture category of the first hand is a palm gesture and the gesture category of the second hand is an indicating gesture, preset content is displayed at the palm position in the palm gesture; Determine an interaction action on the preset content according to the gesture category of the second hand and the key point positions of the second hand; Control the head-mounted display device to execute a corresponding interaction task according to the interaction action.
7. A method for training a control model of a head-mounted display device, characterized in that, The control model includes a gesture recognition network, and the training method includes: Obtain training data, where the training data includes a plurality of images to be recognized and the target recognition results corresponding to the images to be recognized, and the target recognition results include target key point positions and target gesture categories; Extract the hand feature vector of the image to be recognized; Based on the gesture recognition network, perform key point detection on the hand feature vector to obtain the current key point positions of several preset hand key points, and determine the current gesture category corresponding to the image to be recognized according to the hand feature vector; Adjust the model parameters of the control model according to the deviation between the target recognition result and the current recognition result, where the current recognition result includes the current key point positions and the current gesture category.
8. A control device for a head-mounted display device, characterized in that, Includes: An acquisition module, configured to acquire an image to be recognized, where the image to be recognized includes at least part of a hand; An extraction module, configured to extract the hand feature vector of the image to be recognized; A recognition module, configured to perform key point detection on the hand feature vector based on a gesture recognition network to obtain the key point positions of several preset hand key points, and determine the gesture category corresponding to the image to be recognized according to the hand feature vector; A control module, configured to control the head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category.
9. The control device of the head-mounted display device according to claim 7, wherein The gesture recognition network includes a key point detection head. When the recognition module performs key point detection on the hand feature vector based on the gesture recognition network to obtain the key point positions of several preset hand key points, it is used for: Perform key point detection on the hand feature vector based on the key point detection head to obtain the key point positions of several preset hand key points, where the hand feature vector includes finger feature vectors, and the finger feature vectors include thumb feature vectors and / or index finger feature vectors.
10. The control device of the head-mounted display device according to claim 9, characterized in that, The at least part of the hand includes a thumb, and the hand feature vector includes a thumb feature vector. When the recognition module performs key point detection on the hand feature vector based on the key point detection head to obtain the key point positions of several preset hand key points, it is used for: Perform key point detection on the thumb feature vector based on the key point detection head to obtain the thumb key point position corresponding to the thumb key point; and / or The at least part of the hand includes an index finger, and the hand feature vector includes an index finger feature vector. The performing key point detection on the hand feature vector based on the key point detection head to obtain the key point positions of several preset hand key points includes: Based on the key point detection head, perform key point detection on the index finger feature vector to obtain the position of the index finger key points corresponding to the index finger key points.
11. The control device of the head-mounted display device according to claim 9 or 10, characterized in that, When the control module controls the head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category, it is used for: When the gesture category is a palm gesture, determine a target display area according to the key point position and the palm gesture, and the boundary of the target display area is determined according to the key point position; Control the head-mounted display device to display preset content in the target display area.
12. The control device of the head-mounted display device according to claim 11, characterized in that, The preset hand key points include an index finger tip key point and a thumb tip key point. When the gesture category is a palm gesture, when the control module determines the target display area according to the key point position and the palm gesture, it is used for: When the gesture category is a palm gesture, determine the boundary of the gesture area according to the palm gesture, the key point position corresponding to the index finger tip key point, and the key point position corresponding to the thumb tip key point; Determine the target display area according to the boundary of the gesture area and the key point position corresponding to the thumb tip key point.
13. The control device of the head-mounted display device according to claim 9 or 10, characterized in that, When the control module controls the head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category, it is used for: When the image to be recognized includes the hands of the first hand and the second hand, and the gesture category of the first hand is a palm gesture and the gesture category of the second hand is an indicating gesture, display preset content at the palm position in the palm gesture; Determine an interaction action for the preset content according to the gesture category of the second hand and the key point positions of the second hand; Control the head-mounted display device to execute a corresponding interaction task according to the interaction action.
14. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store a computer program; The processor is used to execute the computer program and, when executing the computer program, implement: Obtain an image to be recognized, where the image to be recognized includes at least part of a hand; Extract the hand feature vector of the image to be recognized; Based on a gesture recognition network, perform key point detection on the hand feature vector to obtain the key point positions of several preset hand key points, and determine the gesture category corresponding to the image to be recognized according to the hand feature vector; Control the head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category; Or be used to implement: Obtain training data, where the training data includes multiple images to be recognized and the target recognition results corresponding to the images to be recognized, and the target recognition results include target key point positions and target gesture categories; Extract the hand feature vector of the image to be recognized; Based on the gesture recognition network of the control model, perform key point detection on the hand feature vector to obtain the current key point positions of several preset hand key points, and determine the current gesture category corresponding to the image to be recognized according to the hand feature vector; Adjust the model parameters of the control model according to the deviation between the target recognition result and the current recognition result, where the current recognition result includes the current key point positions and the current gesture category.
15. The computer device according to claim 14, characterized in that, The gesture recognition network includes a key point detection head. When the processor implements the gesture recognition network based on the hand feature vector to detect key points and obtain the key point positions of a plurality of preset hand key points, it is used to implement: Based on the key point detection head, detect key points of the hand feature vector to obtain the key point positions of a plurality of preset hand key points. The hand feature vector includes finger feature vectors, where the finger feature vectors include thumb feature vectors and / or index finger feature vectors.
16. The computer device according to claim 15, characterized in that, At least a part of the hand includes a thumb, and the hand feature vector includes a thumb feature vector. When the processor implements detecting key points of the hand feature vector based on the key point detection head to obtain the key point positions of a plurality of preset hand key points, it is used to implement: Based on the key point detection head, detect key points of the thumb feature vector to obtain the thumb key point position corresponding to the thumb key point; and / or At least a part of the hand includes an index finger, and the hand feature vector includes an index finger feature vector. Detecting key points of the hand feature vector based on the key point detection head to obtain the key point positions of a plurality of preset hand key points includes: Based on the key point detection head, detect key points of the index finger feature vector to obtain the index finger key point position corresponding to the index finger key point.
17. The computer device according to claim 15 or 16, characterized in that, When the processor implements controlling the head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category, it is used to implement: When the gesture category is a palm gesture, determine a target display area according to the key point positions and the palm gesture, and the boundary of the target display area is determined according to the key point positions; Control the head-mounted display device to display preset content in the target display area.
18. The computer device according to claim 17, wherein The preset hand key points include an index finger tip key point and a thumb tip key point. When the processor implements determining the target display area according to the key point positions and the palm gesture when the gesture category is a palm gesture, it is used to implement: When the gesture category is a palm gesture, determine the boundary of the gesture area according to the palm gesture, the key point position corresponding to the index finger tip key point, and the key point position corresponding to the thumb tip key point; Determine the target display area according to the boundary of the gesture area and the key point position corresponding to the thumb tip key point.
19. The computer device according to claim 15 or 16, characterized in that, When the processor implements controlling the head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category, it is used to implement: When the image to be recognized includes the hands of the first hand and the second hand, and the gesture category of the first hand is a palm gesture and the gesture category of the second hand is an indicating gesture, display preset content at the palm position in the palm gesture; Determine an interaction action on the preset content according to the gesture category of the second hand and the key point positions of the second hand; Control the head-mounted display device to perform a corresponding interaction task according to the interaction action.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the control method of the head-mounted display device described in any one of claims 1 to 6, or the steps of the training method of the control model of the head-mounted display device described in claim 7 are implemented.
Citation Information
Patent Citations
Gesture recognition-based distribution network training wearable equipment and interaction method thereof
CN109598998A
Gesture key point detection method and device, computer equipment and storage medium
CN111160288A
Mode-changeable augmented reality interface
CN113196213A
Menu interaction method and related equipment
CN115878013A
Input Device Gesture To Generate Full Screen Change
US20100238123A1
Cited By
Gesture recognition method and device for multi-user scene, equipment and program product
CN121232963A