Control method, model training method and device, equipment and storage medium
By simultaneously outputting key point positions and gesture categories in the gesture recognition network, the problems of large computing volume and high computing power requirements in the prior art are solved, and efficient gesture recognition on head-mounted display devices with low computing power is achieved, improving user interaction experience.
Patent Information
- Application Number
- CN202311805744.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art requires two networks to perform calculations in gesture recognition, resulting in large amounts of calculation and high computing power requirements. It is not suitable for head-mounted display devices with low computing power, and has poor user interaction experience.
By simultaneously outputting key point positions and gesture categories in the gesture recognition network, the dependence on key point detection results is reduced and the calculation amount is reduced. It is suitable for head-mounted display devices with low computing power.
It realizes efficient gesture recognition on head-mounted display devices with low computing power, improving the interactive experience between users and devices.
Smart Images

Figure CN120215682A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of head-mounted display devices, and particularly to a control method, a model training method, a device, a device and a storage medium. Background Art
[0002] Gesture recognition refers to capturing and analyzing human body postures, gesture actions and spatial position information through technologies such as sensors, computer vision and machine learning, so as to achieve natural interaction with a computer system. Applying gesture recognition technology to a head-mounted display device can provide users with a more intuitive, natural and immersive interaction experience. The head-mounted display device can recognize the user's gestures based on a gesture recognition algorithm, present virtual content, and display relevant information according to the received user operations.
[0003] Related technologies need to perform gesture classification according to the operation results of a network for detecting key points, resulting in a reduced speed of the entire operation process and a high requirement for computing power. Therefore, they are not applicable to head-mounted display devices with low computing power, and the interaction experience between users and head-mounted display devices is poor. Summary of the Invention
[0004] The main purpose of the present application is to provide a control method, a model training method, a device, a device and a storage medium, which are applicable to head-mounted display devices with low computing power and improve the interaction experience between users and head-mounted display devices.
[0005] In a first aspect, the present application provides a control method for a head-mounted display device, including:
[0006] Obtaining an image to be recognized, where the image to be recognized includes at least a part of a hand;
[0007] Extracting a hand feature vector of the image to be recognized;
[0008] Based on a gesture recognition network, performing key point detection on the hand feature vector to obtain the key point positions of a plurality of preset hand key points, and determining the gesture category corresponding to the image to be recognized according to the hand feature vector;
[0009] Controlling the head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category.
[0010] In a second aspect, the present application further provides a training method for a control model of a head-mounted display device, where the control model includes a gesture recognition network, and the training method includes:
[0011] Obtaining training data, where the training data includes a plurality of images to be recognized and the target recognition results corresponding to the images to be recognized, and the target recognition results include target key point positions and target gesture categories;
[0012] Extract the hand feature vector of the image to be recognized;
[0013] Based on the gesture recognition network, perform key point detection on the hand feature vector to obtain the current key point positions of several preset hand key points, and determine the current gesture category corresponding to the image to be recognized according to the hand feature vector;
[0014] Adjust the model parameters of the control model according to the deviation between the target recognition result and the current recognition result, where the current recognition result includes the current key point positions and the current gesture category.
[0015] In a third aspect, the present application further provides a control device for a head-mounted display device, including:
[0016] An acquisition module, configured to acquire an image to be recognized, where the image to be recognized includes at least a part of the hand;
[0017] An extraction module, configured to extract the hand feature vector of the image to be recognized;
[0018] A recognition module, configured to perform key point detection on the hand feature vector based on a gesture recognition network to obtain the key point positions of several preset hand key points, and determine the gesture category corresponding to the image to be recognized according to the hand feature vector;
[0019] A control module, configured to control the head-mounted display device to perform a preset task according to the key point positions of at least one of the preset hand key points and the gesture category.
[0020] In a fourth aspect, the present application further provides a computer device, where the computer device includes a memory and a processor;
[0021] The memory is used to store a computer program;
[0022] The processor is configured to execute the computer program and, when executing the computer program, implement the control method for the head-mounted display device as described above, or the training method for the control model of the head-mounted display device as described above.
[0023] In a fifth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the control method for the head-mounted display device as described above, or the steps of the training method for the control model of the head-mounted display device as described above are implemented.
[0024] The present application provides a control method, a model training method, an apparatus, a device, and a storage medium. The control method includes: obtaining an image to be recognized, where the image to be recognized includes at least a part of a hand; extracting a hand feature vector of the image to be recognized; based on a gesture recognition network, performing key point detection on the hand feature vector to obtain key point positions of a plurality of preset hand key points, and determining a gesture category corresponding to the image to be recognized according to the hand feature vector; and controlling a head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category. The present application performs key point detection on an image to be recognized and classifies the gesture of the image to be recognized based on a gesture recognition network. The key point positions and the gesture category are output by the same network at the same time. Since the gesture classification does not depend on the result of key point detection, the calculation amount is reduced, the requirement for computing power is small, and it is applicable to a head-mounted display device with low computing power. According to the key point positions and the gesture category obtained from the network output, the head-mounted display device is controlled to execute a task, so that the head-mounted display device can execute a human-computer interaction task according to the recognized user gesture, enriching the human-computer interaction method and improving the interaction experience between the user and the head-mounted display device. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 It is a schematic flowchart of a control method provided by an embodiment of the present application;
[0027] Figure 2 It is a schematic connection diagram of a server and smart glasses provided by an embodiment of the present application;
[0028] Figure 3 It is a schematic gesture classification diagram provided by an embodiment of the present application;
[0029] Figure 4 It is a schematic hand key point diagram provided by an embodiment of the present application;
[0030] Figure 5 It is a schematic diagram of a palm gesture and an indicating gesture provided by an embodiment of the present application;
[0031] Figure 6 It is a schematic diagram of preset content displayed at the palm position provided by an embodiment of the present application;
[0032] Figure 7 It is a schematic flowchart of a training method of a control model provided by an embodiment of the present application;
[0033] Figure 8 A schematic block diagram of a control device provided by an embodiment of the present application;
[0034] Figure 9 A schematic structural block diagram of a computer device provided by an embodiment of the present application. Specific embodiments
[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0036] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may change according to the actual situation.
[0037] In related technologies, two networks are required for gesture recognition. One network detects key points, and the other network needs to perform gesture classification based on the key point results detected by the previous network to obtain the gesture recognition result. That is, in related technologies, the network for gesture classification needs to perform gesture classification based on the operation results of the network for detecting key points. However, the two network load operations will reduce the speed of the entire operation process and require a large amount of computing power. Therefore, it is not suitable for head-mounted display devices with low computing power, and the interaction experience between users and head-mounted display devices is poor.
[0038] In view of this, the embodiments of the present application provide a control method, a model training method, a device, a device, and a storage medium. Among them, the control method can be applied to a head-mounted display device. It should be noted that the head-mounted display device can be an MR (Mixed Reality) glasses or an AR (Augmented Reality) glasses. In addition, the control method can also be applied to a server, which can be a single server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0039] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0040] See also Figure 1 , Figure 1 This is a flow chart of a control method provided in an embodiment of the present application. It should be noted that the control method provided in an embodiment of the present application can be used in a head mounted display device, and of course can also be used in a server.
[0041] like Figure 2 As shown, the control method is applied to a server, the server and the terminal are connected in communication, and the preset task obtained according to the control method can be sent to the terminal through the server. Of course, it is not limited to this and is not limited here.
[0042] In a specific implementation, the terminal includes a head-mounted display device, which may be MR glasses or AR glasses; the server may be a separate server, a server cluster, or a cloud server providing cloud computing services.
[0043] like Figure 1 As shown, the control method includes steps S101 to S105.
[0044] Step S101: Acquire an image to be recognized, where the image to be recognized includes at least part of a hand.
[0045] The image to be identified may include hands with different gestures, or may include hands in different parts. For example, the image to be identified may include only the palm, the back of the hand, or fingers making different shapes. In addition, the embodiment of the present application may perform pre-processing operations such as image enhancement and image resizing on the image to be identified to meet the needs of image processing and analysis of head-mounted display devices with low computing power requirements. Specifically, the image size is generally 224*224. The embodiment of the present application can reduce the image size to effectively reduce the amount of calculation and running time. Of course, the specific reduction ratio can be determined according to the actual situation, and no specific limitation is made here.
[0046] Step S102: extracting the hand feature vector of the image to be recognized.
[0047] It should be noted that, during the processing of the neural network, the image to be identified including at least part of the hand can be first input into a preset feature extractor to extract the hand feature vector in the image to be identified, for example, extracting the palm feature vector, the feature vectors of multiple fingers, etc. The preset feature extractor of the embodiment of the present application can adopt a convolutional neural network.
[0048] In some embodiments, a convolutional neural network may include an input layer, a convolutional layer, an activation layer, a pooling layer, and a fully connected layer. Specifically, the image to be recognized is input into the input layer of the convolutional neural network; each neuron in the convolutional layer is connected to a local area of the input layer. A convolutional layer may have multiple different convolutional kernels. Each convolutional kernel slides on the input image and processes only a small piece of the image at a time to extract the most basic features in the image to be recognized. By using the convolutional kernel to process the image to be recognized, the feature information of the local area in the image to be recognized can be extracted; the activation layer performs a non-linear mapping on the output result of the convolutional layer; the pooling layer compresses the input feature map to extract the main features, which can reduce the dimension of the data, thereby making the feature extractor more efficient; the fully connected layer is at the end of the convolutional neural network to connect all features, so as to obtain the final output value, that is, the hand feature vector.
[0049] In the preset feature extractor of the embodiment of the present application, the number of channels of each layer structure decreases, that is, the number of feature maps output by each layer decreases. For example, the number of channels of each layer structure can be reduced by a preset number to improve the feature extraction efficiency.
[0050] Step S103: Based on the gesture recognition network, perform key point detection on the hand feature vector to obtain the key point positions of several preset hand key points, and determine the gesture category corresponding to the image to be recognized according to the hand feature vector. Further, the gesture recognition network in the embodiment of the present application may adopt a lightweight network, which is suitable for a head-mounted display device with low computing power. For example, the lightweight network may include computer vision models MobileNets, ShuffleNet, GhostNet. It should be noted that the computer vision model MobileNets can be used to process various training tasks, including analyzing faces, detecting common objects, photo localization, etc.; ShuffleNet is a neural network structure designed specifically for devices with limited computing resources, which greatly reduces the computational overhead while retaining the model accuracy; GhostNet is a neural network architecture, the core of which is to propose a Ghost module, which can generate more feature maps, thereby reducing the amount of computation and improving the operation speed while ensuring the accuracy. In practical applications, according to the operation time consumption of the network on the head-mounted display device, the network with the least time consumption can be selected as the gesture recognition network in the embodiment of the present application, which can improve the efficiency of key point detection and gesture classification.
[0051] Specifically, the embodiment of the present application can perform coordinate mapping processing on the extracted hand feature vectors through a gesture recognition network, so as to obtain the key point positions of a number of preset hand key points. Specifically, the extracted hand feature vectors are segmented to obtain a number of segments of feature vectors. The maximum prediction probability of each segment of feature vectors among the number of segments of feature vectors is statistically calculated. Then, based on the statistically calculated maximum prediction probability, a heat map is drawn. According to the drawn heat map, the X coordinate information and Y coordinate information corresponding to a number of preset hand key points are mapped, so as to obtain the key point positions of the preset hand key points.
[0052] In gesture recognition, key point detection can detect at least one key point of the hand, and the number and positions of these key points vary according to specific applications.
[0053] Exemplarily, the preset hand key points may include at least one of the following: wrist key point, palm key point, finger key point. Among them, the wrist key point can be used as a reference point for the position and orientation of the hand. Fingers are the most commonly used moving parts of the hand. Therefore, detecting finger key points can provide important information about gestures.
[0054] Optionally, the finger key points may include at least one of the following: fingertip key point, knuckle key point, finger root key point. The palm key points may include at least one of the following: palm center key point, key point of the metacarpophalangeal joint, key point at the connection between the palm and the wrist.
[0055] Detecting the key point positions of the fingers includes the fingertips, knuckles, finger roots, etc. The palm is another important part of the hand, which can provide information such as the hand plane direction and hand posture. Detecting the key point positions of the palm includes the palm center and the palm edge, etc. The back of the hand and the back of the wrist can also be used as key point positions for gesture detection, providing information such as the hand orientation and tilt.
[0056] As Figure 4 shown, a general hand may include 21 key points. The preset hand key points in the embodiment of the present application may be key points on individual fingers or key points on the palm. That is, the gesture recognition network in the embodiment of the present application may include a key point detection head, and based on the key point detection head, the key point positions of some of the 21 key points can be output to reduce the calculation amount.
[0057] It should be noted that people do not need to use every part of the hand in their daily work and life. Generally, the frequency of using individual fingers is relatively high. For example, using the thumb and index finger to touch the screen, using the thumb or index finger to edit text, using the index finger to click, etc. Therefore, outputting some of the 21 key points can meet the needs of most scenarios. For example, based on a key point detection head, such as the detection head of a YOLO object detection model, that is, using two fully connected layers as the key point detection head, the hand feature vector of the image to be recognized can be detected, so that the positions of the palm and fingers can be located by detecting key points at specific positions on the palm and fingers. These key points can be finger tip key points, finger knuckle key points, palm center key points, etc. By detecting these key points, the position and orientation information of the palm and fingers corresponding to the key points can be obtained.
[0058] As Figure 3 shown, the gesture categories may include at least one of the following: call, like, ok, palm, stop, etc.
[0059] In some embodiments, the gesture recognition network in the embodiments of the present application may include a classification head. After extracting the hand feature vector of the image to be recognized, through the fully connected layer in the classification head, the gesture category of the image to be recognized can be output, without relying on the key point detection result, thereby reducing the overall operation process and improving the classification efficiency.
[0060] In addition, in the prior art, the dimension of the hand feature vector is generally relatively large. In the gesture recognition network of the embodiments of the present application, the dimension of each fully connected layer can be reduced while ensuring the gesture classification accuracy, so as to reduce the calculation amount.
[0061] Step S104, control the head-mounted display device to execute a preset task according to the key point positions and gesture categories of at least one preset hand key point.
[0062] Gesture recognition is a technology that can recognize user intentions and instructions by analyzing and understanding human actions and gestures, and can realize natural interaction between humans and computer systems. The head-mounted display device may include an MR glasses and an AR glasses. Among them, the gesture recognition technology based on the MR glasses is an application that combines augmented reality, virtual reality and gesture recognition. The MR glasses merge the real world and the virtual world together to create a new environment. Users can see real objects in the real world, and at the same time, they can also see virtual objects in the virtual world, and place the virtual objects in the real world, allowing users to interact with these virtual objects. The gesture recognition technology based on the AR glasses is an application that combines augmented reality and gesture recognition. The AR glasses can superimpose virtual information on the real world, enabling users to interact with virtual information and the real world.
[0063] As Figure 6 shown, in the embodiment of the present application, by obtaining the gesture category, such as the palm gesture or the pointing gesture, the intention and instruction of the user can be recognized, and then the AR glasses can be controlled to execute a preset task. For example, presenting virtual content, operating according to the pointing gesture, etc.
[0064] Applying the gesture recognition technology to the AR glasses can provide a more intuitive, natural and immersive interaction experience for users. By recognizing the user's gestures, the AR glasses can understand the user's intention and instruction, and accordingly present virtual content, perform operations or provide relevant information.
[0065] The control method provided by the above embodiment includes: obtaining an image to be recognized, where the image to be recognized includes at least part of the hand; extracting the hand feature vector of the image to be recognized; based on the gesture recognition network, performing key point detection on the hand feature vector to obtain the key point positions of several preset hand key points, and determining the gesture category corresponding to the image to be recognized according to the hand feature vector; controlling the head-mounted display device to execute a preset task according to the key point positions and gesture category of at least one preset hand key point. In the present application, key point detection is performed on the image to be recognized based on the gesture recognition network, and the gesture of the image to be recognized is classified. The key point positions and the gesture category are output by the network at the same time. Since the gesture classification does not depend on the result of the key point detection, the calculation amount is reduced, the requirement for computing power is small, and it is applicable to the head-mounted display device with low computing power. According to the key point positions and the gesture category obtained by the network output, the head-mounted display device is controlled to execute the task, so that the head-mounted display device can execute the man-machine interaction task according to the recognized user gesture, enriching the man-machine interaction method and improving the interaction experience between the user and the head-mounted display device.
[0066] In an exemplary embodiment, the gesture recognition network includes a key point detection head, and step S103 may specifically include S1030.
[0067] Step S1030: Based on the key point detection head, perform key point detection on the hand feature vector to obtain the key point positions of several preset hand key points. The hand feature vector includes finger feature vectors, where the finger feature vectors include thumb feature vectors and / or index finger feature vectors.
[0068] In the embodiment of the present application, it is not necessary to completely detect 21 key points. The key points of the two fingers that are not easily blocked, namely the user's thumb and / or index finger, can be detected. The user can perform operations such as clicking and swiping by using the thumb or index finger alone. If the user uses the thumb and index finger at the same time, operations such as pointing, zooming in, zooming out, and input can be realized. Therefore, by detecting the key points of the thumb and / or index finger, the user's instruction can be recognized.
[0069] Compared with the computational complexity of the 21 key points of all fingers, only calculating the two key points of the thumb and index finger can reduce the computational complexity. Based on this, the gesture recognition network in the embodiments of the present application includes a key point detection head, and based on the key point detection head, key point detection is performed on the hand feature vector in the image to be recognized. Specifically, the hand feature vector includes finger feature vectors, and only the thumb feature vector or the index finger feature vector can be extracted, or both the thumb feature vector and the index finger feature vector can be extracted. In practical applications, people mainly rely on the thumb and index finger to complete various activities. Therefore, the embodiments of the present application can perform key point detection on the thumb feature vector and the index finger feature vector to reduce the computational complexity.
[0070] In an exemplary embodiment, step S1030 may include step S1031 and / or step S1032.
[0071] Exemplarily, at least a part of the hand includes the thumb, and the hand feature vector includes the thumb feature vector.
[0072] Step S1031: Based on the key point detection head, perform key point detection on the thumb feature vector to obtain the thumb key point position corresponding to the thumb key point.
[0073] Exemplarily, at least a part of the hand includes the index finger, and the hand feature vector includes the index finger feature vector.
[0074] Step S1032: Based on the key point detection head, perform key point detection on the index finger feature vector to obtain the index finger key point position corresponding to the index finger key point.
[0075] In some embodiments, only the key points of the thumb feature vector can be detected, or only the key points of the index finger feature vector can be detected, or both the key points of the thumb feature vector and the index finger feature vector can be detected. Specifically, the thumb and index finger are not easily blocked, which improves the stability of key point detection. Through the key points of the thumb feature vector and / or the key points of the index finger feature vector, it can be determined whether the user's hand makes operations such as input, click, and slide.
[0076] Exemplarily, the key point detection head in the embodiments of the present application may adopt the detection head of the YOLO object detection model, that is, two fully connected layers are used as the key point detection head to detect the thumb feature vector and the index finger feature vector, so as to output four numerical values representing the key point positions corresponding to the thumb key point and the index finger key point. Taking the YOLOv5 object detection model as an example, it uses a key point regressor module to predict the position of the key point. After inputting hand feature vectors such as the thumb feature vector and the index finger feature vector into the key point regressor, the coordinate information corresponding to the thumb key point, that is, the thumb key point position, or the coordinate information corresponding to the index finger key point, that is, the index finger key point position, can be obtained.
[0077] In an exemplary embodiment, step S104 includes step S1041A and step S1042A.
[0078] Step S1041A: When the gesture category is a palm gesture, determine a target display area according to the key point positions and the palm gesture, and the boundary of the target display area is determined according to the key point positions.
[0079] Step S1042A: Control the head-mounted display device to display preset content in the target display area.
[0080] In practical applications, text input on AR glasses is a particularly important part of human-computer interaction. In the embodiments of the present application, a virtual keyboard of the AR glasses can be displayed at the palm center position of one hand of the user. When the AR glasses receive various operations of the fingers of the user's other hand on the virtual keyboard, corresponding content is displayed, such as the virtual keyboard. Without the need for other media, the position of the displayed content can be controlled, which is convenient for use in an environment with a small and crowded space, thereby realizing human-computer interaction and improving the user's human-computer interaction experience.
[0081] Compared with the solution of a floating virtual keyboard, which requires the user to turn the head and use the fingers to tap the buttons floating in the air to select the buttons on the keyboard, the embodiments of the present application can be realized without the user turning the head, which is convenient for controlling the position of the keyboard, enabling the virtual content display to be unrestricted by space, and having high efficiency and convenience in human-computer interaction.
[0082] Compared with the solution of using a mobile phone for cross-screen input, the embodiments of the present application do not require adding other media and can control the position of the keyboard, which is relatively convenient.
[0083] In an exemplary embodiment, the preset hand key points include the index finger tip key point and the thumb tip key point, and step S1041A includes step SA1 and step SB1.
[0084] Step SA1: When the gesture category is a palm gesture, determine the boundary of the gesture area according to the palm gesture, the key point position corresponding to the index finger tip key point, and the key point position corresponding to the thumb tip key point.
[0085] Step SB1: Determine the target display area according to the boundary of the gesture area and the key point position corresponding to the thumb tip key point.
[0086] When the recognized gesture category is the palm gesture, the boundaries of the gesture area that can exactly cover the palm gesture can be determined according to the positions of the key points of the index finger tip and the thumb tip detected, as well as the direction of the palm gesture. Then, according to the key point position corresponding to the thumb tip, the intersection point of two adjacent boundaries of the gesture area closest to the thumb tip key point is determined. Then, taking this intersection point as the reference point, the boundaries of the target display area are determined.
[0087] It can be understood that the shapes of the target display area and the gesture area can both be rectangles. According to the intersection points of two adjacent boundaries of the gesture area, the intersection points of two adjacent boundaries of the target display area can be determined. Finally, according to the size of the boundaries of the gesture area, the boundaries of the target display area can be determined, so as to obtain the target display area. In addition, the size of the target display area in the embodiments of the present application can be adapted to the size of the user's palm center.
[0088] Determining the target display area through the gesture area enables the target display area to be exactly anchored at the palm center position of the user's palm, and the size of the target display area is determined by the size of the gesture area. Therefore, hands of different sizes can be applicable, thereby improving the user experience.
[0089] The following is discussed in combination with an actual application scenario:
[0090] It can be understood that the target display area can be understood as the anchoring position of the input box (virtual keyboard), and the preset content can be text, pictures, or keyboard keys in the input box (virtual keyboard), etc. As Figure 6 shown, when the gesture category is the palm gesture, the boundaries of the target display area can be determined according to the key point positions to control the head-mounted display device to anchor the AR glasses virtual keyboard interface at the palm center position of the palm gesture.
[0091] As Figure 5 shown, the determination process of the boundaries of the target display area is as follows:
[0092] 1. Detect the position of the hand, that is, the rectangular frame ABCD in the figure;
[0093] 2. Recognize the hand. When the palm gesture palm is recognized and the palm position in the palm gesture palm is detected, that is, Figure 5 the rectangular frame (target display area) E-F in, the palm input method is enabled;
[0094] 3. Anchor the palm input box: According to the key point position of the recognized index finger tip, calculate the point among the four corner points of the rectangular frame ABCD that is closest to the index finger tip, that is, Figure 5 midpoint C, and among the sides CD and BC passing through point C, select the side that is closest to the thumb tip point, that is, Figure 5For CD, with D as the reference point, the position of the palm can be estimated; the coordinates of E in the gesture frame ABCD are (1 / 10*CD, 1 / 3*AD), and the coordinates of F are (1 / 2*AB, 4 / 5*AD), where the coefficients can be changed according to the actual application of the product, and this is just an example here.
[0095] 4. When the other hand appears, perform detection and recognition. When the indicating gesture "indicate" is recognized, the click operation of the fingertip of the index finger can be responded to.
[0096] 5. When the fingertip of the index finger of the indicating gesture "indicate" touches the key corresponding to the palm position input box (target display area) of the palm gesture "palm", the corresponding information can be displayed.
[0097] In an exemplary embodiment, step S105 includes step S1051B, step S1052B, and step S1053B.
[0098] Step S1051B: When the image to be recognized includes the hand of the first hand and the hand of the second hand, and the gesture category of the first hand is the palm gesture and the gesture category of the second hand is the indicating gesture, display the preset content at the palm position in the palm gesture.
[0099] Step S1052B: Determine the interaction action on the preset content according to the gesture category of the second hand and the key point positions of the second hand.
[0100] Step S1053B: Control the head-mounted display device to perform the corresponding interaction task according to the interaction action.
[0101] In some embodiments, the image to be recognized may include the hand of the first hand and the hand of the second hand. If the gesture category of the first hand is the palm gesture, the head-mounted display device can be controlled to display the preset content of the target display area, that is, the preset content such as the text, picture, or keyboard key in the input box, at the palm position in the palm gesture. When the gesture category of the second hand is the indicating gesture, the interaction action of the second hand on the preset content can be determined according to the detected key point positions of the finger feature vector of the second hand. For example, when an AR glasses or an MR glasses detects the user's palm gesture, a keyboard is displayed at the palm position of the palm gesture. At this time, the user's other hand can operate on the keyboard at will. If only the key points of the fingertip of the index finger and the key points near the connection position of the index finger and the thumb are detected, the gesture category of this hand can be determined as the indicating gesture. Then, according to the position (such as the positions of numbers 1, 0, 6, etc. on the keyboard) or the moving range (continuously moving to form an unlocking pattern S) on the keyboard corresponding to the detected key points of the fingertip of the index finger, the interaction action of this hand on the keyboard can be determined.
[0102] Such as Figure 6As shown, an input box is displayed at the palm position of one hand, and the input box includes multiple control buttons. The index finger of the other hand clicks on the input box. Under this interaction action, the head-mounted display device performs corresponding interaction tasks. For example, interaction actions such as expansion, editing, deletion, and sliding. The head-mounted display device performs corresponding interaction tasks according to interaction actions such as expansion, editing, deletion, and sliding. For example, expand multi-level directories, display an editing bar, delete the content from the input box, or switch to the next page. The palm of one hand of the user displays preset content such as an input box, and the other hand operates various interaction actions. Therefore, it is possible to control the preset content of the target display area to be displayed at any position, making the position and operation space of human-computer interaction relatively flexible, thereby improving the convenience and experience of user human-computer interaction.
[0103] Please refer to the foregoing embodiments in conjunction with Figure 7 , Figure 7 which is a schematic diagram of a method for training a control model provided by an embodiment of the present application.
[0104] The training method of this control model can be applied to a head-mounted display device, and of course it is not limited thereto. For example, it can be applied to a computer or a server. If the training method of the control model is applied to a server, the control model obtained by the training method can be sent to the head-mounted display device through the server. Of course, it is not limited thereto, and no limitation is made here.
[0105] The control model includes a gesture recognition network, and the training method includes steps S201 to S206.
[0106] Step S201: Obtain training data. The training data includes multiple images to be recognized and the corresponding target recognition results of the images to be recognized. The target recognition results include the target key point positions and the target gesture categories.
[0107] Step S202: Extract the hand feature vectors of the images to be recognized.
[0108] Step S203: Based on the gesture recognition network, perform key point detection on the hand feature vectors to obtain the current key point positions of several preset hand key points, and determine the current gesture category corresponding to the image to be recognized according to the hand feature vectors.
[0109] Step S204: Adjust the model parameters of the control model according to the deviation between the target recognition result and the current recognition result. The current recognition result includes the current key point positions and the current gesture category.
[0110] Model parameters refer to those that control the estimation ability of the model. The model parameters of the control model are crucial for the accuracy of its estimation. During the process of training the control model based on a neural learning network, it is generally necessary to continuously adjust the model parameters to make the current recognition result obtained by the control model the same as the target recognition result, that is, the current key point positions and the target key point positions are the same, and the current gesture category and the target gesture category are the same. Therefore, if the current recognition result obtained through training is very close to the target recognition result corresponding to the training data, it indicates that the estimation accuracy of the control model is relatively high; if there are significant differences between the current recognition result obtained through training and the target recognition result corresponding to the training data, it is necessary to adjust the model parameters of the control model and then repeat the training until the current recognition result is close to the target recognition result or the current recognition result is the same as the target recognition result.
[0111] In one embodiment, step S204 can be specifically: based on a preset loss function, adjust the model parameters of the control model according to the deviation between the target recognition result and the current recognition result. The preset loss function includes a classification loss sub-function and a key point loss sub-function.
[0112] The loss function of the embodiment of the present application can be composed of a classification loss sub-function and a key point loss sub-function. Specifically, the parameters of the preset loss function can be adjusted by combining the loss values of the index finger key points and the middle finger key points. Based on the preset loss function, the model parameters of the control model can be adjusted according to the deviation between the target recognition result and the current recognition result. The trained control model can output the key point positions and gesture categories simultaneously. Gesture classification does not depend on the results of key point detection, reducing the computational amount. It is applicable to head-mounted display devices with low computing power. Moreover, through this control model, the head-mounted display device can be controlled to perform tasks, enabling the head-mounted display device to execute human-computer interaction tasks according to the recognized user gestures, enriching the human-computer interaction method and improving the interaction experience between the user and the head-mounted display device.
[0113] Optionally, a weight ratio can also be configured for the classification loss sub-function and a weight ratio can be configured for the key point loss sub-function. For example, the preset loss function can be the weighted sum of the classification loss sub-function and the key point loss sub-function.
[0114] Optionally, the weight ratios of the classification loss sub-function and the key point loss sub-function can also be adjusted according to the training situation of the control model, so as to obtain a preset loss function that can adapt to the training situation. Of course, the weight ratios can be set according to the actual situation and are not specifically limited here.
[0115] The training method of the control model provided by the above embodiments includes obtaining training data, where the training data includes multiple images to be recognized and the corresponding target recognition results of the images to be recognized. The target recognition results include the positions of target key points and the target gesture categories; extracting the hand feature vectors of the images to be recognized; based on the gesture recognition network, performing key point detection on the hand feature vectors to obtain the current key point positions of several preset hand key points, and determining the current gesture category corresponding to the image to be recognized according to the hand feature vectors; adjusting the model parameters of the control model according to the deviation between the target recognition results and the current recognition results, where the current recognition results include the current key point positions and the current gesture categories. The control model trained in the embodiments of the present application can simultaneously output the key point positions and gesture categories through the same gesture recognition network. Since gesture classification does not depend on the results of key point detection, the amount of calculation is reduced, the requirement for computing power is small, and it is applicable to head-mounted display devices with low computing power. According to the key point positions and gesture categories, the head-mounted display device is controlled to execute tasks, so that the head-mounted display device can execute the tasks of human-computer interaction according to the recognized user gestures, enriching the human-computer interaction method and improving the interaction experience between the user and the head-mounted display device.
[0116] Please refer to Figure 8 , Figure 8 which is a schematic block diagram of a control device provided by an embodiment of the present application. This control device can be configured in a server or an electronic device and is used to execute the foregoing control method.
[0117] As Figure 8 shown, this control device is applied to a head-mounted display device, and the device includes: an acquisition module 110, an extraction module 120, an identification module 130, and a control module 140.
[0118] The acquisition module 110 is used to acquire an image to be recognized, where the image to be recognized includes at least part of a hand.
[0119] The extraction module 120 is used to extract the hand feature vectors of the image to be recognized.
[0120] The identification module 130 is used to perform key point detection on the hand feature vectors based on the gesture recognition network to obtain the key point positions of several preset hand key points, and determine the gesture category corresponding to the image to be recognized according to the hand feature vectors.
[0121] The control module 140 is used to control the head-mounted display device to execute a preset task according to the key point positions and gesture categories of at least one preset hand key point.
[0122] In an exemplary embodiment, the gesture recognition network includes a key point detection head. The recognition module 130 is specifically configured to perform key point detection on the hand feature vector based on the key point detection head to obtain the key point positions of a plurality of preset hand key points. The hand feature vector includes finger feature vectors, where the finger feature vector includes a thumb feature vector and / or an index finger feature vector.
[0123] In an exemplary embodiment, the recognition module 130 includes a first recognition sub-module and / or a second recognition sub-module.
[0124] At least a part of the hand includes a thumb, and the hand feature vector includes a thumb feature vector. The first recognition sub-module is configured to perform key point detection on the thumb feature vector based on the key point detection head to obtain the key point position corresponding to the thumb key point.
[0125] At least a part of the hand includes an index finger, and the hand feature vector includes an index finger feature vector. The second recognition sub-module is configured to perform key point detection on the index finger feature vector based on the key point detection head to obtain the key point position corresponding to the index finger key point.
[0126] In an exemplary embodiment, the control module 140 includes a region determination sub-module and a first control sub-module.
[0127] The region determination sub-module is configured to determine a target display region according to the key point position and the palm gesture when the gesture category is a palm gesture. The boundary of the target display region is determined according to the key point position.
[0128] The first control sub-module is configured to control the head-mounted display device to display preset content in the target display region.
[0129] In an exemplary embodiment, the preset hand key points include an index finger tip key point and a thumb tip key point. The region determination sub-module may include a first determination sub-unit and a second determination sub-unit.
[0130] The first determination sub-unit is configured to determine the boundary of the gesture region according to the palm gesture, the key point position corresponding to the index finger tip key point, and the key point position corresponding to the thumb tip key point when the gesture category is a palm gesture.
[0131] The second determination sub-unit is configured to determine the target display region according to the boundary of the gesture region and the key point position corresponding to the thumb tip key point.
[0132] In an exemplary embodiment, the image to be recognized includes the hands of a first hand and a second hand, and the gesture category of the first hand is a palm gesture. The control module 140 includes a display sub-module, an action determination sub-module, and a second control sub-module.
[0133] A display sub-module, configured to display preset content at the palm position in the palm gesture when the gesture category of the second hand is an indicating gesture.
[0134] An action determination sub-module, configured to determine an interaction action on the preset content according to the gesture category of the second hand and the key point positions of the second hand.
[0135] A second control sub-module, configured to control the head-mounted display device to execute a corresponding interaction task according to the interaction action.
[0136] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described device and each module and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0137] The method of the present application can be used in many general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that execute specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are executed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0138] Exemplarily, the above method and device can be implemented in the form of a computer program, and the computer program can run on a computer device.
[0139] Please refer to Figure 9 , Figure 9 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. The computer device can be a server or an electronic device.
[0140] As Figure 9 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a storage medium and an internal memory.
[0141] The storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can be caused to execute the steps of any control method of a head-mounted display device or a training method of a control model.
[0142] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0143] The internal memory provides an environment for the operation of the computer program in the storage medium. When the computer program is executed by the processor, the processor can be caused to execute the steps of any control method or training method of the control model for a head-mounted display device.
[0144] The network interface is used for network communication, such as sending assigned tasks, etc.
[0145] Those skilled in the art can understand that Figure 9 The structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0146] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0147] Among them, in one embodiment, the processor is used to execute a computer program and implement the following steps when executing the computer program:
[0148] Obtain an image to be recognized, where the image to be recognized includes at least part of a hand;
[0149] Extract the hand feature vector of the image to be recognized;
[0150] Based on a gesture recognition network, perform key point detection on the hand feature vector to obtain the key point positions of several preset hand key points, and determine the gesture category corresponding to the image to be recognized according to the hand feature vector;
[0151] Control the head-mounted display device to execute a preset task according to the key point positions and gesture categories of at least one preset hand key point.
[0152] In one embodiment, the gesture recognition network includes a key point detection head. Based on the gesture recognition network, key point detection is performed on the hand feature vector to obtain the key point positions of a plurality of preset hand key points, including:
[0153] Based on the key point detection head, key point detection is performed on the hand feature vector to obtain the key point positions of a plurality of preset hand key points. The hand feature vector includes finger feature vectors, where the finger feature vectors include thumb feature vectors and / or index finger feature vectors.
[0154] In one embodiment, at least a part of the hand includes a thumb, and the hand feature vector includes a thumb feature vector. Based on the key point detection head, key point detection is performed on the hand feature vector to obtain the key point positions of a plurality of preset hand key points, including:
[0155] Based on the key point detection head, key point detection is performed on the thumb feature vector to obtain the thumb key point position corresponding to the thumb key point; and / or
[0156] At least a part of the hand includes an index finger, and the hand feature vector includes an index finger feature vector. Based on the key point detection head, key point detection is performed on the hand feature vector to obtain the key point positions of a plurality of preset hand key points, including:
[0157] Based on the key point detection head, key point detection is performed on the index finger feature vector to obtain the index finger key point position corresponding to the index finger key point.
[0158] In one embodiment, according to the key point positions of at least one preset hand key point and the gesture category, the head-mounted display device is controlled to execute a preset task, including:
[0159] When the gesture category is a palm gesture, according to the key point positions and the palm gesture, a target display area is determined, and the boundary of the target display area is determined according to the key point positions.
[0160] Control the head-mounted display device to display preset content in the target display area.
[0161] In one embodiment, the preset hand key points include an index finger tip key point and a thumb tip key point. When the gesture category is a palm gesture, according to the key point positions and the palm gesture, a target display area is determined, including:
[0162] When the gesture category is a palm gesture, according to the palm gesture, the key point position corresponding to the index finger tip key point, and the key point position corresponding to the thumb tip key point, the boundary of the gesture area is determined;
[0163] According to the boundary of the gesture area and the key point position corresponding to the thumb tip key point, the target display area is determined.
[0164] In one embodiment, according to the key point positions and gesture categories of at least one preset hand key point, controlling a head-mounted display device to execute a preset task, including:
[0165] When the image to be recognized includes the hands of the first hand and the second hand, and the gesture category of the first hand is a palm gesture and the gesture category of the second hand is an indicating gesture, display preset content at the palm position in the palm gesture;
[0166] Determine an interaction action for the preset content according to the gesture category of the second hand and the key point positions of the second hand;
[0167] Control the head-mounted display device to execute a corresponding interaction task according to the interaction action.
[0168] Correspondingly, in one embodiment, a processor is configured to execute a computer program and implement the following steps when executing the computer program:
[0169] Obtain training data, where the training data includes a plurality of images to be recognized and target recognition results corresponding to the images to be recognized, and the target recognition results include target key point positions and target gesture categories;
[0170] Extract the hand feature vectors of the images to be recognized;
[0171] Based on a gesture recognition network, perform key point detection on the hand feature vectors to obtain the current key point positions of several preset hand key points, and determine the current gesture category corresponding to the image to be recognized according to the hand feature vectors;
[0172] Adjust the model parameters of the control model according to the deviation between the target recognition result and the current recognition result, where the current recognition result includes the current key point positions and the current gesture category.
[0173] It should be noted that those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process of the above-described control method of the head-mounted display device can refer to the corresponding process in the embodiment of the control method of the head-mounted display device described above, and the specific working process of the training method of the control model can refer to the corresponding process in the embodiment of the training method of the control model described above, which will not be elaborated here.
[0174] The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and the method implemented when the computer program is executed by a processor can refer to each embodiment of the control method of the head-mounted display device or the training method of the control model of the head-mounted display device in the present application.
[0175] Among them, the computer-readable storage medium can be an internal storage unit of the computer device in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0176] It should be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0177] It should also be understood that the term "and / or" used in this specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. It should be noted that in this article, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.
[0178] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments. The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art in the technical field disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A control method for a head-mounted display device, characterized in that Including: Obtain an image to be recognized, where the image to be recognized includes at least part of a hand; Extract the hand feature vector of the image to be recognized; Based on a gesture recognition network, perform key point detection on the hand feature vector to obtain the key point positions of several preset hand key points, and determine the gesture category corresponding to the image to be recognized according to the hand feature vector; Control a head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category.
2. The control method of the head-mounted display device according to claim 1, wherein The gesture recognition network includes a key point detection head. Based on the gesture recognition network, performing key point detection on the hand feature vector to obtain the key point positions of several preset hand key points includes: Based on the key point detection head, perform key point detection on the hand feature vector to obtain the key point positions of several preset hand key points. The hand feature vector includes finger feature vectors, where the finger feature vectors include a thumb feature vector and / or an index finger feature vector.
3. The control method of the head-mounted display device according to claim 2, wherein, The at least part of the hand includes a thumb, and the hand feature vector includes a thumb feature vector. Based on the key point detection head, performing key point detection on the hand feature vector to obtain the key point positions of several preset hand key points includes: Based on the key point detection head, perform key point detection on the thumb feature vector to obtain the thumb key point position corresponding to the thumb key point; and / or The at least part of the hand includes an index finger, and the hand feature vector includes an index finger feature vector. Based on the key point detection head, performing key point detection on the hand feature vector to obtain the key point positions of several preset hand key points includes: Based on the key point detection head, perform key point detection on the index finger feature vector to obtain the index finger key point position corresponding to the index finger key point.
4. The control method of the head-mounted display device according to claim 2 or 3, characterized in that, The controlling the head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category includes: When the gesture category is a palm gesture, determine a target display area according to the key point positions and the palm gesture, where the boundary of the target display area is determined according to the key point positions; Control the head-mounted display device to display preset content in the target display area.
5. The control method of the head-mounted display device according to claim 4, characterized in that, The preset hand key points include an index finger tip key point and a thumb tip key point. When the gesture category is a palm gesture, determining the target display area according to the key point positions and the palm gesture includes: When the gesture category is a palm gesture, determine the boundary of the gesture area according to the palm gesture, the key point position corresponding to the index finger tip key point, and the key point position corresponding to the thumb tip key point; Determine the target display area according to the boundary of the gesture area and the key point position corresponding to the thumb tip key point.
6. The control method of the head-mounted display device according to claim 2 or 3, characterized in that, The controlling the head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category includes: The image to be recognized includes the hands of a first hand and a second hand, and when the gesture category of the first hand is a palm gesture and the gesture category of the second hand is an indicating gesture, preset content is displayed at the palm position in the palm gesture; Determine an interaction action for the preset content according to the gesture category of the second hand and the key point positions of the second hand; Control the head-mounted display device to execute a corresponding interaction task according to the interaction action.
7. A training method for a control model of a head-mounted display device, characterized in that, The control model includes a gesture recognition network, and the training method includes: Obtain training data, where the training data includes a plurality of images to be recognized and the target recognition results corresponding to the images to be recognized, and the target recognition results include target key point positions and target gesture categories; Extract the hand feature vector of the image to be recognized; Based on the gesture recognition network, perform key point detection on the hand feature vector to obtain the current key point positions of several preset hand key points, and determine the current gesture category corresponding to the image to be recognized according to the hand feature vector; Adjust the model parameters of the control model according to the deviation between the target recognition result and the current recognition result, where the current recognition result includes the current key point positions and the current gesture category.
8. A control device for a head-mounted display device, characterized in that, Includes: An acquisition module, configured to acquire an image to be recognized, where the image to be recognized includes at least a part of a hand; An extraction module, configured to extract the hand feature vector of the image to be recognized; A recognition module, configured to perform key point detection on the hand feature vector based on a gesture recognition network to obtain the key point positions of several preset hand key points, and determine the gesture category corresponding to the image to be recognized according to the hand feature vector; A control module, configured to control the head-mounted display device to execute a preset task according to the key point positions of at least one of the preset hand key points and the gesture category.
9. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store a computer program; The processor is configured to execute the computer program and, when executing the computer program, implement the control method of the head-mounted display device according to any one of claims 1 to 6, or the training method of the control model of the head-mounted display device according to claim 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the control method of the head-mounted display device according to any one of claims 1 to 6, or the steps of the training method of the control model of the head-mounted display device according to claim 7 are implemented.