A hand gesture recognition method, a head-mounted display device, and a storage medium

Through multiple iterative projections, the distortion of hand image is corrected, the distortion-free image is generated and the coordinates of hand joint nodes are optimized, which solves the problem of inaccurate hand posture recognition in the prior art, and improves the accuracy of hand posture recognition and user interaction experience.

CN120014714BActive Publication Date: 2025-07-11HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510496558.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-11
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

In the prior art, due to the lack of distortion correction for hand images captured by binocular cameras, the predicted 3D hand joint node coordinates are inaccurate, which affects the accuracy of hand posture recognition results.

Method used

The current frame image is corrected by multiple iterative projection directions, a hand image without distortion is generated, and it is input into the pre-trained hand joint node detection model for identification, and the three-dimensional coordinates of the hand joint node are optimized using the objective function.

Benefits of technology

It improves the accuracy of hand joint node recognition and the accuracy of hand posture recognition, is suitable for complex gesture interaction scenarios, and enhances the interactive experience between users and AR/VR devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014714B_ABST
    Figure CN120014714B_ABST
Patent Text Reader

Abstract

The present invention discloses a hand gesture recognition method, a head-mounted display device and a storage medium, relating to the field of human-computer interaction. In the method, the projection direction is iterated multiple times to perform projection correction on the image of the first hand region in the current frame image, so that an image taken due to the inclination of the observation line of sight is corrected into an image with the line of sight perpendicular to the center. In this way, the distortion influence on the hand image fed into the model for inference is overcome, and the positions and relative distances of the hand joint points obtained by the model inference will be more accurate. The problem that the coordinates of the predicted 3D hand joint points are inaccurate due to image distortion is solved by the present invention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-computer interaction technologies, and particularly to a hand gesture recognition method, a head-mounted display device, and a storage medium. Background Art

[0002] In the field of virtual reality (VR) or augmented reality (AR) technologies, three-dimensional gesture technologies are used to enable the wearer of a head-mounted display device to interact naturally and directly with virtual objects in the VR / AR scene using both hands.

[0003] The interaction technology uses a three-dimensional hand model skeleton as the object of gesture analysis, and through the recognition and tracking of multiple hand joint points and degrees of freedom, it discriminates various hand postures of single and both hands, thereby realizing human-computer interaction. Therefore, the extraction of joint points and the determination of degrees of freedom are important research directions in three-dimensional gesture estimation.

[0004] Currently, a binocular camera on a head-mounted display device is used to obtain hand images in real time from different perspectives, and two-dimensional (2D) hand joint points extracted from the hand images are converted into three-dimensional (3D) hand joint points through binocular camera projection.

[0005] Due to the lack of distortion correction for the hand images captured by the binocular camera, the predicted 3D hand joint points have inaccurate coordinates due to image distortion, which affects the subsequent hand gesture recognition results. Summary of the Invention

[0006] Embodiments of this application provide a hand gesture recognition method, a head-mounted display device, and a storage medium, which can obtain a distortion-free hand image from the original binocular image and improve the accuracy of hand joint point recognition results.

[0007] To achieve the above objective, the embodiments of this application adopt the following technical solutions:

[0008] In a first aspect, an embodiment of this application provides a hand gesture recognition method, and the method includes:

[0009] Tracking and obtaining a first hand region of the current frame image in a hand binocular video, and projecting the first hand region through binocular disparity to obtain a first hand bounding box of the first hand region in three-dimensional space;

[0010] Selecting an initial center position of the hand from the first hand bounding box, and calculating a first transformation matrix in the direction from the initial center position to the image optical center;

[0011] Iteratively optimize the first transformation matrix multiple times, and use the iteratively optimized first transformation matrix as the first projection direction;

[0012] According to the first projection direction, perform projection correction on the image of the first hand region in the current frame image to generate a corrected first hand image, where the first hand image is used as input to a pre-trained hand joint point detection model for hand joint point recognition to obtain a recognition result.

[0013] Combined with the first aspect, in a possible design, the iterative process of the first transformation matrix is as follows:

[0014] Based on the first transformation matrix, correct the positions of the hand joint points in the first hand bounding box, and obtain the first direction vector of the corrected hand center point in three-dimensional space;

[0015] Use the product of the first transformation matrix and the first direction vector as the first transformation matrix for the next iterative optimization, and repeat the above iterative process until the first transformation matrix converges and stop the iteration.

[0016] Combined with the first aspect, in a possible design, the method of correcting the positions of the hand joint points in the first hand bounding box based on the first transformation matrix and obtaining the first direction vector of the corrected hand center point in three-dimensional space includes:

[0017] Based on the first transformation matrix, transform the positions of the hand joint points in the first hand bounding box from the first three-dimensional positions to the second three-dimensional positions;

[0018] Obtain the position information of the projection center point of the second three-dimensional position on the celestial equator plane to obtain the target center position;

[0019] Back-project the target center position through the celestial equator plane to obtain the first direction vector of the corrected hand center point in three-dimensional space.

[0020] Combined with the first aspect, in a possible design, the method of performing projection correction on the image of the first hand region in the current frame image according to the first projection direction to generate a corrected first hand image includes:

[0021] According to the first projection direction, project the hand joint points in the first hand bounding box onto the celestial equator plane to obtain the normalized projection coordinates of the hand joint points;

[0022] Based on the normalized projection coordinates, perform an affine transformation on the image of the first hand region in the current frame image to generate a corrected first hand image.

[0023] Combined with the first aspect, in a possible design, after generating the corrected first hand image, the method further includes:

[0024] Obtain the normalized coordinates of the hand joint points in the previous frame image;

[0025] Input the normalized coordinates of the hand joint points in the previous frame image and the first hand image into the hand joint point detection model for inference to obtain the recognition result, where the recognition result includes the heat map of the hand joint points and the relative distance heat map in the current frame image.

[0026] Combined with the first aspect, in a possible design, after obtaining the recognition result, the method further includes:

[0027] Construct an objective function, which is equal to the sum of the projection error and distance error terms and the smoothness constraint term, where the projection error and distance error terms are obtained from the sum of the projection error values of 26 hand joint points and the distance error values of 26 hand joint points; the projection error value of each hand joint point is equal to the square of the norm of the difference between the normalized pixel position in the heat map corresponding to the hand joint point and the pixel position after the three-dimensional coordinates are projected by the camera, and the distance error value of each hand joint point is equal to the square of the difference between the norm of the difference between the three-dimensional coordinates of the hand joint point and the three-dimensional coordinates of the middle finger proximal metacarpal joint divided by the palm distance and the normalized relative distance of the hand joint point in the relative distance heat map;

[0028] The smoothness constraint term is obtained from the sum of the squares of the norms of the differences between the three-dimensional coordinates of each hand joint point in the t-th frame and the three-dimensional coordinates in the (t - 1)-th frame among 26 hand joint points.

[0029] The expression of the objective function is E(P)=E t (P)+E s (P), and the expression of E t (P) is:

[0030] , in the expression, E t (P) is the projection error and distance error term, u i is the normalized pixel position of the i-th hand joint point in the heat map, p i is the three-dimensional coordinates of the i-th hand joint point, Π is the camera projection model, p midproxi is the three-dimensional coordinates of the middle finger proximal metacarpal joint, I h is the palm distance, I i is the normalized relative distance of the i-th hand joint point in the relative distance heat map, and the value range is [-1, 1];

[0031] E s (P) is expressed as:

[0032] , in the expression, Es (P) is the smoothness constraint term, is the three-dimensional coordinate of the i-th hand joint point in the (t - 1)-th frame, is the three-dimensional coordinate of the i-th hand joint point in the t-th frame;

[0033] Take the result of iteratively optimizing the objective function as the three-dimensional coordinate of the hand joint point optimized based on the recognition result in the current frame image.

[0034] Combined with the first aspect, in a possible design, selecting the initial center position of the hand from the first hand bounding box and calculating the first transformation matrix in the direction from the initial center position to the optical center of the image includes:

[0035] Take the position of the proximal phalangeal joint of the middle finger in the first hand bounding box as the initial center position, and calculate the first transformation matrix in the direction from the initial center position to the optical center of the image.

[0036] Combined with the first aspect, in a possible design, tracking and obtaining the first hand region of the current frame image in the binocular hand video includes:

[0037] Obtain the current frame image in the binocular hand video;

[0038] If the previous frame image of the current frame image tracks the hand, obtain the first hand region in the current frame image, where the first hand region is the hand region predicted in the current frame image based on the positions of the hand joint points in the previous frame image;

[0039] The method further includes:

[0040] If the previous frame image of the current frame image does not track the hand, input the current frame image into the hand detection model for inference to obtain the hand center point position and the second hand region in the current frame image;

[0041] Calculate the maximum angle among the angles between the hand center point position and the four sides of the second hand region to obtain the second transformation matrix;

[0042] Project the second hand region through binocular disparity to obtain the second hand bounding box of the second hand region in three-dimensional space;

[0043] Based on the second transformation matrix, correct the positions of the hand joint points in the second hand bounding box and obtain the second direction vector of the corrected hand center point in three-dimensional space;

[0044] Obtain the second projection direction according to the product of the second transformation matrix and the second direction vector;

[0045] Calculate the projection radius based on the maximum angle, and perform projection correction on the image of the second hand region in the current frame image through the projection radius and the second projection direction to generate a corrected second hand image.

[0046] Combined with the first aspect, in a possible design, projecting the hand region through binocular disparity includes:

[0047] Calculate the overlap degree of the hand region corresponding to the left hand and the hand region corresponding to the right hand in the current frame image;

[0048] When the overlap degree is less than the threshold, project the hand region corresponding to the left hand and the hand region corresponding to the right hand through binocular disparity;

[0049] When the overlap degree is greater than or equal to the threshold and one of the left hand and the right hand is tracked to the hand in the current frame image, project the hand region corresponding to the other hand through binocular disparity;

[0050] Wherein, the hand region is the first hand region or the second hand region.

[0051] In a second aspect, an embodiment of the present application provides a hand gesture recognition device, and the device includes:

[0052] A projection module, configured to track and obtain the first hand region of the current frame image in the binocular video of the hand, and project the first hand region through binocular disparity to obtain a first hand bounding box of the first hand region in three-dimensional space;

[0053] A matrix calculation module, configured to select an initial center position of the hand from the first hand bounding box and calculate a first transformation matrix of the direction from the initial center position to the optical center of the image;

[0054] A projection direction calculation module, configured to perform multiple iterative optimizations on the first transformation matrix and use the iteratively optimized first transformation matrix as the first projection direction;

[0055] An image correction module, configured to perform projection correction on the image of the first hand region in the current frame image according to the first projection direction to generate a corrected first hand image, wherein the first hand image is used as input to a pre-trained hand joint point detection model for hand joint point recognition to obtain a recognition result.

[0056] In a third aspect, an embodiment of the present application provides a head-mounted display device, including at least one processor, a memory, a camera module, and a display module, and the at least one processor is coupled to the memory for reading and executing instructions in the memory to execute the method of the first aspect and its possible design manners.

[0057] Fourthly, an embodiment of the present application provides a storage medium, in which a computer program is stored. The computer program is configured to execute the method according to the first aspect and its possible design manners when running.

[0058] Compared with the prior art, a hand gesture recognition method, a head-mounted display device and a storage medium provided by an embodiment of the present application. The method can be applied to the head-mounted display device and includes: tracking and obtaining a first hand region of a current frame image in a binocular video of a hand, projecting the first hand region through binocular disparity to obtain a first hand bounding box of the first hand region in a three-dimensional space, then selecting an initial center position of the hand from the first hand bounding box, and calculating a first transformation matrix in the direction from the initial center position to the optical center of the image. Performing multiple iterative optimizations on the first transformation matrix, and using the iteratively optimized first transformation matrix as the first projection direction. Then, according to the first projection direction, performing projection correction on the image of the first hand region in the current frame image to generate a corrected first hand image. The first hand image is used to be input into a pre-trained hand joint point detection model for hand joint point recognition to obtain a recognition result. The present application projects the tracked hand joint points into a three-dimensional space and performs multiple iterative optimizations on the first transformation matrix, so as to gradually adjust the line of sight of the current frame image to be parallel to the direction of the optical center of the image. In this way, compared with the original current frame image, the hand posture in the corrected first hand image is standardized. Therefore, the first hand image obtained by the present application is beneficial for the model to more accurately infer the position and distance of the hand joint points, thereby improving the accuracy of subsequent hand gesture recognition.

[0059] Details of one or more embodiments of the present application are set forth in the following drawings and description, so that other features, objects and advantages of the present application will become more comprehensible. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0061] Figure 1 shows an application environment diagram of a hand gesture recognition method provided by an embodiment of the present application;

[0062] Figure 2 shows a flowchart of a hand gesture recognition method provided by an embodiment of the present application;

[0063] Figure 3 shows a flowchart of a method for validating the effectiveness of a box selection result provided by an embodiment of the present application;

[0064] Figure 4Shows a flowchart of another hand gesture recognition method provided by an embodiment of the present application;

[0065] Figure 5 Shows a flowchart of another method for validating the effectiveness of a box selection result provided by an embodiment of the present application;

[0066] Figure 6 Shows a schematic diagram of the hardware structure of a hand gesture recognition device provided by an embodiment of the present application;

[0067] Figure 7 Shows a block diagram of the hardware structure of a head-mounted display device provided by an embodiment of the present application. Detailed implementation manners

[0068] For a clearer understanding of the purpose, technical solution, and advantages of the present application, the present application will be described and explained below with reference to the accompanying drawings and embodiments.

[0069] Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the general meaning understood by those with ordinary skills in the technical field to which the present application belongs. In the present application, words such as "a", "one", "a kind of", "the", "these", etc. do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "comprising", "having" and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "coupled", etc. involved in the present application do not limit to physical or mechanical connections, but may include electrical connections, whether directly or indirectly connected. The "plurality" involved in the present application means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0070] 3D gesture estimation (or hand pose recognition) is crucial in the AR / VR field as it provides a natural and intuitive way of interaction, greatly enhancing the user experience and immersion. Users can directly control virtual objects through gestures without relying on traditional controllers, enabling more flexible and precise operations, especially suitable for tasks that require high-precision input. Additionally, it supports complex human-computer interactions, adapts to diverse application scenarios such as education, medical, and entertainment, simplifies hardware requirements, and makes the interaction more free.

[0071] 3D gesture estimation also promotes social interactions, allowing users to have rich non-verbal communications in a multi-person environment, such as shaking hands or hugging, while providing a concealed control method, enhancing security and privacy protection for use in public places. With technological advancements, gesture recognition has become more accurate and reliable, driving the development of innovative applications such as gesture-driven games and teaching software, injecting new vitality into the AR / VR industry. In short, 3D gesture estimation is one of the key technologies for realizing immersive AR / VR experiences.

[0072] In a human-computer interaction scenario, it is necessary to obtain the three-dimensional coordinate information of hand joint points in order to use this information to distinguish the hand poses of users, including clicking, swiping, grasping, stroking, flipping, making a fist, applauding, etc., so as to generate various interactions with virtual objects. Therefore, obtaining the three-dimensional coordinates of hand joint points is the core of hand pose recognition.

[0073] Currently, binocular images of the hand can be obtained through a binocular camera, the two-dimensional coordinates of hand joint points can be recognized from the binocular images, and then the three-dimensional coordinates of the hand can be obtained through camera projection. Since obtaining the three-dimensional coordinates based on a binocular camera is achieved through three-dimensional reconstruction, when the image is distorted, the obtained three-dimensional coordinates are not accurate enough, and there are serious defects in gesture interaction scenarios with high precision requirements.

[0074] In view of this, the embodiments of the present application provide a hand pose recognition method, which corrects the projection of the image of the first hand region in the current frame image by iteratively projecting the direction multiple times, so that an image taken with an inclined observation line of sight is corrected into an image with a line of sight perpendicular to the center. In this way, the hand image fed into the model for inference overcomes the influence of distortion, and then the positions and relative distances of the hand joint points obtained by model inference will be more accurate.

[0075] The solution of the present application can be applied in human-computer interaction scenarios, such as the human-computer interaction of a terminal and the human-computer interaction scenario of a vehicle-mounted system. Figure 1 The application environment diagram of a hand pose recognition method provided by the embodiments of the present application is shown, as Figure 1As shown, the terminal 102 communicates with the server 104 via a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or on other network servers. The terminal 102 can execute the method provided in the embodiments of the present application. Or the server 104 executes the method and returns the recognition result to the terminal 102. Or the server 104 trains the tracking model, the hand detection model, and the hand joint point detection model. After the training is completed, the server 104 sends the models to the terminal 102, and the terminal performs inference based on the deployed models.

[0076] Among them, the terminal can specifically include a virtual reality headset (VR headset), a head-mounted display device, an electronic display screen, and a mixed reality (MR) device, etc. Among them, the MR device can be an MR glasses, an MR helmet, an MR camera, etc. The head-mounted display device includes augmented reality glasses (AR glasses), MR glasses, etc. The vehicle-mounted system can be a vehicle-mounted chip, a vehicle-mounted device (such as a car computer, a vehicle-mounted computer, etc. with the ability to construct an augmented reality space).

[0077] Figure 2 The flowchart of a hand gesture recognition method provided by the embodiments of the present application is shown. Figure 2 The method shown can be applied to the scenario where the user wears a head-mounted display device and performs gesture interaction with the head-mounted display device. The head-mounted display device can be AR glasses. The method includes steps S201 to S204.

[0078] Step S201, track and obtain the first hand region of the current frame image in the binocular video of the hand, and project the first hand region through binocular disparity to obtain the first hand bounding box of the first hand region in the three-dimensional space.

[0079] Among them, the binocular video of the hand refers to the hand motion video captured synchronously by the left and right cameras. Specifically, the head-mounted display device contains two cameras (or binocular cameras) that capture the hand from different perspectives. The two cameras collect images in real time and synchronously to simulate the binocular observation of an object by a person.

[0080] The current frame image includes the left-eye image and the right-eye image of the current frame. Among them, the left-eye image (or the right-eye image) may contain the left hand and the right hand, or only contain one hand, or not contain a hand.

[0081] In this step, a tracking algorithm is used to obtain the first hand region. Specifically, it is determined whether hand joint points are recognized in the previous frame image of the current frame image. If hand joint points are recognized in the previous frame image, it means that the hand is tracked. Then, based on the positions of the hand joint points in the previous frame image, the positions of the hand joint points in the current frame image are predicted.

[0082] It can be understood that since there is no restrictive requirement on the number of hands in the binocular images, in this step, it is divided into four cases: determining whether the left hand joint points are recognized in the left-eye image of the previous frame, determining whether the right hand joint points are recognized in the left-eye image of the previous frame, determining whether the left hand joint points are recognized in the right-eye image of the previous frame, and determining whether the right hand joint points are recognized in the right-eye image of the previous frame.

[0083] In the four cases, if hand joint points of any hand are recognized in any one-eye image of the previous frame, the tracking calculation method is used to determine the possible positions of the hand in the current frame image of this eye image. Exemplarily, if the hand joint points of the left hand and the right hand are recognized in the left-eye image of the previous frame, the tracking calculation method is used to determine the possible positions of the left hand and the right hand in the left-eye image of the current frame respectively.

[0084] The specific calculation method is as follows:

[0085] Record the hand joint point position at time t - 1 as joints t-1 The time stamp at time t - 1 is ts t-1 The hand joint point position at time t is joint st The time stamp at time t is t st The time stamp to be predicted is t s+1 α is a coefficient to adjust the tracking prediction effect. Then, the hand joint point position in the predicted next frame image is:

[0086] ,

[0087] Through the left and right eye camera back-projection models, project the positions of the predicted hand joint points onto the left and right eye images to obtain:

[0088] 。

[0089] Among them, is the left eye camera back-projection model, is the 3D coordinate of the left hand joint point, is the right eye camera back-projection model, is the 3D coordinate of the right hand joint point.

[0090] Taking the left hand as an example (the right hand is the same): Calculate the pixel average value of all 3D hand joint points:

[0091] 。

[0092] Take the coordinate point corresponding to the pixel average value as the center point for tracking the hand area. At the same time, calculate the maximum value of each 3D hand joint point to :

[0093] 。

[0094] The side length of the hand area bounding box is:

[0095]

[0096] The first hand area can be obtained through the side length of the hand bounding area and the center point of the hand area.

[0097] In this step, the position information of the hand joints in the current frame image is predicted by using the position information of the hand joints in the previous frame image through mathematical modeling, avoiding complex model inference for each frame image. This can reduce the computational load, have low requirements for hardware resources, and be more suitable for AR / VR devices.

[0098] After obtaining the first hand area, perform a stereo projection on it. Specifically, the binocular camera captures two images of the hand from different perspectives, and the pixel position difference of the same hand area in these two images is called binocular disparity. Based on the binocular disparity, the distance between the left and right cameras, and the focal length, calculate the Z-axis coordinate of the hand, and then based on the Z-axis coordinate and the pixel coordinates of the hand in the image, calculate the coordinates of the first hand area in the three-dimensional space through the camera internal parameter matrix. In this step, only the images of the local area (i.e., the first hand area) need to be projected, so the computational load of 3D reconstruction is reduced.

[0099] In some embodiments, the method further includes performing cross-validation on the bounding results to identify the validity of the bounding results. For hands with invalid bounding, no projection is performed to save computational resources.

[0100] Specifically, the method includes steps S301 to S303 as Figure 3 shown.

[0101] Step S301, calculate the overlap degree of the first hand area corresponding to the left hand and the first hand area corresponding to the right hand in the current frame image.

[0102] The current frame image includes a left-eye image and a right-eye image. In this step, the overlap degrees of the first hand regions corresponding to the left hand and the right hand in the left-eye image and in the right-eye image are calculated. Among them, the overlap degree (IoU) can be the ratio of the intersection area to the union area of the first hand regions corresponding to the left and right hands. The intersection area refers to the pixel area of the overlapping part of the two first hand regions, and the union area refers to the difference between the total coverage area of the two first hand regions and the intersection area.

[0103] Step S302: When the overlap degree is less than the threshold, project the first hand regions corresponding to the left hand and the right hand through binocular disparity.

[0104] The threshold is a preset critical value, and the threshold is a number in the range of [0, 1]. For example, the threshold can be 0.3, 0.4, 0.8, etc.

[0105] The threshold is used to determine whether the hand regions significantly overlap. In this step, if the IoU is less than the threshold, it is considered that the two hands are independent (or non-overlapping). Then, after separately stereographically projecting the regions of the two hands in each eye image, the following step S202 is executed.

[0106] In some embodiments, in order to adapt to complex hand interaction scenarios, a dynamic threshold can be adopted to adapt to the dynamic transformation of frequent switching of the two hands. For example, when the hand movements are frequent, the threshold is reduced to improve sensitivity; when the movements are stable, the threshold is increased to reduce misjudgment.

[0107] Step S303: When the overlap degree is greater than or equal to the threshold, and one of the left hand and the right hand has its hand tracked in the current frame image, project the first hand region corresponding to the other hand through binocular disparity.

[0108] In this step, if the IoU is greater than or equal to the threshold, it is considered that the two hands may overlap or cross. Then, when one hand has been stably located in the current frame image (any eye image) through the tracking algorithm, only the first hand region of the other hand is stereographically projected by binocular disparity, avoiding repeated calculation of the stably tracked hand, which can reduce resource consumption.

[0109] In addition, when the overlap degree is greater than or equal to the threshold and neither of the two hands has its hand tracked in the current frame image, the calculation of this frame image is abandoned.

[0110] Through steps S301 to S303, the valid first hand regions selected by bounding boxes are screened out. For the valid first hand regions selected by bounding boxes, after stereographic projection, the following step S202 is executed on the first hand bounding box.

[0111] It should be noted that it is assumed here that the left hand is valid, and the corresponding hand may be seen in the left-eye image or the right-eye image. Therefore, the following processing flow will be performed 1 to 2 times for each hand, depending on the possible valid results of hand region detection. The following describes the process of image correction once, and the others are similar, with a maximum of four times possible.

[0112] Step S202: Select the initial center position of the hand from the first hand bounding box, and calculate the first transformation matrix in the direction from the initial center position to the optical center of the image.

[0113] In this step, the initial center position refers to the position coordinates of the hand center point before iterative optimization. In some embodiments, the position of the proximal palmar joint of the middle finger is used as the initial center position.

[0114] Among them, the optical center direction of the image is (0, 0, -1). In this step, the transformation relationship from the local point to (0, 0, -1) is calculated. Subsequently, based on this transformation relationship, all points of the hand can be mapped to (0, 0, -1). For specific reference, see step S203.

[0115] Step S203: Perform multiple iterative optimizations on the first transformation matrix, and use the iteratively optimized first transformation matrix as the first projection direction.

[0116] The iterative process of the first transformation matrix is as follows: Based on the first transformation matrix, correct the positions of the hand joint points in the first hand bounding box, and obtain the first direction vector of the corrected hand center point in the three-dimensional space; take the product of the first transformation matrix and the first direction vector as the first transformation matrix for the next iterative optimization, and repeat the above iterative process until the first transformation matrix converges and stop the iteration.

[0117] Correction means mapping the positions of the hand joint points from the original positions to the transformed positions with the optical axis direction of the image as the observation line of sight through the first transformation matrix. That is, each hand joint point is rotated to the [0, 0, -1] direction through the inverse transformation of the first transformation matrix (T_mid_o) to obtain a new position in the three-dimensional space. The purpose of correcting the positions of the 3D hand joint points is to standardize the originally diverse hand postures so that the palm of the hand is vertically oriented towards the observation line of sight, which is beneficial for subsequent model inference.

[0118] Since the image will be distorted during projection, there is an error between the initial center position based on which the current correction is made and the actual center position of the hand. Therefore, in this step, the corrected hand coordinate points are projected onto the equatorial plane. Among them, the hand is represented by 26 degrees of freedom for the three-dimensional coordinates of 21 hand joint points. In the embodiments of the present application, for the sake of convenience of expression, it is simply referred to as 26 joint points. In the equatorial plane coordinates of the 26 joint points, the average value of the maximum coordinate point and the minimum coordinate point is taken to obtain the target center position of the 26 joint points on the equatorial plane. The target center position is back-projected through the equatorial plane to obtain the first direction vector V of the corrected hand center point in the three-dimensional space. dir The corrected hand center point approaches the actual center position step by step through multiple iterations. Then the direction of the new center point in the original proximal metacarpal space of the middle finger is dir_mid_new = T_mid_o × V dir .

[0119] Then use dir_mid_new to replace the first transformation matrix, and repeat step S203 until the direction converges. Take the final dir_mid_new as the first projection direction. The purpose of such multiple iterations is to gradually adjust the projection direction so that the projection direction is parallel to the image optical center direction, that is, looking straight at the image center, which can reduce the influence of image distortion caused by the perspective tilt.

[0120] In step S203, each iteration includes position correction, equatorial plane projection, and matrix update. The first transformation matrix (T_mid_o) is adjusted through multiple iterations to gradually correct the position deviation of the hand joint points in the three-dimensional space. Through gradient descent optimization, gradually approach the true pose of the hand in the three-dimensional space and reduce the projection distortion.

[0121] Step S204: According to the first projection direction, perform projection correction on the image of the first hand region in the current frame image to generate a corrected first hand image, where the first hand image is used as input to a pre-trained hand joint point detection model for hand joint point recognition to obtain a recognition result.

[0122] This step performs geometric correction on the hand image in the current frame based on the finally optimized projection direction to generate a distortion-free two-dimensional image. Among them, the geometric correction can be an affine transformation. In the geometrically corrected image, the palm region is stretched or rotated to a frontal view angle, so the image distortion caused by the perspective change is eliminated, which can improve the recognition accuracy of the hand joint point detection model, that is, it can make the coordinates and distances of the hand joint points obtained by reasoning more accurate.

[0123] In some of these embodiments, step S204 includes: projecting the hand joint points in the first hand bounding box onto the equatorial plane according to the first projection direction to obtain the normalized projection coordinates of the hand joint points; and performing an affine transformation on the image of the first hand region in the current frame image based on the normalized projection coordinates to generate the corrected first hand image.

[0124] In this embodiment, in addition to unifying the observation perspective of the hand, the coordinates of the hand joint points are also normalized to unify the size of the hand. Specifically, the equatorial plane projection coordinates of the joint points within the projection radius are mapped from [-stereographic_radius, stereographic_radius] to [-1, 1]. If it is the right hand, the X coordinate needs to be reversed, that is, a left-right hand mirror image is made. Here, stereographic_radius refers to the projection radius.

[0125] Step S204 corrects and crops the image of the hand region position (i.e., the first hand region) from the original image (i.e., the current frame image) through binocular stereovision correction to obtain the image (i.e., the first hand image) that meets the size requirements and is corrected and cropped for input to the model.

[0126] In the embodiments described in the above steps S201 to S204, the hand joint points obtained by tracking are projected into three-dimensional space, and the first transformation matrix is optimized through multiple iterations to gradually adjust the line of sight of the current frame image to be parallel to the direction of the image optical center. In this way, compared with the original current frame image, the corrected first hand image has a standardized hand posture. Therefore, the first hand image obtained by using this embodiment is beneficial for the model to more accurately infer the positions and distances of the hand joint points, and thus the subsequent recognition of hand postures is more accurate. It should be noted that this embodiment is not simply a 2D rotation of a planar image. This type of 2D rotation correction method is only applicable to simple gestures with the palm extended and cannot correct 3D postures. This embodiment eliminates projection errors through iterative optimization and standardizes the hand posture to the [0, 0, -1] direction through 3D rotation transformation, so as to achieve affine transformation of complex gestures such as finger bending, occlusion, and rotation, improve the robustness and accuracy of the model's posture recognition for complex gestures, and is more suitable for complex gesture interaction scenarios between users and AR / VR devices.

[0127] In some of these embodiments, after obtaining the first hand image, it can be used for model inference.

[0128] Specifically, after step S204, the method provided in this embodiment further includes:

[0129] Obtain the normalized coordinates of the hand joint points in the previous frame image, and input the normalized coordinates of the hand joint points in the previous frame image and the first hand image into the hand joint point detection model for inference to obtain the recognition result, where the recognition result includes the heat map of the hand joint points and the relative distance heat map in the current frame image.

[0130] It can be understood that the normalized coordinates of the hand joint points in the input previous frame image provide a reference for the output result of the hand joint point detection model. The first hand image is a standardized image with a fixed size after correction and cropping. Both can improve the accuracy of the positions and relative distances of the hand joint points in the recognition result output by the model.

[0131] In this embodiment, the hand joint point detection model is pre-trained, and its training process can be carried out in the server, that is, the server sends the trained hand joint point detection model to the head-mounted display device, and the head-mounted display device directly runs the hand joint point detection model. Or, the hand joint point detection model is trained and tested in the head-mounted display device, and this application does not limit this.

[0132] The above is described by taking tracking to obtain hand joint points as an example. In other embodiments, the hand was not tracked in the previous frame of the current frame image. For the correction method of the current frame image, refer to the following description.

[0133] Figure 4 The flowchart of another hand gesture recognition method provided by the embodiment of the present application is shown. Figure 4 The method shown can be applied to the scenario of gesture interaction with the head-mounted display device after the user wears the head-mounted display device. The head-mounted display device can be an AR glasses. The method includes steps S401 to step S407.

[0134] Step S401, obtain the current frame image in the binocular hand video.

[0135] Step S402, if the hand is not tracked in the previous frame image of the current frame image, input the current frame image into the hand detection model for inference to obtain the hand center point position and the second hand area of the current frame image.

[0136] This step is used to locate the hand through the model when the hand tracking fails (such as the hand moves quickly or is blocked) or when detecting the first frame image of the video, to prevent the process from interrupting.

[0137] In this step, it is determined whether hand joint points are recognized in the previous frame image of the current frame image. If no hand joint points are recognized in the previous frame image, it means that the hand is not tracked. Then, through a pre-trained hand detection model, the current frame image is inferred to identify the region where hand joint points exist, and the position of the hand center point and the second hand region are output.

[0138] The hand detection model refers to an object detection model based on a convolutional neural network. Its input is a binocular image that meets the resolution requirements, and its output is information on whether there are left and right hands in each image, the pixel positions of the center points of the left and right hands (i.e., the position of the hand center point), and the pixel sizes of the left and right hands (i.e., the second hand region). The second hand region refers to a rectangular region with the position of the hand center point as the reference, composed of four sides. The hand detection model can be networks such as YOLO, Faster R-CNN, etc. As an example, the backbone network of the hand detection model uses MobileNet.

[0139] It can be understood that since there is no restrictive requirement on the number of hands in the binocular image, in this step, it is divided into four cases: determining whether the left hand joint points are recognized in the left-eye image of the previous frame, determining whether the right hand joint points are recognized in the left-eye image of the previous frame, determining whether the left hand joint points are recognized in the right-eye image of the previous frame, and determining whether the right hand joint points are recognized in the right-eye image of the previous frame. For specific reference, please refer to the explanation in step S201 here, and this application will not repeat it here.

[0140] After step S402, the method provided in this embodiment further includes performing cross-validation on the box selection result to identify the validity of the recognition result. For the hand with invalid box selection, the following steps S403 to S407 are not executed to save computing resources.

[0141] Specifically, the method includes steps S501 to S503 as shown in Figure 5 Figure.

[0142] Step S501: Calculate the overlap degree of the second hand region corresponding to the left hand and the second hand region corresponding to the right hand in the current frame image.

[0143] Step S502: When the overlap degree is less than the threshold, project the second hand region corresponding to the left hand and the second hand region corresponding to the right hand through binocular disparity.

[0144] Step S503: When the overlap degree is greater than or equal to the threshold and one of the left hand and the right hand is tracked in the current frame image, project the second hand region corresponding to the other hand through binocular disparity.

[0145] For the explanations of steps S501 to S503, reference can be made to steps S301 to S303 above, and this application will not repeat them here.

[0146] Through steps S501 to S503, the second hand region with valid bounding is screened out, and for the second hand region with valid bounding, step S403 below is executed.

[0147] Step S403: Calculate the maximum angle among the angles between the position of the hand center point and the four sides of the second hand region to obtain the second transformation matrix.

[0148] The four sides of the second hand region respectively form angles with the center point connection line, and the largest of these angles is selected as the stretching direction of the hand in the plane. This can make the hand stretch at the maximum angle in space, that is, the posture of the hand can be observed to the greatest extent. Therefore, the maximum angle is used as the second transformation matrix for transformation.

[0149] The second transformation matrix is an affine transformation matrix, representing the linear transformation relationship of the second hand region rotating from the original direction to the maximum angle direction.

[0150] Step S404: Project the second hand region through binocular disparity to obtain the second hand bounding box of the second hand region in three-dimensional space.

[0151] Using the disparity information of the binocular camera, the two-dimensional hand region is mapped from the image plane to three-dimensional space to determine the position and size of the hand in three-dimensional space.

[0152] This step can refer to step S201 above and will not be elaborated here.

[0153] Step S405: Based on the second transformation matrix, correct the positions of the hand joint points in the second hand bounding box, and obtain the second direction vector of the corrected hand center point in three-dimensional space.

[0154] In this step, the three-dimensional coordinates of each hand joint point are rotated and translated through the second transformation matrix to correct its 3D position. For details, reference can be made to the introduction of step S203.

[0155] Step S406: Obtain the second projection direction according to the product of the second transformation matrix and the second direction vector.

[0156] Among them, the second transformation matrix can be represented by T_mid_o, and the second direction vector can be represented by V dir Then the second projection direction is expressed as: dir_mid_new = T_mid_o × V dir .

[0157] Step S407: Calculate the projection radius based on the maximum angle, and perform projection correction on the image of the second hand region in the current frame image through the projection radius and the second projection direction to generate a corrected second hand image.

[0158] Specifically, the following formula is used to calculate the projection radius: sin(angle) / (1 - cos(angle)). Where angle represents the maximum angle.

[0159] The projection direction and the projection radius are used to correct the observation line of sight of the hand in the image, so that an image with an oblique line of sight is corrected into an image with a perpendicular line of sight. This can eliminate the hand pose difference, reduce the projection error, and thus reduce the influence of image distortion caused by the perspective change on the model inference result.

[0160] Through steps S401 to S407, the embodiments of the present application provide another hand pose recognition method. When the tracking fails, the hand key points of the current frame image are obtained through model inference to ensure the coherence of hand pose recognition. At the same time, the image of the second hand region obtained by model inference is also subjected to image correction processing to reduce image distortion. Therefore, the subsequent inference result of the second hand image will be more accurate.

[0161] In some of these embodiments, after obtaining the second hand image, it can be used for model inference.

[0162] Specifically, after step S407, the method provided by this embodiment further includes:

[0163] Input the second hand image into the hand joint point detection model for inference to obtain the recognition result, where the recognition result includes the heat map of the hand joint points and the relative distance heat map in the current frame image.

[0164] It can be understood that the second hand image is a standardized image with a fixed size after correction and cropping. Therefore, the positions and relative distances of the hand joint points in the recognition result output by the model will be more accurate.

[0165] In this embodiment, the hand joint point detection model is pre-trained, and its training process can be carried out on the server. That is, the server sends the trained hand joint point detection model to the head-mounted display device, and the head-mounted display device directly runs the hand joint point detection model. Or, the hand joint point detection model is trained and tested in the head-mounted display device. The present application does not limit this.

[0166] The above text introduced how to correct and crop hand images, including correcting and cropping the images obtained by the hand tracking method and the images obtained by the hand detection model. In addition, it also introduced how to send the corrected hand images into the hand joint point detection model for inference. Next, a method for calculating the 3D coordinates of hand joint points provided by the embodiments of the present application will be introduced, which can perform non-linear optimization on the results obtained by model inference, making the hand joint points jitter less and the joint point positions more accurate.

[0167] Specifically, the method includes: inputting the normalized coordinates of the hand joint points of the previous frame image and the hand image into the hand joint point detection model for inference to obtain the recognition result, where the recognition result includes the heat map of the hand joint points and the relative distance heat map in the current frame image. The hand image is the first hand image or the second hand image.

[0168] It can be understood that for the case where the hand is tracked in the previous frame image, the input item in the hand joint point detection model includes the normalized coordinates of the hand joint points of the previous frame image. For the case where the hand is not tracked in the previous frame image, the information of the previous frame image does not need to be input.

[0169] The hand image is the first hand image or the second hand image. The hand image refers to the current frame hand region image after preprocessing (such as affine transformation, cropping, gray normalization, etc.), and the size is fixed to meet the model input requirements. Among them, the purpose of normalizing the hand joint points is to eliminate the scale difference.

[0170] The hand key point detection model includes a heat map branch and a relative distance branch, where the heat map branch outputs the coordinates of each hand joint point. Therefore, the heat map of the hand joint points refers to the possible positions of each hand joint point on the binocular image. The relative distance branch outputs the distance offsets of each hand joint point relative to the hand center (i.e., the proximal metacarpal bone of the middle finger). Therefore, the relative distance heat map refers to the relative distance of each hand joint point from the hand center in the Z-axis direction. Among them, the relative distance can be normalized to a value within the range of [-1, 1].

[0171] After that, the method further includes: constructing an objective function, which is equal to the sum of a projection error and a distance error term and a smoothness constraint term, where the projection error and the distance error term are obtained from the sum of the projection error values of 26 hand joint points and the distance error values of 26 hand joint points; the projection error value of each hand joint point is equal to the square of the norm of the difference between the normalized pixel position in the heat map corresponding to the hand joint point and the pixel position after the three-dimensional coordinates are projected by the camera, and the distance error value of each hand joint point is equal to the square of the difference between the norm of the difference between the three-dimensional coordinates of the hand joint point and the three-dimensional coordinates of the middle finger proximal metacarpal joint divided by the palm distance and the normalized relative distance of the hand joint point in the relative distance heat map;

[0172] The smoothness constraint term is obtained from the sum of the squares of the norms of the differences between the three-dimensional coordinates of each of the 26 hand joint points in the t-th frame and the three-dimensional coordinates in the (t - 1)-th frame.

[0173] The expression of the objective function is E(P)=E t (P)+E s (P).

[0174] E t (P) is the projection error and the distance error term, and the expression of E t (P) is:

[0175] , in the expression, assuming u i is [sg_x, sg_y]u i is the normalized pixel position of the i-th hand joint point in the heat map of the hand joint point, p i is the three-dimensional coordinates of the i-th hand joint point, Π is the camera projection model, p midproxi is the three-dimensional coordinates of the middle finger proximal metacarpal joint, I h is the palm distance, I i is the normalized relative distance of the i-th hand joint point in the relative distance heat map.

[0176] The error projection ensures that the projection of the 3D joint points on the image plane is consistent with the actual observations.

[0177] The distance error is to ensure that the relative distance of the heat map of the 3D joint points is consistent with the information of the true relative distance, that is, the relative distance output by the model is as close as possible to the relative distance calculated by the joint points. Among them, It is the relative distance, which is represented by the ratio of the difference between the three-dimensional coordinates of the i-th hand joint point and the three-dimensional coordinates of the proximal metacarpal bone of the middle finger to the palm distance. This value reflects the relative distance between the i-th hand joint point and the center of the hand. This formula takes into account the different palm sizes of different users, so the relative distance will be affected by the palm size. Therefore, the palm distance is introduced, and in this way, a better optimization effect can be achieved for the hand joint points of each user.

[0178] E s (P) is the smoothness constraint term, E s (P) is expressed as:

[0179] , in the expression, is the three-dimensional coordinate of the i-th hand joint point in the (t - 1)-th frame, is the three-dimensional coordinate of the i-th hand joint point in the t-th frame.

[0180] Among them, is to ensure that the changes of the 3D joint points between adjacent frames are smooth.

[0181] Aiming at minimizing the objective function, iterative optimization is carried out, and the result of the iterative optimization of the objective function is used as the 3D coordinates of the hand joint points optimized based on the recognition result in the current frame image. The 3D coordinates of each optimized hand joint point are more accurate, so the jitter of the fingers is smaller, and the recognition effect of the hand gesture is better.

[0182] The following uses a specific example to illustrate the method provided by the embodiment of the present application.

[0183] Optional step 1: Train a hand detection model that can be used or use a trained model. The input of the model is an image corresponding to the resolution requirement, and the output of the model is the information of whether there are left and right hands in each image, the pixel positions of the center points of the left and right hands, and the pixel sizes of the sizes of the left and right hands.

[0184] Optional step 2: Train a hand joint point detection model, or use a pre-trained model. The input parameter requirements of the model are the cropped and corrected hand image, whether to use the joint point positions of the previous frame, and the normalized joint point positions of the previous frame (if using the joint point positions of the previous frame). The model output is the heat map of the joint points of the current frame, the heat map of the relative distance, the inference valid flag bit, and the bending degree of each finger.

[0185] After obtaining the hand detection model and the hand joint point detection model, execute step 3.

[0186] Step 3: Obtain the possible position information and the presence or absence of hand information of the hand position in the current frame image in the left and right eye images.

[0187] Specifically, according to whether the hand is tracked in the previous frame image, step 3 includes step 3.1 and step 3.2.

[0188] Step 3.1: The hand joint points are not correctly solved simultaneously in the previous frame image, specifically including the cases where neither hand is solved or only one hand is solved. In these cases, the left and right eye images are sent to the hand detection model for inference after preprocessing. The output results are whether there is information about the left and right hands in each image, the pixel positions of the center points of the left and right hands, and the pixel sizes of the left and right hand regions.

[0189] Step 3.2: If the hand joint points of a certain hand are correctly solved in the previous frame, then the position where the hand may appear in the current frame is determined by using the tracking calculation method. Specifically, refer to step S201 above. The bounding box area of the hand is obtained in step 3.2.

[0190] Step 3.3: Cross-validation is performed on the bounding box results in 3.1 and 3.2 respectively to identify whether the hands boxed in step 3.1 or 3.2 are valid. If a certain hand is invalid, the subsequent workflow does not process the corresponding hand. If both hands are invalid, wait for the next frame and restart step 3. Step 3.3 corresponds to steps S301 to S303 above, and corresponds to steps S501 to S503 above.

[0191] Step 4: Detect the hand joint points for the hands with valid bounding boxes and their views.

[0192] Specifically, according to whether the hand is tracked in a frame image, step 4 includes step 4.1 and step 4.2.

[0193] Step 4.1: If the current hand area (equivalent to the first hand area above) is obtained by the method of step 3.1, then the pixel positions of the four vertices on the image can be inversely calculated through the central area of the hand and the hand area size. This step is equivalent to steps S403 to S406 above.

[0194] Step 4.2: If the current hand area (equivalent to the second hand area above) is obtained by the method of step 3.2, then the transformation angle is calculated by the following method.

[0195] Step 4.2.1: Rotate the hand area to the direction of [0, 0, -1] (the direction of the optical center of the image) for projection correction. The specific method is to select the proximal palmar joint of the middle finger and calculate the coordinate transformation relationship T_mid_o from the proximal palmar joint of the middle finger to [0, 0, -1].

[0196] Step 4.2.2: Rotate each joint point to the direction of [0, 0, -1] through the inverse transformation of T_mid_o, and calculate the coordinates [sg_x, sg_y] of each joint point on the equatorial plane through equatorial plane projection.

[0197] Step 4.2.3: Obtain the center position by taking the average of the maximum coordinate point and the minimum coordinate among the equatorial plane coordinates of the 26 joint points.

[0198] Step 4.2.4: Back-project the recalculated center position through the equatorial plane to obtain the direction vector V of the center point. dir Then the direction of the new center point in the original middle finger proximal metacarpal space is dir_mid_new = T_mid_o × V. dir .

[0199] Step 4.2.5: Replace T_mid_o in 4.2.2 with dir_mid_new, and repeat the calculation process from Step 4.2.2 to Step 4.2.4 until the direction converges.

[0200] Step 4.3: Calculate the stereographic projection of the 3D joint points on the equatorial plane using the converged new direction (taking the left eye image as an example, the 3D joint points are points in the OpenXR coordinate system of the left eye optical center, represented in real physical scale).

[0201] Step 4.4: Map the equatorial plane projection coordinates of the joint points within the projection radius from [-stereographic_radius, stereographic_radius] to [-1, 1]. If it is the right hand, the X needs to be reversed, that is, this step needs to perform left - right hand mirroring.

[0202] Step 4.5: Through binocular stereo vision correction, select the image of the hand region position from the original image for correction and cropping, and obtain the corrected and cropped image of the size required for inputting into the inference model.

[0203] Step 4.6: Send it into the hand joint point detection model for inference. The input is the normalized coordinates of the hand joint points in the previous frame image and the corrected and cropped image of the current frame. What is obtained is the heat map of the joint point normalized coordinates, the heat map of the relative distance, the inference valid flag bit, and the bending degree of each finger.

[0204] Step 5: Calculate the 3D coordinates of the finger joints using the method of nonlinear optimization.

[0205] Specifically, in step 4.7, the image coordinates (heat map of joint normalized coordinates) of the hand joint points of the left or right hand under each camera and the relative distance information of the joint points (heat map of relative distance) are obtained. For the information output by the model, in step 5, the 3D coordinates of each hand joint point are calculated by means of non-linear optimization.

[0206] The optimized objective function can be written as E(P)=E t (P)+E s (P). The specific optimization process will not be elaborated here. The result of iterative optimization is the precise 3D coordinates of each finger joint of the corresponding hand.

[0207] The above has elaborated in detail on a hand gesture recognition method provided by an embodiment of the present application. Next, a hand gesture recognition device will be introduced.

[0208] Please refer to Figure 6 , Figure 6 which shows a schematic hardware structure diagram of a hand gesture recognition device provided by an embodiment of the present application. As Figure 6 shown, the device includes:

[0209] A projection module 601, configured to track and obtain a first hand region of the current frame image in the hand binocular video, and project the first hand region through binocular disparity to obtain a first hand bounding box of the first hand region in three-dimensional space.

[0210] A matrix calculation module 602, configured to select an initial center position of the hand from the first hand bounding box and calculate a first transformation matrix in the direction from the initial center position to the image optical center.

[0211] A projection direction calculation module 603, configured to perform multiple iterative optimizations on the first transformation matrix and use the iteratively optimized first transformation matrix as the first projection direction.

[0212] An image correction module 604, configured to perform projection correction on the image of the first hand region in the current frame image according to the first projection direction to generate a corrected first hand image, where the first hand image is used to be input into a pre-trained hand joint point detection model for hand joint point recognition to obtain a recognition result.

[0213] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combined form.

[0214] An embodiment of the present application further provides a head-mounted display device, such asFigure 7 As shown in Figure 7 , the head-mounted display device includes a processor 701, a memory 702, a shooting module 703, a display module 704, an interaction module 705, an IMU module 706, a gyroscope 707, a geomagnetic sensor 708, and a wireless communication module 709.

[0215] The memory 702 can be used to store computer programs, for example, software programs and modules of application software. The processor 701 executes various functional applications and data processing by running the computer programs stored in the memory 702, such as implementing the method provided in the embodiments of the present application.

[0216] The shooting module 703 is mainly responsible for shooting the picture information of the real environment, such as obtaining the binocular video of the user's hand. The shooting module can be a camera.

[0217] The display module 704 is used to display virtual images and augmented reality content, and it includes a micro-projection system and optical elements.

[0218] The interaction module 705 includes a microphone, an eye tracking sensor, a ring, a wristband, etc.

[0219] The wireless communication module 709 includes: 4G / 5G full netcom, WiFi, Bluetooth, etc.

[0220] Those of ordinary skill in the art can understand that Figure 7 the structure shown is only schematic and does not limit the structure of the MR glasses. For example, the head-mounted display device may further include more or fewer components than those shown in Figure 7 or have a different configuration from that shown in Figure 7 the figure.

[0221] Next, a simple introduction to the application scenarios of the head-mounted display device will be given.

[0222] In the gesture interaction scenario of using the head-mounted display device, the user can directly control virtual objects through gestures, such as picking up a water cup in the virtual space, waving to a digital human in the virtual space, etc., so as to improve the user's immersion.

[0223] In addition, in combination with the method provided in the above embodiments, a storage medium can also be provided in this embodiment to implement. A computer program is stored on the storage medium; when the computer program is executed by the processor, any one of the hand gesture recognition methods in the above embodiments is implemented.

[0224] It should be understood that the specific embodiments described here are only used to explain this application, rather than to limit it. According to the embodiments provided in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the protection scope of the present application.

[0225] Obviously, the accompanying drawings are only some examples or embodiments of the present application. For those of ordinary skill in the art, the present application can also be applied to other similar situations based on these drawings without creative efforts. Additionally, it can be understood that although the work done during the development process here may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be regarded as insufficient disclosure of the present application.

[0226] The term "embodiment" in the present application means that the specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.

[0227] The above embodiments only represent several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be construed as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A hand gesture recognition method, characterized in that, The method includes: Tracking and obtaining the first hand region of the current frame image in the binocular hand video, projecting the first hand region through binocular disparity to obtain the first hand bounding box of the first hand region in the three-dimensional space; Selecting the initial center position of the hand from the first hand bounding box, and calculating the first transformation matrix in the direction from the initial center position to the optical center of the image; Performing multiple iterative optimizations on the first transformation matrix, and using the iteratively optimized first transformation matrix as the first projection direction; Wherein, the iterative process of the first transformation matrix includes: based on the first transformation matrix, correcting the positions of the hand joint points in the first hand bounding box, and obtaining the first direction vector of the corrected hand center point in the three-dimensional space; taking the product of the first transformation matrix and the first direction vector as the first transformation matrix for the next iterative optimization, repeating the above iterative process until the first transformation matrix converges, and stopping the iteration; According to the first projection direction, performing projection correction on the image of the first hand region in the current frame image to generate the corrected first hand image, wherein the first hand image is used as input to a pre-trained hand joint point detection model for hand joint point recognition to obtain the recognition result.

2. The hand gesture recognition method according to claim 1, wherein The correcting the positions of the hand joint points in the first hand bounding box based on the first transformation matrix and obtaining the first direction vector of the corrected hand center point in the three-dimensional space includes: Based on the first transformation matrix, transforming the positions of the hand joint points in the first hand bounding box from the first three-dimensional positions to the second three-dimensional positions; Obtaining the position information of the projection center point of the second three-dimensional position on the equatorial plane to obtain the target center position; Back-projecting the target center position through the equatorial plane to obtain the first direction vector of the corrected hand center point in the three-dimensional space.

3. The hand gesture recognition method according to claim 1, characterized in that, The performing projection correction on the image of the first hand region in the current frame image according to the first projection direction to generate the corrected first hand image includes: According to the first projection direction, projecting the hand joint points in the first hand bounding box onto the equatorial plane to obtain the normalized projection coordinates of the hand joint points; Performing an affine transformation on the image of the first hand region in the current frame image based on the normalized projection coordinates to generate the corrected first hand image.

4. The hand gesture recognition method according to claim 1, wherein After generating the corrected first hand image, the method further includes: Obtaining the normalized coordinates of the hand joint points of the previous frame image; Inputting the normalized coordinates of the hand joint points of the previous frame image and the first hand image into the hand joint point detection model for inference to obtain the recognition result, wherein the recognition result includes the heat map and the relative distance heat map of the hand joint points in the current frame image.

5. The hand gesture recognition method according to claim 4, wherein After obtaining the recognition result, the method further includes: Construct an objective function that is equal to the sum of the projection error and the distance error term and the smoothness constraint term. The projection error and the distance error term are obtained from the sum of the projection error values of 26 hand joint points and the distance error values of 26 hand joint points. The projection error value of each hand joint point is equal to the square of the norm of the difference between the normalized pixel position in the heat map corresponding to the hand joint point and the pixel position after the 3D coordinates are projected by the camera. The distance error value of each hand joint point is equal to the square of the difference between the norm of the difference between the 3D coordinates of the hand joint point and the 3D coordinates of the middle finger proximal phalanx metacarpal joint divided by the palm distance and the normalized relative distance of the hand joint point in the relative distance heat map. The smoothness constraint term is obtained from the sum of the squares of the norms of the differences between the 3D coordinates of each of the 26 hand joint points in the current frame image and the 3D coordinates in the previous frame image of the current frame image.

6. The hand gesture recognition method according to claim 1, wherein, The tracking to obtain the first hand region of the current frame image in the binocular hand video includes: Obtain the current frame image in the binocular hand video; If the previous frame image of the current frame image tracks the hand, obtain the first hand region in the current frame image. The first hand region is the hand region in the current frame image predicted based on the positions of the hand joint points in the previous frame image. The method further includes: If the previous frame image of the current frame image does not track the hand, input the current frame image into a hand detection model for inference to obtain the hand center point position of the current frame image and the second hand region; Calculate the maximum angle among the angles between the hand center point position and the four sides of the second hand region to obtain a second transformation matrix; Project the second hand region through binocular disparity to obtain a second hand bounding box of the second hand region in 3D space; Based on the second transformation matrix, correct the positions of the hand joint points in the second hand bounding box and obtain the second direction vector of the corrected hand center point in 3D space; Obtain a second projection direction according to the product of the second transformation matrix and the second direction vector; Calculate a projection radius based on the maximum angle, and perform projection correction on the image of the second hand region in the current frame image through the projection radius and the second projection direction to generate a corrected second hand image.

7. The hand gesture recognition method according to claim 1 or claim 6, characterized in that Projecting the hand region through binocular disparity includes: Calculate the overlap degree of the hand region corresponding to the left hand and the hand region corresponding to the right hand in the current frame image; In the case where the overlap degree is less than the threshold, project the hand region corresponding to the left hand and the hand region corresponding to the right hand through binocular disparity; In the case where the overlap degree is greater than or equal to the threshold and one of the left hand and the right hand is tracked in the current frame image, project the hand region corresponding to the other hand through binocular disparity; Wherein, the hand region is the first hand region or the second hand region.

8. A head-mounted display device, characterized in that It includes at least one processor, a memory, a camera module, and a display module. The at least one processor is coupled to the memory and is used to read and execute instructions in the memory to perform the hand pose recognition method according to any one of claims 1 to 7.

9. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is configured to execute the hand gesture recognition method according to any one of claims 1 to 7 when running.

Citation Information

Patent Citations

  • Method for detecting hand joints in glasses equipment, storage medium, electronic equipment and product

    CN119672766A

  • Position estimating device, position estimating method, and program

    WO2018235923A1