Hand posture recognition method, head-mounted display device and storage medium
By projecting correction of the hand image, and using the iteratively optimized transformation matrix to eliminate image distortion, the problem of inaccurate coordinates of 3D hand joint nodes in the prior art is solved, and the accuracy of hand posture recognition is improved.
Patent Information
- Application Number
- CN202510496558.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
In the prior art, due to the lack of distortion correction of the hand images captured by the binocular camera, the predicted 3D hand joint node coordinates are inaccurate, which affects the subsequent hand posture recognition results.
By tracking and obtaining the current frame image in the binocular video of the hand, and projecting the hand area using binocular parallax, obtaining the enclosing box of the hand area in three-dimensional space. Then, the initial center position of the hand is selected, the transformation matrix from the initial center position to the center direction of the image is calculated, and the transformation matrix is optimized through multiple iterations to correct the hand image as the projection direction.
By inputting the corrected hand image into the pre-trained hand joint detection model for identification, the accuracy of the hand joint recognition results is significantly improved and the error caused by image distortion is reduced.
Smart Images

Figure CN120014714A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of human-computer interaction technology, and in particular to a hand gesture recognition method, a head-mounted display device, and a storage medium. Background Art
[0002] In the field of virtual reality (VR) or augmented reality (AR) technology, three-dimensional gesture technology is used to enable the wearer of a head-mounted display device to interact naturally and directly with virtual objects in the VR / AR scene with both hands.
[0003] The interactive technology uses the 3D hand model skeleton as the gesture analysis object, and identifies and tracks multiple hand joints and degrees of freedom to distinguish various hand gestures of one or both hands, thereby achieving human-computer interaction. Therefore, the extraction of joints and the determination of degrees of freedom are important research directions in 3D gesture estimation.
[0004] Currently, the binocular camera on the head-mounted display device is used to obtain hand images from different perspectives in real time, and the two-dimensional (2D) hand joint points extracted from the hand images are converted into three-dimensional (3D) hand joint points through binocular camera projection.
[0005] Due to the lack of distortion correction for the hand images taken by the binocular camera, the predicted 3D hand joints have inaccurate coordinates due to image distortion, which affects the subsequent hand posture recognition results. Summary of the invention
[0006] The embodiments of the present application provide a hand gesture recognition method, a head-mounted display device, and a storage medium, which can obtain an undistorted hand image from the original binocular image and improve the accuracy of the hand joint point recognition results.
[0007] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions: In a first aspect, an embodiment of the present application provides a hand gesture recognition method, the method comprising: Tracking and acquiring a first hand region of a current frame image in a binocular video of a hand, projecting the first hand region through binocular parallax, and obtaining a first hand bounding box of the first hand region in a three-dimensional space; From the first hand bounding box, select the initial center position of the hand, and calculate the first transformation matrix from the initial center position to the direction of the image optical center; Perform multiple iterations of optimization on the first transformation matrix, and use the first transformation matrix after the iteration optimization as the first projection direction; According to the first projection direction, projection correction is performed on the image of the first hand area in the current frame image to generate a corrected first hand image, wherein the first hand image is used to input into a pre-trained hand joint point detection model for hand joint point recognition to obtain a recognition result.
[0008] In combination with the first aspect, in a possible design, the iteration process of the first transformation matrix is as follows: Based on the first transformation matrix, the positions of the hand joint points in the first hand bounding box are corrected, and a first direction vector of the corrected hand center point in the three-dimensional space is obtained; The product of the first transformation matrix and the first direction vector is used as the first transformation matrix for the next iterative optimization, and the above iterative process is repeated until the first transformation matrix converges and the iteration is stopped.
[0009] In combination with the first aspect, in a possible design manner, based on the first transformation matrix, correcting the position of the hand joint point in the first hand bounding box, and obtaining the first direction vector of the corrected hand center point in the three-dimensional space, includes: Based on the first transformation matrix, transform the position of the hand joint point in the first hand bounding box from the first three-dimensional position to the second three-dimensional position; Obtaining the position information of the projection center point of the second three-dimensional position on the equatorial plane to obtain the target center position; The target center position is back-projected through the stereographic plane to obtain the first direction vector of the corrected hand center point in three-dimensional space.
[0010] In combination with the first aspect, in a possible design manner, performing projection correction on the image of the first hand area in the current frame image according to the first projection direction to generate a corrected first hand image includes: According to the first projection direction, the hand joint points in the first hand bounding box are projected onto the equatorial plane to obtain normalized projection coordinates of the hand joint points; An affine transformation is performed on the image of the first hand region in the current frame image based on the normalized projection coordinates to generate a corrected first hand image.
[0011] In combination with the first aspect, in a possible design manner, after generating the corrected first hand image, the method further includes: Get the normalized coordinates of the hand joint points of the previous frame image; The normalized coordinates of the hand joint points in the previous frame image and the first hand image are input into the hand joint point detection model for inference to obtain a recognition result, wherein the recognition result includes a heat map and a relative distance heat map of the hand joint points in the current frame image.
[0012] In combination with the first aspect, in a possible design manner, after obtaining the recognition result, the method further includes: Construct an objective function, which is equal to the sum of the projection error and the distance error term and the smoothness constraint term, wherein the projection error and the distance error term are obtained by the sum of the projection error values of 26 hand joints and the distance error values of 26 hand joints; wherein the projection error value of each hand joint is equal to the square of the norm of the difference between the normalized pixel position in the heat map corresponding to the hand joint and the pixel position after the three-dimensional coordinates are projected by the camera, and the distance error value of each hand joint is equal to the square of the difference between the norm of the difference between the three-dimensional coordinates of the hand joint and the three-dimensional coordinates of the proximal metacarpal joint of the middle finger divided by the palm distance and the normalized relative distance of the hand joint in the relative distance heat map; The smoothness constraint term is obtained by the sum of the squares of the differences between the three-dimensional coordinates of each of the 26 hand joints in the t-th frame and the three-dimensional coordinates in the t-1-th frame.
[0013] The expression of the objective function is E(P)=E t (P)+E s (P), E t The expression of (P) is: , in the expression, E t (P) is the projection error and distance error term, u i is the normalized pixel position of the ith hand joint point in the heat map, p i is the three-dimensional coordinate of the i-th hand joint point, Π is the camera projection model, p midproxi is the three-dimensional coordinate of the proximal metacarpal joint of the middle finger, I h is the palm distance, I i is the normalized relative distance of the ith hand joint point in the relative distance heat map, and its value range is [-1,1]; E s The expression of (P) is: , in the expression, E s (P) is the smoothness constraint, is the 3D coordinate of the i-th hand joint point in the t-1-th frame, is the three-dimensional coordinate of the i-th hand joint point in the t-th frame; The result of iterative optimization of the objective function is used as the three-dimensional coordinates of the hand joint points in the current frame image obtained based on the optimization of the recognition results.
[0014] In combination with the first aspect, in a possible design, selecting the initial center position of the hand from the first hand bounding box and calculating the first transformation matrix from the initial center position to the optical center direction of the image includes: The position of the proximal metacarpal joint of the middle finger in the first hand bounding box is taken as the initial center position, and the first transformation matrix from the initial center position to the direction of the image optical center is calculated.
[0015] In combination with the first aspect, in a possible design manner, the step of tracking and acquiring a first hand region of a current frame image in a binocular video of a hand includes: Get the current frame image in the hand binocular video; If the previous frame image of the current frame image tracks a hand, a first hand region in the current frame image is obtained, where the first hand region is a hand region in the current frame image predicted based on the positions of the hand joints in the previous frame image; The method further comprises: If the previous frame image of the current frame image does not track the hand, the current frame image is input into the hand detection model for inference to obtain the hand center point position and the second hand area of the current frame image; Calculate the maximum angle between the hand center point and the four sides of the second hand region to obtain a second transformation matrix; Projecting the second hand region through binocular parallax to obtain a second hand bounding box of the second hand region in three-dimensional space; Based on the second transformation matrix, the position of the hand joint point in the second hand bounding box is corrected, and a second direction vector of the corrected hand center point in the three-dimensional space is obtained; Obtain a second projection direction according to the product of the second transformation matrix and the second direction vector; The projection radius is calculated based on the maximum angle, and the image of the second hand area in the current frame image is projected and corrected using the projection radius and the second projection direction to generate a corrected second hand image.
[0016] In combination with the first aspect, in a possible design, projection is performed on the hand area through binocular parallax, including: Calculate the overlap between the hand area corresponding to the left hand and the hand area corresponding to the right hand in the current frame image; When the overlap is less than a threshold, the hand area corresponding to the left hand and the hand area corresponding to the right hand are projected by binocular parallax; When the overlap is greater than or equal to the threshold and one of the left and right hands is tracked in the current frame image, the hand area corresponding to the other hand is projected by binocular parallax; The hand region is the first hand region or the second hand region.
[0017] In a second aspect, an embodiment of the present application provides a hand gesture recognition device, the device comprising: A projection module is used to track and obtain a first hand region of a current frame image in a binocular video of a hand, and to project the first hand region through binocular parallax to obtain a first hand bounding box of the first hand region in a three-dimensional space; A matrix calculation module, used to select the initial center position of the hand from the first hand bounding box, and calculate a first transformation matrix from the initial center position to the direction of the optical center of the image; A projection direction calculation module is used to perform multiple iterations of optimization on the first transformation matrix, and use the first transformation matrix after iteration optimization as the first projection direction; An image correction module is used to perform projection correction on the image of the first hand area in the current frame image according to a first projection direction to generate a corrected first hand image, wherein the first hand image is used to input into a pre-trained hand joint point detection model for hand joint point recognition to obtain a recognition result.
[0018] In a third aspect, an embodiment of the present application provides a head-mounted display device, comprising at least one processor, a memory, a camera module, and a display module, wherein the at least one processor is coupled to the memory and is used to read and execute instructions in the memory to execute the method of the first aspect and its possible design methods.
[0019] In a fourth aspect, an embodiment of the present application provides a storage medium, in which a computer program is stored, wherein the computer program is configured to execute the method of the first aspect and its possible design manner when running.
[0020] Compared with the prior art, the embodiments of the present application provide a hand gesture recognition method, a head-mounted display device, and a storage medium, wherein the method can be applied to a head-mounted display device, including: tracking and acquiring the first hand area of the current frame image in the binocular video of the hand, projecting the first hand area through binocular parallax, obtaining the first hand bounding box of the first hand area in three-dimensional space, and then selecting the initial center position of the hand from the first hand bounding box, and calculating the first transformation matrix from the initial center position to the optical center direction of the image. The first transformation matrix is iteratively optimized for multiple times, and the first transformation matrix after iterative optimization is used as the first projection direction. Then, according to the first projection direction, the image of the first hand area in the current frame image is projected and corrected to generate a corrected first hand image, wherein the first hand image is used to input into a pre-trained hand joint point detection model for hand joint point recognition to obtain a recognition result. The present application projects the hand joint points obtained by tracking into three-dimensional space and iteratively optimizes the first transformation matrix multiple times, thereby gradually adjusting the line of sight of the current frame image to be parallel to the direction of the optical center of the image. Compared with the original current frame image, the hand posture in the corrected first hand image is standardized. Therefore, the use of the first hand image obtained by the present application is conducive to the model to more accurately infer the position and distance of the hand joint points, thereby improving the accuracy of subsequent hand posture recognition.
[0021] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A diagram showing an application environment of a hand gesture recognition method provided by an embodiment of the present application is shown; Figure 2 A flow chart of a hand gesture recognition method provided by an embodiment of the present application is shown; Figure 3 A flowchart of a method for verifying the validity of a box selection result provided by an embodiment of the present application is shown; Figure 4 A flowchart of another hand gesture recognition method provided by an embodiment of the present application is shown; Figure 5 A flowchart of another method for verifying the validity of a box selection result provided by an embodiment of the present application is shown; Figure 6 A schematic diagram of the hardware structure of a hand gesture recognition device provided in an embodiment of the present application is shown; Figure 7 A hardware structure block diagram of a head-mounted display device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0023] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0024] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by people with general skills in the technical field to which this application belongs. The words "one", "a", "the", "these" and the like in this application do not indicate a quantitative limitation, and they may be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" may mean: A exists alone, A and B exist at the same time, and B exists alone. Generally, the character " / " indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", "third", etc. in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0025] 3D gesture estimation (or hand gesture recognition) is crucial in the AR / VR field because it provides a natural and intuitive way of interaction, greatly enhancing the user experience and immersion. Users can control virtual objects directly through gestures without relying on traditional controllers, thus achieving more flexible and precise operations, which is particularly suitable for tasks that require high-precision input. In addition, it supports complex human-computer interaction, adapts to various application scenarios such as education, medical care, and entertainment, simplifies hardware requirements, and makes interaction more free.
[0026] 3D gesture estimation also facilitates social interaction, allowing users to engage in rich non-verbal communication in multi-person environments, such as handshakes or hugs, while providing a covert control method to enhance safety and privacy protection for use in public places. As technology advances, gesture recognition becomes more accurate and reliable, driving the development of innovative applications such as gesture-driven games and teaching software, injecting new vitality into the AR / VR industry. In short, 3D gesture estimation is one of the key technologies to achieve immersive AR / VR experience.
[0027] In human-computer interaction scenarios, the three-dimensional coordinate information of the hand joints is needed to identify the user's hand gestures, including clicking, sliding, grabbing, stroking, flipping, clenching fists, clapping, etc., so as to generate various interactions with virtual objects. Therefore, obtaining the three-dimensional coordinates of the hand joints is the core of hand gesture recognition.
[0028] Currently, binocular images of the hand can be obtained through binocular cameras, the two-dimensional coordinates of the hand joints can be identified from the binocular images, and the three-dimensional coordinates of the hand can be obtained through camera projection. Since the three-dimensional coordinates obtained based on the binocular camera are achieved through three-dimensional reconstruction, when the image is distorted, the obtained three-dimensional coordinates are not accurate enough, which has serious defects in gesture interaction scenarios with high precision requirements.
[0029] In view of this, the embodiment of the present application provides a hand gesture recognition method, which performs projection correction on the image of the first hand area in the current frame image by iterating the projection direction multiple times, so that an image taken due to the tilt of the observation line of sight is corrected to an image with the line of sight perpendicular to the center. In this way, the hand image sent to the model reasoning overcomes the distortion effect, and the position and relative distance of the hand joint points obtained by the model reasoning will be more accurate.
[0030] The solution of the present application can be applied in human-computer interaction scenarios, such as human-computer interaction of terminals and human-computer interaction scenarios of vehicle-mounted systems. Figure 1 FIG. 1 shows an application environment diagram of a hand gesture recognition method provided by an embodiment of the present application, such as Figure 1 As shown, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal 102 can execute the method provided in the embodiment of the present application. Alternatively, the server 104 executes the method and returns the recognition result to the terminal 102. Alternatively, the server 104 trains the tracking model, the hand detection model, and the hand joint point detection model. After the training is completed, the server 104 sends the model to the terminal 102, and the terminal performs reasoning based on the deployed model.
[0031] The terminal may specifically include a virtual reality headset (VR headset), a head-mounted display device, an electronic display screen, and a mixed reality (MR) device, etc. MR devices may include MR glasses, MR helmets, MR cameras, etc., and head-mounted display devices include augmented reality glasses (AR glasses), MR glasses, etc. The vehicle-mounted system may be a vehicle-mounted chip, a vehicle-mounted device (such as a vehicle machine with the ability to build an augmented reality space, a vehicle-mounted computer, etc.), etc.
[0032] Figure 2 A flow chart of a hand gesture recognition method provided by an embodiment of the present application is shown. Figure 2 The method shown can be applied to a scenario where a user wears a head-mounted display device and performs gesture interaction with the head-mounted display device. The head-mounted display device may be AR glasses. The method includes steps S201 to S204.
[0033] Step S201, tracking and acquiring a first hand region of a current frame image in a binocular video of a hand, projecting the first hand region through binocular parallax, and obtaining a first hand bounding box of the first hand region in a three-dimensional space.
[0034] Among them, the hand binocular video refers to the hand movement video captured synchronously by the left and right cameras. Specifically, the head-mounted display device contains two cameras (or binocular cameras) that shoot the hands from different perspectives. The two cameras synchronously collect images in real time to simulate the binocular observation of objects by humans.
[0035] The current frame image includes a left eye image and a right eye image of the current frame, wherein the left eye image (or the right eye image) may include the left hand and the right hand, or only include one hand, or include no hand.
[0036] In this step, a tracking algorithm is used to obtain the first hand area. Specifically, it is determined whether the hand joint points are recognized in the previous frame image of the current frame image. If the hand joint points are recognized in the previous frame image, it means that the hand is tracked. Then, based on the position of the hand joint points in the previous frame image, the position of the hand joint points in the current frame image is predicted.
[0037] It can be understood that since there is no restrictive requirement on the number of hands in the binocular image, this step is divided into four cases: judging whether the left hand joint points are recognized in the left eye image of the previous frame, judging whether the right hand joint points are recognized in the left eye image of the previous frame, judging whether the left hand joint points are recognized in the right eye image of the previous frame, and judging whether the right hand joint points are recognized in the right eye image of the previous frame.
[0038] In the four cases, if the hand joints of any hand are identified in any eye image of the previous frame, the possible position of the hand in the eye image in the current frame image is determined by tracking calculation. For example, if the hand joints of the left hand and the right hand are identified in the left eye image of the previous frame, the possible positions of the left hand and the right hand in the left eye image of the current frame are determined by tracking calculation.
[0039] The specific calculation method is as follows: Record the position of the hand joints at time t-1 as joints t-1 , the timestamp at time t-1 is ts t-1 , the position of the hand joint at time t is joint st , the timestamp at time t is t st , the predicted timestamp is t s+1 , α is the coefficient for adjusting the tracking prediction effect, then the predicted position of the hand joint point in the next frame image is: , Through the left and right camera back-projection model, the predicted positions of the hand joints are projected onto the left and right camera images to obtain: .
[0040] in, is the back-projection model of the left camera, is the 3D coordinate of the left hand joint point, is the back-projection model of the right camera, are the 3D coordinates of the right hand joint points.
[0041] Let's take the left hand as an example (the same applies to the right hand): calculate the pixel average of all 3D hand joints: .
[0042] The coordinate point corresponding to the pixel average value is used as the center point of the hand tracking area. At the same time, the coordinates of each 3D hand joint point are calculated. The maximum value of: .
[0043] The side length of the hand area selection is:
[0044] The first hand region can be obtained through the side length of the hand selection region and the center point of the hand region.
[0045] This step uses the position information of the hand joints in the previous frame image to predict the position information of the hand joints in the current frame image through mathematical modeling, avoiding complex model reasoning for each frame image. This can reduce the amount of calculation, has low hardware resource requirements, and is more suitable for AR / VR devices.
[0046] After the first hand area is obtained, it is stereo projected. Specifically, the binocular camera captures two images of the hand from different perspectives, and the difference in pixel position of the same hand area in the two images is called binocular parallax. Based on the binocular parallax, the distance between the left and right cameras, and the focal length, the Z-axis coordinate of the hand is calculated, and then based on the Z-axis coordinate and the pixel coordinate of the hand on the image, the coordinates of the first hand area in the three-dimensional space are calculated through the camera intrinsic parameter matrix. In this step, only the image of the local area (i.e., the first hand area) needs to be projected, thereby reducing the amount of calculation for three-dimensional reconstruction.
[0047] In some of the embodiments, the method further includes cross-validating the box selection results to identify the validity of the box selection results, and not performing projection on hands with invalid box selection to save computing resources.
[0048] Specifically, the method includes: Figure 3 Steps S301 to S303 are shown.
[0049] Step S301: Calculate the degree of overlap between a first hand region corresponding to the left hand and a first hand region corresponding to the right hand in the current frame image.
[0050] The current frame image includes a left eye image and a right eye image. In this step, the overlap of the first hand area corresponding to the left hand and the first hand area corresponding to the right hand in the left eye image and the overlap in the right eye image are calculated. The overlap (IoU) can be the ratio of the intersection area and the union area of the first hand areas corresponding to the left and right hands, where the intersection area refers to the pixel area of the overlapping part of the two first hand areas, and the union area refers to the difference between the total coverage area of the two first hand areas and the intersection area.
[0051] Step S302: When the overlap degree is less than a threshold, projecting the first hand region corresponding to the left hand and the first hand region corresponding to the right hand through binocular parallax.
[0052] The threshold is a pre-set critical value, and the threshold is a number in the range of [0, 1]. For example, the threshold can be 0.3, 0.4, 0.8, etc.
[0053] The threshold is used to determine whether the areas of the two hands overlap significantly. In this step, if the IoU is less than the threshold, the two hands are considered independent (or non-overlapping). Then, the areas of the two hands in each eye image can be stereo projected separately, and then the following step S202 is executed.
[0054] In some embodiments, in order to adapt to complex hand interaction scenarios, a dynamic threshold can be used to adapt to the dynamic changes of frequent switching of both hands. For example, when the hands move frequently, the threshold is lowered to increase sensitivity; when the movements are stable, the threshold is increased to reduce misjudgment.
[0055] Step S303: when the overlap degree is greater than or equal to the threshold and one of the left hand and the right hand is tracked in the current frame image, a first hand area corresponding to the other hand is projected through binocular parallax.
[0056] In this step, if IoU is greater than or equal to the threshold, it is considered that the two hands may overlap or cross. In this case, when one of the hands has been stably located in the current frame image (either eye image) through the tracking algorithm, only the first hand area of the other hand is projected with binocular disparity to avoid repeated calculation of the stably tracked hand, which can reduce resource consumption.
[0057] In addition, when the overlap is greater than or equal to the threshold and both hands are not tracked in the current frame image, the calculation of the frame image is abandoned.
[0058] Through steps S301 to S303, a first hand region with valid frame selection is screened out, and after stereoscopic projection is performed on the first hand bounding box, the following step S202 is executed on the first hand bounding box.
[0059] It should be noted that here, assuming that the left hand is valid, the corresponding hand may be seen in the left or right eye image, so each hand will have the following processing flow 1 to 2 times, depending on the valid result of the hand area detection. The following describes the process of one image correction process, and the others are similar, and there may be up to four times.
[0060] Step S202: Select the initial center position of the hand from the first hand bounding box, and calculate the first transformation matrix from the initial center position to the optical center direction of the image.
[0061] In this step, the initial center position refers to the position coordinates of the center point of the hand before iterative optimization. In some embodiments, the position of the proximal metacarpal joint of the middle finger is used as the initial center position.
[0062] Among them, the direction of the optical center of the image refers to (0, 0, -1). This step calculates the transformation relationship from the local point to (0, 0, -1). Subsequently, based on this transformation relationship, all points on the hand can be mapped to (0, 0, -1). For details, refer to step S203.
[0063] Step S203: perform multiple iterations of optimization on the first transformation matrix, and use the first transformation matrix after the iteration optimization as the first projection direction.
[0064] The iterative process of the first transformation matrix is as follows: based on the first transformation matrix, the position of the hand joint points in the first hand bounding box is corrected, and the first direction vector of the hand center point in the three-dimensional space after correction is obtained; the product of the first transformation matrix and the first direction vector is used as the first transformation matrix for the next iterative optimization, and the above iterative process is repeated until the first transformation matrix converges and the iteration is stopped.
[0065] Correction refers to mapping the position of the hand joints from the original position to the transformed position with the image light direction as the observation line of sight through the first transformation matrix. That is, each hand joint is rotated to the [0, 0, -1] direction through the inverse transformation of the first transformation matrix (T_mid_o) to obtain a new position in three-dimensional space. The purpose of correcting the position of the 3D hand joints is to standardize the original hand postures of various shapes so that the palm of the hand is vertically facing the observation line of sight, which is conducive to subsequent model reasoning.
[0066] Since the image will be distorted during projection, there is an error between the initial center position on which this correction is based and the actual center position of the hand. Therefore, this step projects the corrected hand coordinate points to the equatorial plane. Among them, the hand is represented by 26 degrees of freedom to represent the three-dimensional coordinates of 21 hand joints. In the embodiment of the present application, for the sake of ease of expression, it is referred to as 26 joints. In the equatorial plane coordinates of the 26 joints, the maximum coordinate point and the minimum coordinate point are averaged to obtain the target center position of the 26 joints on the equatorial plane. The target center position is back-projected through the equatorial plane to obtain the first direction vector V of the center point of the hand in three-dimensional space after correction. dir After correction, the center point of the hand approaches the actual center position step by step through multiple iterations. Then the direction of the new center point in the original middle finger proximal metacarpal space is dir_mid_new = T_mid_o × V dir .
[0067] Then, dir_mid_new is used to replace the first transformation matrix, and step S203 is repeated until the direction converges, and the final dir_mid_new is used as the first projection direction. The purpose of multiple iterations is to gradually adjust the projection direction so that the projection direction is parallel to the direction of the optical center of the image, that is, the line of sight is level with the center of the image, which can reduce the image distortion caused by the tilt of the viewing angle.
[0068] In step S203, each iteration includes position correction, projection on the polar plane and matrix update, and the first transformation matrix (T_mid_o) is adjusted in multiple iterations to gradually correct the position deviation of the hand joints in the three-dimensional space. Through gradient descent optimization, the real posture of the hand in the three-dimensional space is gradually approached to reduce projection distortion.
[0069] Step S204: According to the first projection direction, projection correction is performed on the image of the first hand area in the current frame image to generate a corrected first hand image, wherein the first hand image is used to input into a pre-trained hand joint point detection model for hand joint point recognition to obtain a recognition result.
[0070] This step performs geometric correction on the hand image in the current frame based on the final optimized projection direction to generate a distortion-free two-dimensional image. The geometric correction can be an affine transformation. The palm area in the geometrically corrected image is stretched or rotated to a normal viewing angle, thereby eliminating image distortion caused by changes in viewing angles, and improving the recognition accuracy of the hand joint point detection model, that is, making the inferred coordinates and distances of the hand joint points more accurate.
[0071] In some of the embodiments, step S204 includes: projecting the hand joint points in the first hand bounding box onto the equatorial plane according to the first projection direction to obtain normalized projection coordinates of the hand joint points; performing an affine transformation on the image of the first hand area in the current frame image based on the normalized projection coordinates to generate a corrected first hand image.
[0072] In this embodiment, in addition to unifying the observation angle of the hand, the coordinates of the hand joints are normalized to unify the size of the hand. Specifically, the stereographic plane projection coordinates of the joints within the projection radius are mapped from [-stereographic_radius, stereographic_radius] to [-1, 1]. If it is the right hand, the X coordinate should be reversed, that is, the left and right hands are mirrored. Among them, stereographic_radius refers to the projection radius.
[0073] Step S204 uses binocular stereo vision correction to select an image of the hand area position (i.e., the first hand area) from the original image (i.e., the current frame image) for correction and cropping, thereby obtaining a corrected and cropped image (i.e., the first hand image) that meets the size required by the input model.
[0074] In the embodiment recorded in the above steps S201 to S204, the hand joint points obtained by tracking are projected into three-dimensional space, and the first transformation matrix is optimized through multiple iterations, and the line of sight of the current frame image is gradually adjusted to be parallel to the direction of the optical center of the image. In this way, the first hand image after correction is compared with the original current frame image. The hand posture is standardized, so the first hand image obtained by using this embodiment is conducive to the model to make more accurate inferences about the position and distance of the hand joint points, so that the subsequent recognition of the hand posture is more accurate. It should be noted that this embodiment is not just a simple 2D rotation of the plane image. The 2D rotation correction method is only applicable to simple gestures of unfolded palms and cannot correct 3D postures. This embodiment eliminates projection errors through iterative optimization, and standardizes the hand posture to the [0, 0, -1] direction through 3D rotation transformation. In this way, affine transformation of complex gestures such as finger bending, occlusion, and rotation can be achieved, and the robustness and accuracy of the model's posture recognition of complex gestures are improved, which is more suitable for complex gesture interaction scenarios between users and AR / VR devices.
[0075] In some of the embodiments, after the first hand image is obtained, it can be used for model reasoning.
[0076] Specifically, after step S204, the method provided in this embodiment further includes: The normalized coordinates of the hand joints of the previous frame image are obtained, and the normalized coordinates of the hand joints of the previous frame image and the first hand image are input into the hand joint detection model for inference to obtain the recognition result, wherein the recognition result includes the heat map and relative distance heat map of the hand joints in the current frame image.
[0077] It is understandable that the normalized coordinates of the hand joint points of the previous frame of the input image provide a reference for the results output by the hand joint point detection model. The first hand image is a standardized image with a fixed size after correction and cropping. Both can improve the accuracy of the position and relative distance of the hand joint points in the recognition results output by the model.
[0078] In this embodiment, the hand joint point detection model is pre-trained, and its training process can be performed in the server, that is, the server sends the trained hand joint point detection model to the head-mounted display device, and the head-mounted display device directly runs the hand joint point detection model. Alternatively, the hand joint point detection model is trained and tested in the head-mounted display device, and this application does not limit this.
[0079] The above is explained by taking tracking of hand joint points as an example. In other embodiments, the previous frame of the current frame image does not track the hand. For the correction method of the current frame image, refer to the following description.
[0080] Figure 4A flowchart of another hand gesture recognition method provided in an embodiment of the present application is shown. Figure 4 The method shown can be applied to a scenario where a user wears a head-mounted display device and performs gesture interaction with the head-mounted display device. The head-mounted display device may be AR glasses. The method includes steps S401 to S407.
[0081] Step S401: Obtain the current frame image in the hand binocular video.
[0082] Step S402: If the previous frame image of the current frame image does not track the hand, the current frame image is input into the hand detection model for inference to obtain the hand center point position of the current frame image and the second hand area.
[0083] This step is used to locate the hand through the model to prevent process interruption when hand tracking fails (such as when the hand moves quickly or is blocked) or when detecting the first image frame.
[0084] In this step, it is determined whether the hand joints are recognized in the previous frame image of the current frame image. If the hand joints are not recognized in the previous frame image, it means that the hand is not tracked. Then, the pre-trained hand detection model is used to infer the current frame image, identify the area where the hand joints exist, and output the hand center point position and the second hand area.
[0085] The hand detection model refers to a target detection model based on a convolutional neural network. Its input is a binocular image with the required resolution. Its output is information about whether there are left and right hands in each image, the pixel position of the center points of the left and right hands (i.e., the position of the hand center points), and the pixel size of the left and right hands (i.e., the second hand area). The second hand area refers to a rectangular area based on the position of the hand center point, which consists of four sides. The hand detection model can be a network such as YOLO and Faster R-CNN. As an example, the backbone network of the hand detection model uses MobileNet.
[0086] It is understandable that, since there is no restrictive requirement for the number of hands in the binocular image, this step is divided into four cases: judging whether the left hand joint point is recognized in the left eye image of the previous frame, judging whether the right hand joint point is recognized in the left eye image of the previous frame, judging whether the left hand joint point is recognized in the right eye image of the previous frame, and judging whether the right hand joint point is recognized in the right eye image of the previous frame. For details, please refer to the explanation of this point in step S201, and this application will not make a cumbersome explanation here.
[0087] After step S402, the method provided in this embodiment further includes cross-validating the box selection results to identify the validity of the results. For hands with invalid box selection, the following steps S403 to S407 are not executed to save computing resources.
[0088] Specifically, the method includes: Figure 5 Steps S501 to S503 are shown.
[0089] Step S501: Calculate the degree of overlap between the second hand region corresponding to the left hand and the second hand region corresponding to the right hand in the current frame image.
[0090] Step S502: When the overlap degree is less than a threshold, projecting the second hand region corresponding to the left hand and the second hand region corresponding to the right hand through binocular parallax.
[0091] Step S503: when the overlap degree is greater than or equal to the threshold and one of the left hand and the right hand is tracked in the current frame image, a second hand area corresponding to the other hand is projected through binocular parallax.
[0092] For the explanation of steps S501 to S503 , please refer to the above steps S301 to S303 , and this application will not provide a redundant description here.
[0093] Through steps S501 to S503, the second hand area with valid frame selection is screened out, and the following step S403 is executed for the second hand area with valid frame selection.
[0094] Step S403: Calculate the maximum angle between the hand center point position and the four sides of the second hand region to obtain a second transformation matrix.
[0095] The four sides of the second hand area form angles with the line connecting the center points respectively, and the largest angle is selected as the extension direction of the hand in the plane. This can make the hand extend at the maximum angle in space, that is, the hand posture can be observed to the greatest extent, so the maximum angle is used as the second transformation matrix for transformation.
[0096] The second transformation matrix is an affine transformation matrix, which represents the linear transformation relationship of the second hand region rotating from the original direction to the direction of the maximum angle.
[0097] Step S404: project the second hand region through binocular parallax to obtain a second hand bounding box of the second hand region in three-dimensional space.
[0098] Using the parallax information of the binocular camera, the two-dimensional area of the hand is mapped from the image plane to the three-dimensional space to determine the position and size of the hand in the three-dimensional space.
[0099] This step may refer to the above step S201 and will not be described again here.
[0100] Step S405: Based on the second transformation matrix, correct the positions of the hand joint points in the second hand bounding box, and obtain the second direction vector of the corrected hand center point in the three-dimensional space.
[0101] This step rotates and translates the three-dimensional coordinates of each hand joint point through the second transformation matrix to correct its 3D position. For details, please refer to the introduction of step S203.
[0102] Step S406: Obtain a second projection direction according to the product of the second transformation matrix and the second direction vector.
[0103] The second transformation matrix can be represented by T_mid_o, and the second direction vector can be represented by V dir Indicates that the second projection direction is expressed as: dir_mid_new=T_mid_o×V dir .
[0104] Step S407: Calculate the projection radius based on the maximum angle, and perform projection correction on the image of the second hand area in the current frame image by using the projection radius and the second projection direction to generate a corrected second hand image.
[0105] Specifically, the projection radius is calculated using the following formula: sin(angle) / (1-cos(angle)), where angle represents the maximum angle.
[0106] The projection direction and projection radius are used to correct the hand's observation line of sight in the image, so that an image viewed at an angle can be corrected to an image viewed vertically. This can eliminate the difference in hand postures and reduce projection errors, thereby reducing the image distortion caused by changes in viewing angle and the impact on the model inference results.
[0107] Through step S401 to step S407, the embodiment of the present application provides another hand gesture recognition method, which obtains the hand key points of the current frame image through model reasoning when tracking fails, so as to ensure the consistency of hand gesture recognition. At the same time, the image of the second hand area obtained by model reasoning is also corrected to reduce image distortion, so that the subsequent reasoning result of the second hand image will be more accurate.
[0108] In some of the embodiments, after the second hand image is obtained, it can be used for model reasoning.
[0109] Specifically, after step S407, the method provided in this embodiment further includes: The second hand image is input into the hand joint point detection model for inference to obtain a recognition result, wherein the recognition result includes a heat map and a relative distance heat map of the hand joint points in the current frame image.
[0110] It can be understood that the second hand image is a standardized image with a fixed size after correction and cropping, so the position and relative distance of the hand joint points in the recognition result output by the model will be more accurate.
[0111] In this embodiment, the hand joint point detection model is pre-trained, and its training process can be performed in the server, that is, the server sends the trained hand joint point detection model to the head-mounted display device, and the head-mounted display device directly runs the hand joint point detection model. Alternatively, the hand joint point detection model is trained and tested in the head-mounted display device, and this application does not limit this.
[0112] The above describes how to correct and crop hand images, including correcting and cropping images obtained using a hand tracking method and correcting and cropping images obtained using a hand detection model. In addition, it also describes how to send the corrected hand image to a hand joint point detection model for reasoning. The following describes a method for calculating the 3D coordinates of hand joint points provided in an embodiment of the present application, which can perform nonlinear optimization on the results obtained by model reasoning, so that the hand joint points have less jitter and the joint point positions are more accurate.
[0113] Specifically, the method includes: inputting the normalized coordinates of the hand joint points of the previous frame image and the hand image into the hand joint point detection model for inference to obtain a recognition result, wherein the recognition result includes a heat map and a relative distance heat map of the hand joint points in the current frame image. The hand image is the first hand image or the second hand image.
[0114] It can be understood that, if the hand is tracked in the previous frame image, the input items in the hand joint point detection model include the normalized coordinates of the hand joint points of the previous frame image. If the hand is not tracked in the previous frame image, there is no need to input the information of the previous frame image.
[0115] The hand image is the first hand image or the second hand image. The hand image refers to the hand region image of the current frame after preprocessing (such as affine transformation, cropping, grayscale normalization, etc.), and the size is fixed as the model input requirement. The purpose of normalizing the hand joints is to eliminate scale differences.
[0116] The hand key point detection model includes a heat map branch and a relative distance branch, where the heat map branch outputs the coordinates of each hand joint point. Therefore, the heat map of the hand joint points refers to the possible position of each hand joint point on the binocular image. The relative distance branch outputs the distance offset of each hand joint point relative to the hand center (i.e., the proximal metacarpal bone of the middle finger). Therefore, the relative distance heat map refers to the relative distance of each hand joint point to the hand center in the Z-axis direction. Among them, the relative distance can be normalized to a value in the range of [-1, 1].
[0117] After that, the method also includes: constructing an objective function, which is equal to the sum of the projection error and the distance error term, and the smoothness constraint term, wherein the projection error and the distance error term are obtained by the sum of the projection error values of 26 hand joints and the distance error values of 26 hand joints; wherein the projection error value of each hand joint is equal to the square of the norm of the difference between the normalized pixel position in the heat map corresponding to the hand joint and the pixel position after the three-dimensional coordinate is projected by the camera, and the distance error value of each hand joint is equal to the square of the difference between the norm of the difference between the three-dimensional coordinate of the hand joint and the three-dimensional coordinate of the proximal metacarpal joint of the middle finger divided by the palm distance and the normalized relative distance of the hand joint in the relative distance heat map; The smoothness constraint term is obtained by the sum of the squares of the differences between the three-dimensional coordinates of each of the 26 hand joints in the t-th frame and the three-dimensional coordinates in the t-1-th frame.
[0118] The expression of the objective function is E(P)=E t (P)+E s (P).
[0119] E t (P) is the projection error and distance error term, E t The expression of (P) is: , in the expression, assuming u i is [sg_x, sg_y]u i is the normalized pixel position of the ith hand joint point in the heat map of the hand joint points, p i is the three-dimensional coordinate of the i-th hand joint point, Π is the camera projection model, p midproxi is the three-dimensional coordinate of the proximal metacarpal joint of the middle finger, I h is the palm distance, I i is the normalized relative distance of the i-th hand joint point in the relative distance heat map.
[0120] The error projection Ensure that the projection of the 3D joint points on the image plane is consistent with the actual observations.
[0121] Distance Error This is to ensure that the relative distance of the heat map of the 3D joint points is consistent with the actual relative distance information, that is, the relative distance output by the model is as close as possible to the relative distance calculated by the joint points. is the relative distance, which is expressed by the ratio of the difference between the 3D coordinates of the i-th hand joint point and the 3D coordinates of the proximal metacarpal bone of the middle finger to the palm distance. This value reflects the relative distance between the i-th hand joint point and the hand center. This formula takes into account the different sizes of palms of different users, so the relative distance will be affected by the size of the palm, so the palm distance is introduced, so that the hand joints of each user have a better optimization effect in this way.
[0122] E s (P) is the smoothness constraint term, E s The expression of (P) is: , in the expression, is the 3D coordinate of the i-th hand joint point in the t-1-th frame, is the three-dimensional coordinate of the i-th hand joint point in the t-th frame.
[0123] Among them This is to ensure that the changes of 3D joint points between adjacent frames are smooth.
[0124] With the goal of minimizing the objective function, iterative optimization is performed, and the result of iterative optimization of the objective function is used as the 3D coordinates of the hand joint points in the current frame image obtained based on the recognition result optimization. The optimized 3D coordinates of each hand joint point are more accurate, so the finger shake is smaller and the recognition effect of the hand posture is better.
[0125] The method provided in the embodiment of the present application is described below with a specific example.
[0126] Optional step 1: Train a usable hand detection model or use an already trained model. The input of the model is the image of the required resolution, and the output of the model is the information of whether there are left and right hands in each image, the pixel position of the center point of the left and right hands, and the pixel size of the left and right hands.
[0127] Optional step 2: Train a hand joint detection model, or use a pre-trained model. The model's input parameters require cropped and rectified hand images, whether to use the joint positions of the previous frame, and the normalized joint positions of the previous frame (if the joint positions of the previous frame are used). The model outputs the heat map of the joints in the current frame, the heat map of the relative distance, the inference valid flags, and the curvature of each finger.
[0128] After obtaining the hand detection model and the hand joint point detection model, execute step 3.
[0129] Step 3: Obtain the possible position information of the hand position in the current frame image in the left and right eye images and whether there is hand information.
[0130] Specifically, according to whether the hand is tracked in the previous frame image, step 3 includes step 3.1 and step 3.2.
[0131] Step 3.1: In the previous frame, the hand joints are not correctly solved at the same time, including the case where both hands are missing or only one hand is solved. In these cases, the left and right eye images are sent to the hand detection model for inference after preprocessing. The output results are the information of whether there are left and right hands in each image, the pixel position of the center point of the left and right hands, and the pixel size of the left and right hand area.
[0132] Step 3.2: If the hand joints of a certain hand are correctly solved in the previous frame, then the position where the hand may appear on the image in the current frame is determined by tracking calculation. For details, please refer to step S201 above. In step 3.2, the frame selection area of the hand is obtained.
[0133] Step 3.3: Cross-validate the selection results in 3.1 and 3.2 to identify whether the hand selected in step 3.1 or step 3.2 is valid. If a hand is invalid, the subsequent workflow will not be performed for the corresponding hand. If both hands are invalid, wait for the next frame and restart step 3. Step 3.3 corresponds to steps S301 to S303 above, and to steps S501 to S503 above.
[0134] Step 4: Detect hand joints for the selected valid hand and its view.
[0135] Specifically, according to whether the hand is tracked in a frame of image, step 4 includes step 4.1 and step 4.2.
[0136] Step 4.1: If the current hand region (equivalent to the first hand region above) is obtained by the method of step 3.1, the pixel positions of the four vertices on the image can be obtained by reverse calculation through the central region of the hand and the size of the hand region. This step is equivalent to steps S403 to S406 above.
[0137] Step 4.2: If the current hand region (equivalent to the second hand region above) is obtained by the method in step 3.2, then the transformation angle is calculated by the following method.
[0138] Step 4.2.1: Rotate the hand region to [0, 0, -1] (the direction of the image optical center) for projection correction. The specific method is to select the proximal metacarpal joint of the middle finger and calculate the coordinate transformation relationship T_mid_o from the proximal metacarpal joint of the middle finger to [0, 0, -1].
[0139] Step 4.2.2: Rotate each joint point to the direction [0, 0, -1] through the inverse transformation of T_mid_o, and calculate the coordinates [sg_x, sg_y] of each joint point on the equatorial plane through projection on the equatorial plane.
[0140] Step 4.2.3: In the stellar plane coordinates of the 26 joint points, the center position is obtained by averaging the maximum coordinate point and the minimum coordinate point.
[0141] Step 4.2.4: Back-project the recalculated center position through the equatorial plane to obtain the direction quantity V of the center point dir Then the direction of the new center point in the original middle finger proximal metacarpal space is dir_mid_new = T_mid_o × V dir .
[0142] Step 4.2.5: Use dir_mid_new to replace T_mid_o in 4.2.2, and repeat the calculation process from step 4.2.2 to step 4.2.4 until the direction converges.
[0143] Step 4.3: Use the converged new direction to calculate the stereoscopic projection of the 3D joint point on the stereoscopic plane (taking the left eye image as an example, the 3D joint point is the point in the OpenXR coordinate system of the left eye center, expressed in real physical scale).
[0144] Step 4.4: Map the stereographic plane projection coordinates of the joint points within the projection radius from [-stereographic_radius, stereographic_radius] to [-1, 1]. If it is right-handed, X needs to be reversed, that is, this step needs to be mirrored for the left and right hands.
[0145] Step 4.5: Select the image of the hand area from the original image through binocular stereo vision correction, correct it and crop it, and obtain the corrected and cropped image of the size required for input into the inference model.
[0146] Step 4.6: Send to the hand joint detection model for inference. The input is the normalized coordinates of the hand joints in the previous frame and the corrected cropped image of the current frame. The obtained results are the heat map of the normalized coordinates of the joints, the heat map of the relative distance, the valid inference flags, and the curvature of each finger.
[0147] Step 5: Use nonlinear optimization to calculate the 3D coordinates of the finger joints.
[0148] Specifically, in step 4.7, the image coordinates of the hand joints of the left or right hand under each camera (thermal map of the normalized coordinates of the joints) and the relative distance information of the joints (thermal map of the relative distance) are obtained. Based on the information output by the model, step 5 calculates the 3D coordinates of each hand joint through nonlinear optimization.
[0149] The optimization objective function can be written as E(P)=E t (P)+E s (P). The specific optimization process is not described here. The result of iterative optimization is the precise 3D coordinates of each finger joint of the corresponding hand.
[0150] A hand gesture recognition method provided in an embodiment of the present application is described in detail above. A hand gesture recognition device is introduced below.
[0151] Please refer to Figure 6 , Figure 6 FIG. 1 shows a hardware structure diagram of a hand gesture recognition device provided in an embodiment of the present application. Figure 6 As shown, the device comprises: The projection module 601 is used to track and obtain the first hand area of the current frame image in the hand binocular video, and project the first hand area through binocular parallax to obtain a first hand bounding box of the first hand area in three-dimensional space.
[0152] The matrix calculation module 602 is used to select the initial center position of the hand from the first hand bounding box and calculate the first transformation matrix from the initial center position to the optical center direction of the image.
[0153] The projection direction calculation module 603 is used to perform multiple iterations of optimization on the first transformation matrix, and use the first transformation matrix after iteration optimization as the first projection direction.
[0154] The image correction module 604 is used to perform projection correction on the image of the first hand area in the current frame image according to the first projection direction to generate a corrected first hand image, wherein the first hand image is used to input into a pre-trained hand joint point detection model to perform hand joint point recognition to obtain a recognition result.
[0155] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0156] The present application also provides a head mounted display device, such as Figure 7As shown, the head-mounted display device includes a processor 701, a memory 702, a shooting module 703, a display module 704, an interaction module 705, an IMU module 706, a gyroscope 707, a geomagnetic sensor 708 and a wireless communication module 709.
[0157] The memory 702 can be used to store computer programs, such as software programs and modules of application software. The processor 701 executes various functional applications and data processing by running the computer programs stored in the memory 702, such as implementing the method provided in the embodiments of the present application.
[0158] The shooting module 703 is mainly responsible for shooting the image information of the real environment, such as obtaining the binocular video of the user's hand, etc. The shooting module can be a camera.
[0159] The display module 704 is used to display virtual images and augmented reality content, and includes a micro-projection system and optical elements.
[0160] The interactive module 705 includes a microphone, an eye tracking sensor, a finger ring, a wristband, etc.
[0161] The wireless communication module 709 includes: 4G / 5G full network access, WiFi, Bluetooth, etc.
[0162] It can be understood by those skilled in the art that Figure 7 The structure shown is for illustration only and does not limit the structure of the MR glasses. Figure 7 More or fewer components as shown, or with Figure 7 Different configurations shown.
[0163] The following is a brief introduction to the application scenarios of head-mounted display devices.
[0164] In the gesture interaction scenario using a head-mounted display device, users can directly control virtual objects through gestures, such as picking up a water cup in the virtual space, waving to a digital person in the virtual space, etc., thereby enhancing the user's sense of immersion.
[0165] In addition, in combination with the method provided in the above embodiment, a storage medium may be provided in this embodiment to implement the method. The storage medium stores a computer program; when the computer program is executed by a processor, any one of the hand gesture recognition methods in the above embodiment is implemented.
[0166] It should be understood that the specific embodiments described herein are only used to explain the application, rather than to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of this application.
[0167] Obviously, the drawings are only some examples or embodiments of the present application. For ordinary technicians in the field, the present application can also be applied to other similar situations based on these drawings without creative work. In addition, it is understandable that although the work done in this development process may be complicated and lengthy, for ordinary technicians in the field, certain changes in design, manufacturing or production based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient content disclosed in this application.
[0168] The term "embodiment" in this application refers to a specific feature, structure or characteristic described in conjunction with the embodiment that can be included in at least one embodiment of the present application. The appearance of this phrase in various locations in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is clearly or implicitly understood by those of ordinary skill in the art that the embodiments described in this application can be combined with other embodiments without conflict.
[0169] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of patent protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the attached claims.
Claims
1. A hand gesture recognition method, characterized in that: Methods include: Tracking and acquiring a first hand region of a current frame image in a binocular video of a hand, projecting the first hand region through binocular parallax, and obtaining a first hand bounding box of the first hand region in a three-dimensional space; From the first hand bounding box, select the initial center position of the hand, and calculate the first transformation matrix from the initial center position to the direction of the image optical center; Perform multiple iterations of optimization on the first transformation matrix, and use the first transformation matrix after the iteration optimization as the first projection direction; According to the first projection direction, projection correction is performed on the image of the first hand area in the current frame image to generate a corrected first hand image, wherein the first hand image is used to input into a pre-trained hand joint point detection model for hand joint point recognition to obtain a recognition result.
2. The hand gesture recognition method according to claim 1, characterized in that: The iterative process of the first transformation matrix is as follows: Based on the first transformation matrix, the positions of the hand joint points in the first hand bounding box are corrected, and a first direction vector of the corrected hand center point in the three-dimensional space is obtained; The product of the first transformation matrix and the first direction vector is used as the first transformation matrix for the next iterative optimization, and the above iterative process is repeated until the first transformation matrix converges and the iteration is stopped.
3. The hand gesture recognition method according to claim 2, characterized in that: The method of correcting the position of the hand joint point in the first hand bounding box based on the first transformation matrix and obtaining the first direction vector of the hand center point in the three-dimensional space after correction includes: Based on the first transformation matrix, transform the position of the hand joint point in the first hand bounding box from the first three-dimensional position to the second three-dimensional position; Obtaining the position information of the projection center point of the second three-dimensional position on the equatorial plane to obtain the target center position; The target center position is back-projected through the stereographic plane to obtain the first direction vector of the corrected hand center point in three-dimensional space.
4. The hand gesture recognition method according to claim 1, characterized in that: The step of performing projection correction on the image of the first hand region in the current frame image according to the first projection direction to generate a corrected first hand image includes: According to the first projection direction, the hand joint points in the first hand bounding box are projected onto the equatorial plane to obtain normalized projection coordinates of the hand joint points; An affine transformation is performed on the image of the first hand region in the current frame image based on the normalized projection coordinates to generate a corrected first hand image.
5. The hand gesture recognition method according to claim 1, characterized in that: After generating the corrected first hand image, the method further includes: Get the normalized coordinates of the hand joint points of the previous frame image; The normalized coordinates of the hand joint points in the previous frame image and the first hand image are input into the hand joint point detection model for inference to obtain a recognition result, wherein the recognition result includes a heat map and a relative distance heat map of the hand joint points in the current frame image.
6. The hand gesture recognition method according to claim 5, characterized in that: After obtaining the recognition result, the method further includes: Construct an objective function, which is equal to the sum of the projection error and the distance error term and the smoothness constraint term, wherein the projection error and the distance error term are obtained by the sum of the projection error values of the 26 hand joints and the distance error values of the 26 hand joints; wherein the projection error value of each hand joint is equal to the square of the norm of the difference between the normalized pixel position in the heat map corresponding to the hand joint and the pixel position after the three-dimensional coordinates are projected by the camera, and the distance error value of each hand joint is equal to the square of the difference between the norm of the difference between the three-dimensional coordinates of the hand joint and the three-dimensional coordinates of the proximal metacarpal joint of the middle finger divided by the palm distance and the normalized relative distance of the hand joint in the relative distance heat map; The smoothness constraint term is obtained by summing the square of the difference between the three-dimensional coordinates of each of the 26 hand joints in the current frame image and the three-dimensional coordinates in the previous frame image of the current frame image.
7. The hand gesture recognition method according to claim 1, characterized in that: The step of tracking and acquiring a first hand region of a current frame image in a binocular video of a hand includes: Get the current frame image in the hand binocular video; If the previous frame image of the current frame image tracks a hand, a first hand region in the current frame image is obtained, where the first hand region is a hand region in the current frame image predicted based on the positions of the hand joints in the previous frame image; The method further comprises: If the previous frame image of the current frame image does not track the hand, the current frame image is input into the hand detection model for inference to obtain the hand center point position and the second hand area of the current frame image; Calculate the maximum angle between the hand center point and the four sides of the second hand region to obtain a second transformation matrix; Projecting the second hand region through binocular parallax to obtain a second hand bounding box of the second hand region in three-dimensional space; Based on the second transformation matrix, the position of the hand joint point in the second hand bounding box is corrected, and a second direction vector of the corrected hand center point in the three-dimensional space is obtained; Obtain a second projection direction according to the product of the second transformation matrix and the second direction vector; The projection radius is calculated based on the maximum angle, and the image of the second hand area in the current frame image is projected and corrected using the projection radius and the second projection direction to generate a corrected second hand image.
8. The hand gesture recognition method according to claim 1 or claim 7, characterized in that: Projection of the hand area through binocular parallax, including: Calculate the overlap between the hand area corresponding to the left hand and the hand area corresponding to the right hand in the current frame image; When the overlap is less than a threshold, the hand area corresponding to the left hand and the hand area corresponding to the right hand are projected by binocular parallax; When the overlap is greater than or equal to the threshold and one of the left and right hands is tracked in the current frame image, the hand area corresponding to the other hand is projected by binocular parallax; The hand region is the first hand region or the second hand region.
9. A head mounted display device, characterized in that: It includes at least one processor, a memory, a camera module and a display module. The at least one processor is coupled to the memory and is used to read and execute instructions in the memory to perform the hand gesture recognition method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the hand gesture recognition method according to any one of claims 1 to 8 when running.
Citation Information
Patent Citations
Method and device for determining 6DOF pose based on 3D gesture recognition
CN114967943A
Calibration method, image correction method and device, electronic equipment and storage medium
CN116721161A
Facial expression driving method and device based on face 3D key points
CN118230394A
Method for detecting hand joints in glasses equipment, storage medium, electronic equipment and product
CN119672766A
Correction method of projection image, projector, projection system and program
JP2023179887A
Cited By
Finger joint point spatial position estimation method and device and head-mounted display equipment
CN121213646A