Gesture recognition method, device, electronic device and storage medium based on Lie group
Through the Lie group-based gesture recognition method, the forward kinematics model and nonlinear optimization technology are used to solve the problem of inconsistent shapes before and after gesture recognition results, and the accuracy and continuity of gesture recognition are achieved.
Patent Information
- Application Number
- CN202111057359.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-09
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2041-09-09
AI Technical Summary
Existing gesture recognition technology produces inconsistent gesture recognition results for the same user, resulting in discontinuous interactions.
A Lie group-based gesture recognition method is adopted. By acquiring continuous frame images captured by the camera, the key points of the hand are detected and the three-dimensional coordinates are calculated. The forward kinematics model and Lie group are used to describe the position transformation of the key points of the hand. Combined with the Jacobian matrix and nonlinear optimization, the parameters of the key points of the hand are output to ensure the consistency of the front and back hand shapes.
Improves the accuracy and continuity of gesture recognition, ensuring the stability and consistency of gesture interaction.
Smart Images

Figure CN115798031B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of image processing technology, and in particular to a gesture recognition method, device, electronic device, and storage medium based on Lie groups. Background Art
[0002] In the field of human-computer interaction, gesture interaction has become a hot topic and a key technical focus due to its natural and intuitive nature. Accurate and rapid hand pose estimation is a key technology in gesture interaction. Hand pose estimation has a wide range of applications. In smart homes, gestures can be made in front of a camera to control home appliances, such as fast-forwarding, rewinding, pausing, and playing videos on a display or TV. In AR / VR, the interaction of the user's hands with the real or virtual world greatly enhances the user's immersive experience. In sign language, hand movement recognition can make communication between hearing and hearing-mute people more convenient and accurate.
[0003] The data foundation of hand posture estimation technology is hand motion data. It can be said that the accuracy of hand motion data collection determines whether subsequent gesture interaction can be successfully implemented. In existing solutions, hand motion data can be obtained from a variety of different types of input devices, such as data gloves, accelerometers, touch screens, monocular cameras, multi-cameras, depth cameras, etc. Different types of input devices have their own advantages and disadvantages in obtaining hand motion data. Among them, monocular cameras and multi-cameras directly capture images, analyze the image information of bare hands from the images, obtain hand motion data, and then identify the type or semantics of the gesture and calculate the position of the hand in real three-dimensional space. In terms of ease of use, this technology does not require additional equipment and is more in line with human interaction habits.
[0004] When performing gesture recognition, the raw data collected by monocular and multi-camera cameras is 2D image data. After calculating the three-dimensional spatial position of the hand key points from the 2D image data, three-dimensional gesture posture estimation is performed through deep neural networks, etc. However, this estimation is completely based on the model trained with graphic samples. For the recognition results of continuous graphic processing of the same user, the consistency of the previous and subsequent hand shapes cannot be guaranteed. Summary of the Invention
[0005] The present invention provides a Lie group-based gesture recognition method, device, electronic device and storage medium to solve the technical problem of inconsistent shapes of gesture recognition results before and after for the same user.
[0006] In a first aspect, an embodiment of the present invention provides a gesture recognition method based on Lie groups, comprising:
[0007] Acquire continuous frame images captured by a camera, detect hand key points in the continuous frame images, and obtain a first key point position corresponding to each hand key point in each frame image;
[0008] Calculating the three-dimensional coordinates of each of the hand key points in a corresponding first coordinate system, where the first coordinate system is a three-dimensional coordinate system determined by adjacent key points of the hand key point away from the fingertips;
[0009] The three-dimensional coordinates are input into a preset forward kinematics model to obtain a pose transformation set of the hand key points corresponding to each frame of the image. The forward kinematics model is the origin of multiple three-dimensional coordinate systems connected by rigid bodies, based on the kinematic description of Lie groups. Among two adjacent hand key points, the hand key point farthest from the fingertips is used as the origin to determine a three-dimensional coordinate system. The two adjacent hand key points are connected by a rigid body, and the pose transformation set is the set of poses of the three-dimensional coordinate systems corresponding to all key points.
[0010] Calculating the Jacobian matrix corresponding to each frame of the image and the predicted three-dimensional coordinates corresponding to each hand key point in each frame of the image based on the pose transformation set, and calculating the corresponding second key point position based on the projection of the predicted three-dimensional coordinates in the image;
[0011] Based on the first key point position, the second key point position and the Jacobian matrix, nonlinearly optimize the pose transformation set;
[0012] According to the forward kinematics model and the optimized pose transformation set, preset hand key point parameters are output.
[0013] In a second aspect, an embodiment of the present invention further provides a gesture recognition device based on Lie groups, comprising:
[0014] A key point detection unit is used to obtain continuous frame images captured by the camera, detect hand key points in the continuous frame images, and obtain a first key point position corresponding to each hand key point in each frame image;
[0015] A coordinate calculation unit, configured to calculate the three-dimensional coordinates of each of the hand key points in a corresponding first coordinate system, where the first coordinate system is a three-dimensional coordinate system determined by adjacent key points of the hand key point away from the fingertips;
[0016] A model calculation unit is configured to input the three-dimensional coordinates into a preset forward kinematics model to obtain a pose transformation set corresponding to each frame of the hand key points; the forward kinematics model is the origin of multiple three-dimensional coordinate systems connected by rigid bodies, and is based on the kinematic description of Lie groups; a three-dimensional coordinate system is determined with the hand key point farther away from the fingertips as the origin, and the two adjacent hand key points are connected by a rigid body, and the pose transformation set is the set of poses of the three-dimensional coordinate systems corresponding to all key points;
[0017] a position calculation unit, configured to calculate the Jacobian matrix corresponding to each frame of the image and the predicted three-dimensional coordinates corresponding to each hand key point in each frame of the image based on the pose transformation set, and calculate the corresponding second key point position based on the projection of the predicted three-dimensional coordinates in the image;
[0018] a nonlinear optimization unit, configured to perform nonlinear optimization on the pose transformation set based on the first key point position, the second key point position, and the Jacobian matrix;
[0019] A parameter output unit is used to output preset hand key point parameters based on the forward kinematics model and the optimized pose transformation set.
[0020] In a third aspect, an embodiment of the present invention further provides an electronic device, including:
[0021] one or more processors;
[0022] a memory for storing one or more programs;
[0023] When the one or more programs are executed by the one or more processors, the electronic device implements the Lie group-based gesture recognition method as described in the first aspect.
[0024] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the Lie group-based gesture recognition method as described in the first aspect.
[0025] The above-mentioned gesture recognition method, device, electronic device and storage medium based on Lie group, in which the continuous frame images captured by the camera are obtained, the hand key points in the continuous frame images are detected to obtain the first key point position corresponding to each hand key point in each frame image; the three-dimensional coordinates of each hand key point in the corresponding first coordinate system are calculated, and the first coordinate system is a three-dimensional coordinate system determined by the adjacent key points of the hand key point away from the fingertips; the three-dimensional coordinates are input into a preset forward kinematics model to obtain the posture transformation set of the hand key point corresponding to each frame image, and the forward kinematics model is the origin of multiple three-dimensional coordinate systems connected by rigid bodies, based on the kinematic description of Lie group; two adjacent Among the hand key points, a three-dimensional coordinate system is determined with the hand key point farthest from the fingertips as the origin. Two adjacent hand key points are connected by a rigid body. The pose transformation set is the set of poses of the three-dimensional coordinate systems corresponding to all key points. Based on the pose transformation set, the Jacobian matrix corresponding to each frame of the image and the predicted three-dimensional coordinates corresponding to each hand key point in each frame of the image are calculated, and the corresponding second key point position is calculated based on the projection of the predicted three-dimensional coordinates in the image. Based on the first key point position, the second key point position, and the Jacobian matrix, the pose transformation set is nonlinearly optimized. Based on the forward kinematics model and the optimized pose transformation set, the preset hand key point parameters are output. The positions of the hand key points are detected through the image, and the pose transformation set is obtained and the positions of the hand key points are calculated based on the constructed forward kinematics model. The positions obtained by the two methods are used to optimize the pose transformation set. The relevant parameters of the hand in each frame of the image are finally confirmed based on the pose transformation set. The hand parameters obtained through the forward kinematics model ensure the consistency of the hand shape processing before and after, thereby improving the accuracy of the processing results. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 A flowchart of a gesture recognition method based on Lie groups provided by an embodiment of the present invention;
[0027] Figure 2 Schematic diagram of the key points and degrees of freedom of the hand;
[0028] Figure 3 Schematic diagram of rigid body transformation;
[0029] Figure 4 A schematic diagram of the recognition results of hand key points in a frame of image provided by an embodiment of the present invention;
[0030] Figure 5 A schematic diagram of the calculation process of the pose transformation set of a frame of image provided by an embodiment of the present invention;
[0031] Figure 6A structural diagram of a gesture recognition device based on Lie groups provided by an embodiment of the present invention;
[0032] Figure 7 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0033] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended to explain the present invention, not to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.
[0034] It should be noted that due to space limitations, this application specification does not enumerate all optional implementation methods. After reading this application specification, those skilled in the art should be able to understand that as long as the technical features do not contradict each other, any combination of technical features can constitute an optional implementation method.
[0035] Each embodiment is described in detail below.
[0036] Figure 1 A flowchart of a gesture recognition method based on Lie groups provided in an embodiment of the present invention is provided. The gesture recognition method based on Lie groups is used in an electronic device. As shown in the figure, the gesture recognition method based on Lie groups includes:
[0037] Step S110: Acquire continuous frame images captured by the camera, detect hand key points in the continuous frame images, and obtain the first key point position corresponding to each hand key point in each frame image.
[0038] In the gesture interaction scenario used in this solution, gesture recognition relies on continuous frame images captured by a camera. This can be a monocular camera or multiple cameras (for example, a binocular camera composed of two monocular cameras, or a combination of more monocular cameras). Each camera captures a continuous frame of images, and gesture interaction recognition is achieved by analyzing the image information of the bare hand in these continuous frames.
[0039] Gesture recognition is ultimately achieved through the transformation and judgment of the relative position relationship of the hand key points. Generally speaking, the hand key points are defined based on the hand bone structure. The human hand skeleton is mainly composed of carpal bones, metacarpal bones, and phalanges. The phalanges are composed of proximal phalanges, middle phalanges, and distal phalanges. The carpometacarpal joint (i.e. wrist joint), metacarpophalangeal joint (MCP), proximal interphalangeal joint (PIP), and distal interphalangeal joint (DIP) are formed between the carpal bones, metacarpal bones, proximal phalanges, middle phalanges, and distal phalanges. According to the physiological characteristics of the human palm, different joints have different degrees of freedom (DoF, Degree of Freedom). For example, it is generally believed that the MCP joint has 2DoF of flexion / extension and adduction / abduction, while the PIP and DIP joints only have 1DoF of flexion / extension. The key points of the hand actually correspond to the joints of the hand bones and the fingertips, such as Figure 2 As shown in FIG, the hand key points are represented by black dots. The number of hand key points of a normal and complete hand is 21, and the DoF of different joint types are also marked.
[0040] Specifically in this solution, the hand key points in the image are first obtained based on content detection. For example, a neural network is pre-trained and each frame of the image is input into the neural network to obtain the observation values of the hand key points. The observation values of the hand key points can also be obtained through computer vision methods, such as feature point extraction and feature point matching. The key point recognition based on neural networks or computer vision methods has been implemented in the prior art and will not be described in detail here.
[0041] Step S120: Calculate the three-dimensional coordinates of each of the hand key points in the corresponding first coordinate system, where the first coordinate system is a three-dimensional coordinate system determined by adjacent key points of the hand key point away from the fingertips.
[0042] The hand key point calculation process based on the forward kinematics model requires extracting the three-dimensional coordinates of each key point from each frame of the image. However, not all three-dimensional coordinates of the key points are based on the same three-dimensional coordinate system. Instead, each hand key point corresponds to a first coordinate system. In terms of the connection relationship between the hand key points, each hand key point has one or two adjacent hand key points, and one of them is far away from the fingertips. The first coordinate system corresponding to each hand key point is the three-dimensional coordinate system determined with the hand key point far away from the fingertips as the origin. In the specific calculation process, based on the previous multiple frames of images, the structure from motion calculation method is used to obtain the three-dimensional coordinates of each hand key point.
[0043] Step S130: Input the three-dimensional coordinates into a preset forward kinematics model to obtain a pose transformation set of the hand key points corresponding to each frame of the image. The forward kinematics model is the origin of multiple three-dimensional coordinate systems connected by rigid bodies, based on the kinematic description of Lie groups. Among two adjacent hand key points, the hand key point far away from the fingertips is used as the origin to determine a three-dimensional coordinate system. The two adjacent hand key points are connected by a rigid body. The pose transformation set is a collection of poses of the three-dimensional coordinate systems corresponding to all key points.
[0044] In this solution, in order to achieve gesture recognition, the ultimate goal is to obtain a description of the hand movement process. Based on the analysis of the hand structure, the hand movement process can be regarded as the rigid body movement of the key points of the hand connected by rigid bodies in three-dimensional space. The rigid body movement in three-dimensional space is described by the rotation matrix. R can be defined as SO(3), which is called a special orthogonal group. The motion of a rigid body in three-dimensional space consists of two parts: movement and rotation, which can be described as follows: the position of the motion reference system B fixed on the rigid body relative to the fixed coordinate system A is denoted by T ab (like Figure 3 Set the unknown vector from the origin of coordinate system A to the origin of moving coordinate system B. The posture of the moving coordinate system B relative to the fixed coordinate system A is R ab ∈SO(3), the system’s position is given by p ab and R ab Determine that the configuration space of the system is R ab The product space of SO(3) is denoted as SE(3). When the homogeneous expression is used, the rigid body transformation in three-dimensional space is denoted as T ab , expressed as a 4×4 matrix as follows:
[0045]
[0046] In set-related definitions, a Lie group is a (non-empty) A subset of G, which satisfies: (1) G is a group; (2) G is an embedded The manifold in (3); the multiplication operation and its inverse operation on the group are smooth. According to the definition of Lie group, SO(3) and SE(3) above both belong to Lie group.
[0047] In three-dimensional space, a rotation has a maximum of 3 degrees of freedom, and a rigid body transformation has a maximum of 6 degrees of freedom. SE(3) can express a rigid body transformation with 6 degrees of freedom (3DoF, rotation has 3DoF), and SO(3) expresses a rotation with 3DoF.
[0048] Every Lie group has a corresponding Lie algebra, which describes the local properties of the Lie group. The Lie algebra of SE(3) is defined in The six-dimensional vector on , where three dimensions represent translation and three dimensions represent rotation. The Lie algebra of SO(3) is defined in A three-dimensional vector on a Lie algebra represents a rotation. An important reason for introducing Lie algebras is that their exponential product mappings facilitate differential calculations. There is a mutual transformation (computational) relationship between Lie groups and their Lie algebras: Lie algebras are mapped to Lie groups via exponentials, and Lie groups are mapped to Lie algebras via logarithms.
[0049] The motion relationship between each key point of the hand is based on the rigid connection between each key point of the hand, and the relatively fixed characteristics of the rigid connection. Given the relative positions of the two adjacent rods that make up the kinematic pair, the position of the end hand is determined by forward kinematics.
[0050] Specifically for gesture pose verification, each joint is connected to a coordinate system, and the pose of that coordinate system relative to the reference coordinate system is the set of all these poses, called a pose transformation set. The fingertips (i.e., the fingertips) transform from their current position to their target position as each joint undergoes a pose transformation.
[0051] The forward kinematics model in this scheme is established based on the world coordinate system as a whole, which requires obtaining the camera model parameters, the initial gesture and the rigid body connection length between the hand key points; the forward kinematics model is based on the camera model parameters, the initial gesture and the rigid body connection length, and is obtained using the synthesis rules of the coordinate transformation of the hand key points.
[0052] Camera model parameters include intrinsic and extrinsic parameters. This patent is applicable to a variety of cameras besides depth cameras. If using one or more monocular cameras, the intrinsic calibration of each monocular camera must be completed in advance. If using multiple monocular cameras, the corresponding extrinsic calibration must also be completed in advance. For multi-camera systems, both intrinsic and extrinsic calibrations are already completed in advance. Camera calibration can be performed using various calibration plates, such as the AprilGrid or checkerboard calibration plates. Obtaining camera model parameters essentially involves obtaining the intrinsic parameters, or both intrinsic and extrinsic parameters, obtained during pre-calibration. The extrinsic parameters are the transformation matrices between each camera. In this solution, for multi-camera systems, one camera must be identified as the reference camera and its corresponding pose in the world coordinate system must be obtained. The initial gesture is the initialization gesture, typically corresponding to the rotation angles of all hand key points being zero. Alternatively, the reference gesture can be set to correspond to the state of a naturally extended hand. Alternatively, the reference gesture can be flexibly set based on the actual situation, and joint pose transformations are then generated based on this reference gesture. The initial gesture is described by the corresponding pose transformation set. The length of the rigid body connection between the key points of the hand is determined based on the basic size of the hand. For example, the average length between the joints of multiple hands is taken. The specific length can also be set according to the actual size of the current user.
[0053] When building a forward kinematics model, to better align it with the general laws of hand motion, the degrees of freedom of each hand keypoint are further determined, and motion constraints on these degrees of freedom are set within the forward kinematics model. The degrees of freedom of a hand keypoint are essentially the degrees of freedom of the joints corresponding to those keypoints. Due to the skeletal and muscular limitations of the human hand, different joints have different degrees of freedom. The current consensus is that the human wrist has 6 degrees of freedom, and the finger joints have no more than 2 degrees of freedom: abduction / adduction and flexion / extension. Figure 2 Here's a schematic diagram of the hand's key points defined as 26 DoF, for illustrative purposes only: a) Wrist joint: 6 DoF, i.e., 3 DoF for rotation + 3 DoF for translation. b) Index, middle, ring, and pinky fingers: These four fingers have the same degree of freedom, each with 4 DoF: 2 DoF (abduction / adduction, flexion / extension) at the MCP joint, 1 DoF (flexion / extension) at the PIP joint, and 1 DoF (flexion / extension) at the DIP joint. c) Thumb: 1 DoF at the thumb IP joint, 1 DoF at the thumb MCP joint, and 2 DoF at the thumb CMC joint, for a total of 4 DoF for the thumb.
[0054] It should be noted that there is no consensus in the industry on the number of DoF for the thumb. The main controversy focuses on the MCP and CMC joints of the thumb. Figure 2The 4DoF used in this paper is just an example choice. In practice, the CMC joint can have 3DoF and the MCP joint can have 2DoF. Moreover, 26DoF is just an optional solution. When designing the forward kinematics model, the maximum DoF can be designed for each joint. Figure 2 The design has a total of 26 DoFs. The vector of these 26 DoFs is called the state vector, which is the quantity to be optimized / estimated. Once the value of the state vector is estimated, the position of the key points between the fingers can be determined. Of course, the dimensions of the state vector will change accordingly depending on the DoF settings of the hand key points, but the basic calculation process remains the same.
[0055] The initial value of the pose transformation set of the currently processed image is the pose transformation set corresponding to the previous frame image; or, it is the pose transformation set initialized based on the currently processed image.
[0056] The pose transformation set includes the rotation angle and translation of each hand key point;
[0057] When initializing the pose transformation set, the rotation angle is initialized to 0. The translation of the hand key point corresponding to the wrist joint is calculated through multi-view geometry and motion recovery structure combined with the camera model parameters of the camera. The translation of other hand key points is calculated based on the translation of the hand key point corresponding to the wrist joint and the rigid body connection length of the hand key point.
[0058] The forward kinematics model in this scheme considers the carpal bones, metacarpal bones, proximal phalanges, mid-phalanges and distal phalanges of the hand as rigid links, among which the carpometacarpal joint connecting the carpal bones and metacarpal bones (referred to as the wrist joint or wrist joint in this article) is equivalent to a spherical joint, and the other joints are equivalent to revolute joints. Therefore, each finger of a person is composed of a set of rigid links connected by revolute joints, and a coordinate system is fixed at each joint to represent the relative posture of the joint. The relative postures of all the key points of the hand at the same moment constitute the posture transformation set at that moment. Each element in the posture transformation set is a three-dimensional rigid body transformation represented by a Lie group, which is used to indicate the rotation angle and translation amount of each key point of the hand (for example, the rotation angle of the corresponding joints of the wrist, fingers, etc. around the XYZ axis).
[0059] During the specific calculation process, considering that the rotation angle and translation can be calculated to obtain a homogeneous transformation matrix, and the homogeneous transformation matrix can also be used to calculate the rotation angle and translation, this solution can express hand movements based on the rotation angle and translation, or it can express hand movements through the homogeneous transformation matrix. Generally speaking, the rotation angle and translation can be more intuitively presented as the rotation angle around the XYZ axis and the translation along the XYZ axis, respectively. In this embodiment, the rotation angle and translation are specifically used for calculation.
[0060] Step S140: Based on the pose transformation set, calculate the Jacobian matrix corresponding to each frame image and the predicted three-dimensional coordinates corresponding to each hand key point in each frame image, and calculate the corresponding second key point position based on the projection of the predicted three-dimensional coordinates in the image.
[0061] When the hand moves, mathematically, the rotation angle and translation values are perturbed based on the initial gesture, and the values of the pose transformation set change accordingly. During this process, the fingertips move with the transformation of each joint. That is, the end of the skeleton transforms to the target position based on the initial position of the pose transformation set. Therefore, based on the initial pose and relative pose, the position of each key point of the hand can be calculated using the synthesis rules of coordinate transformation. In this case, the predicted 3D coordinates are obtained. Based on the camera model parameters, the predicted 3D coordinates are projected onto the corresponding image to obtain the second key point pose. Simultaneously, the corresponding Jacobian matrix is calculated based on the current pose transformation set.
[0062] Step S150: Based on the first key point position, the second key point position and the Jacobian matrix, nonlinearly optimize the pose transformation set.
[0063] During the calculation, the position of the first key point (i.e., the position of the hand key point directly identified from the image) and the position of the second key point (i.e., the position of the hand key point obtained by calculation) cannot completely coincide, that is, the difference between the two is not 0. At this time, it is necessary to optimize this calculation error to obtain the optimal posture parameters and the coordinates of the three-dimensional space point. Specifically, this error is handled through nonlinear optimization.
[0064] In the specific implementation process, step S150 can be specifically implemented through steps S151 to S154:
[0065] Step S151: Calculate a reprojection error using the first key point position and the second key point position.
[0066] The reprojection error is the L2 norm of the first keypoint position and the second keypoint position.
[0067] Step S152: Calculate the temporal error using the pose transformation set between the currently processed image and the previous frame image.
[0068] The time domain error is expressed by e temporal =||θ t -θ t-1 || 2 Calculate, where e temporal represents the time domain error, θ t Represents the Lie algebra corresponding to the pose transformation set of the currently processed image; θ t-1Represents the Lie algebra corresponding to the pose transformation set of the previous frame image.
[0069] Step S153: Calculate the sum of the reprojection error and the time domain error to obtain an error vector.
[0070] Step S154: performing nonlinear optimization on the pose transformation set according to the error vector and the Jacobian matrix.
[0071] The error vector and Jacobian matrix are input into a preset nonlinear optimization model, and the optimized pose transformation set is output. In the specific processing process, the pose transformation set of each frame image is calculated independently. This calculation method may cause jitter in the pose calculation results. Therefore, based on the optimized pose transformation set corresponding to the current multiple frames of images, the pose transformation set of a preset number of images before and after is obtained with reference to the currently processed image. Based on the obtained pose transformation set, the pose transformation set of the currently processed image is smoothed in the time domain. Specifically, with the currently processed image as the center, the pose transformation set of a preset number of images before and after is taken, and then the rotation angle and translation amount are processed using an adaptive low-pass filtering algorithm to alleviate jitter and achieve time domain smoothing.
[0072] Specifically in this solution, the calculation formulas for the error vector and the Jacobian matrix are nonlinear, so the corresponding optimization problem can be converted into a nonlinear least squares problem and solved using a nonlinear optimization algorithm. The commonly used definition of a nonlinear least squares problem is:
[0073] Given a vector function f, (which contains n unknown parameters and m observations, where m ≥ n), our goal is to find the minimum value of ‖f(x)‖, which is equivalent to finding
[0074]
[0075] in The vector function f is composed of the error between the computational effort and the observed quantity for each keypoint of the hand. x is the unknown vector to be optimized, representing the pose of the hand joints. This can be understood as x being the state vector in the robot's state estimation problem. Based on this, commonly used nonlinear least squares methods can be employed, such as gradient descent, Newton's method, Gauss-Newton method, and Levenberg–Marquardt algorithm.
[0076] Step S160: Outputting preset hand key point parameters according to the forward kinematics model and the optimized pose transformation set.
[0077] Based on the determined pose transformation set, various preset hand keypoint parameters can be output based on the forward kinematics model as needed. These preset hand keypoint parameters may include the 3D coordinates of the keypoints, their rotation angles and translations, and one or more of the pose transformation sets for the hand keypoints. Calculating these various parameters using the forward kinematics model, based on the determined pose transformation set, is a standard forward kinematics calculation process.
[0078] Please refer to Figure 4 and Figure 5 , the following is combined Figure 4 and Figure 5 A description of the specific recognition of a frame of image by the scheme. For the current processed image, there are Figure 4 This is a schematic diagram of the hand key points detected in the current processing image, where there are 21 corresponding numbers 0-20, and the position corresponding to each hand key point is the first key point position. Figure 5 As shown, assuming that the number of the current processed image in the entire continuous frame image is t, the corresponding first key point position is expressed as i represents the number of the hand key point, t represents the t-th frame image currently being calculated, (u, v) represents the position coordinates of the hand key point in the image, and 2d1 represents the position of the first key point identified. During initialization, as mentioned above, the camera model parameters, the initial gesture, and the rigid body connection length between the hand key points are initialized. The camera model parameters include the intrinsic parameters of each monocular camera, which are used to convert the position (x, y, z) in three-dimensional space into (u, v) of a 2D image, as well as the poses of the left and right cameras relative to the world coordinate system, respectively, denoted as w T cL , w T cR .
[0079] Forward kinematic modeling based on Lie groups actually determines the kinematic relationship of each transformation matrix in the pose transformation set, and uses the synthesis rule to obtain the pose relative to the reference coordinate system. Figure 5 The middle finger is used as an example for detailed explanation. The methods for other fingers are the same. Figure 4 The pose of the hand key point numbered 0 in the camera coordinate system and initial value r stands for wrist, c stands for camera, Indicates the position of the wrist joint relative to the camera when the rotation angle of the wrist joint is 0. For the position transformation of the phalanx of the middle finger, r T mcp Indicates the relative position relationship between MCP and wrist joint (corresponding to Figure 4The position of the metacarpophalangeal joint (MCP) numbered 9 relative to the wrist joint coordinate system); Indicates the relative position relationship of the MCP relative to the wrist when the rotation angle of the MCP joint is 0; mcp T pip Indicates the relative position relationship between PIP and MCP, Indicates the relative position relationship of PIP relative to MCP when the rotation angle of PIP joint motion is 0; pip T dip Indicates the relative position relationship between DIP and PIP, Indicates the relative position relationship of DIP relative to PIP when the rotation angle of DIP joint is 0. In addition, it is also necessary to obtain the middle fingertip (corresponding to Figure 4 Number 12) Now for the three-dimensional coordinates of the distal finger joint coordinate system dip p tip Next, we can calculate the position of the hand key points (i.e., the end-effector in the robot) relative to the wrist (shown in 3D1) based on the constructed forward kinematics model. Taking the tip of the middle finger as an example, using the synthesis rule and the definition of three-dimensional rigid body transformation, the coordinates of the fingertip in the camera coordinate system are:
[0080] c p tip = c T r · r T mcp · mcp T pip · pip T dip · dip p tip
[0081] The calculation method for other key points is similar. If there are multiple monocular cameras, the above formula only needs to be multiplied by the camera extrinsic parameter to obtain the coordinates of the fingertip in the corresponding camera coordinate system. The above formula shows that the end of the bone moves with the movement of the preset joint point; the entire hand skeleton can move according to this pose transformation set to reach the target position from the initial position. Based on this calculation process, the three-dimensional coordinates of all 21 hand key points can be obtained (shown as 3d2). If there are two monocular cameras, 21×2=42 three-dimensional coordinates can be calculated and integrated for nonlinear optimization.
[0082] According to forward kinematics, the hand pose is further calculated. When the hand moves relative to the initial gesture, mathematically corresponding disturbances are generated, and the pose transformation set needs to be updated:
[0083]
[0084] Since the Lie group SE(3) is used to represent the posture in this scheme, δ(·) represents the perturbation on the Lie group, which corresponds to matrix multiplication. That is, δ(·) in the formula represents the perturbation amount.
[0085] According to the camera projection transformation (represented by Proj()), the 2D position of the fingertip in the image plane (shown as 2d2) is calculated:
[0086] uv p tip =Proj( c p tip )
[0087] If there are multiple cameras, the projection transformation of the corresponding camera can be used. The positions of other hand key points are calculated in the same way.
[0088] Calculating the Jacobian matrix of the pose transformation set represented by the Lie group has been implemented. Specifically in this solution, for each finger, the following Jacobian matrix needs to be calculated:
[0089] J p :express uv p tip Relative to c p tip The Jacobian matrix of ;
[0090] r J p :express c p tip Relative to the carpometacarpal joint position c T r The Jacobian matrix of ;
[0091] mcp J r : Indicates the position of the carpometacarpal joint c T r Relative to the metacarpophalangeal joint r T mcp The Jacobian matrix of ;
[0092] pip J mcp :Indicates the metacarpophalangeal joint posture r T mcp Relative to mcp T pip The Jacobian matrix of ;
[0093] dip J pip :express mcp T pip Relative to pip T dip The Jacobian matrix of .
[0094] Based on the obtained position of the hand key points and the Jacobian matrix, the error is further processed. The error function is defined as the sum of the reprojection error and the time domain error E = e 2D +e temporal . Reprojection error: e 2D It represents the difference between the calculated updated position of the hand key point (i.e., the second key point position) and the detected position of the hand key point (i.e., the first key point position), and the updated error vector (shown as 2d3). represents the 2D coordinates of the i-th key point in the image space of the j-th camera, obtained by detection; uv p i The calculation is based on the camera projection transformation, as mentioned above, where Proj j () represents the function of projecting a point in 3D space to the image space of the jth camera. Time domain error: e temporal =||θ t -θ t-1 || 2 Here, the state variable θ contains 26 degrees of freedom (DoF), of which 23 components are rotation angles and 3 are translations. In the Lie group representation, θ is obtained by taking the logarithm of each element in the pose transformation set 1, that is, θ is the Lie algebra corresponding to the Lie group.
[0095] Based on the Jacobian matrix and error vector calculated in the previous article, the state vector of the key points of the hand is optimized and solved: specifically, the error and Jacobian function are put into a general optimization algorithm (such as the nonlinear optimization algorithm Levenberg-Marquardt algorithm, Gauss-Newton algorithm) as input, and the iteration termination condition is set to obtain the posture transformation set after optimization and update (shown in 3d3).
[0096] The optimized posture transformation set is obtained by independently processing a single frame image, which may also cause jitter in the description of the posture between multiple needle images. For this reason, a time domain window of length k frames is set in this scheme. After the optimization calculation of the t-th frame image, a total of k frames of images are taken before and after it (for example, the posture transformation set 3d3` optimized for the t-1 frame and the posture transformation set 3d3 optimized for the t-th frame) are taken as the center. Then, an adaptive low-pass filtering algorithm is used to process the rotation and translation to obtain a smoothed posture transformation set 3d4 to alleviate jitter. Here, taking the first-order low-pass filter as an example, the following formula is applied to each component of the 26 DoFs:
[0097]
[0098] The value of α is between 0.2 and 0.9, which is determined according to the pixel ratio of the hand key points to the entire 2D image.
[0099] After the smoothing process is completed, a variety of values can be output, including but not limited to (the following only takes the middle finger as an example):
[0100] The three-dimensional coordinates of the fingertip in the world coordinate system: in, w T c Represents the camera's pose in the world coordinate system.
[0101] The three-dimensional coordinates of the distal finger joint: in To take the updated pose transformation set The same applies to other key points.
[0102] The three-dimensional coordinates of the proximal knuckle: in To take the updated pose transformation set The amount of translation is enough;
[0103] The same applies to the metacarpophalangeal joints and wrist joints.
[0104] The rotation and translation of the joint point, that is, θ, is obtained by taking the logarithm of each element in the updated posture transformation set 1, that is, θ is the Lie algebra corresponding to the Lie group.
[0105] The updated pose transformation set represents the transformation of the hand from the initial pose to the updated pose.
[0106] The position of the key points of the hand on the 2D image: just perform a projection transformation on the 3D points obtained above. Take the distal finger joint as an example:
[0107] This solution does not require model training based on a large amount of sample data collection. Accurate recognition results can be obtained through the hand motion model constructed based on forward kinematics, which improves processing efficiency and accuracy.
[0108] The above method obtains the continuous frame images captured by the camera, detects the hand key points in the continuous frame images, and obtains the first key point position corresponding to each hand key point in each frame image; calculates the three-dimensional coordinates of each of the hand key points in the corresponding first coordinate system, and the first coordinate system is a three-dimensional coordinate system determined by the adjacent key points of the hand key point away from the fingertips; inputs the three-dimensional coordinates into a preset forward kinematics model to obtain the posture transformation set of the hand key point corresponding to each frame image, and the forward kinematics model is the origin of multiple three-dimensional coordinate systems connected by rigid bodies, based on the kinematic description of Lie groups; among two adjacent hand key points, the hand away from the fingertips is the first key point. A three-dimensional coordinate system is determined with the hand key points as the origin, and two adjacent hand key points are connected by a rigid body. The pose transformation set is the set of poses of the three-dimensional coordinate system corresponding to all key points; based on the pose transformation set, the Jacobian matrix corresponding to each frame image and the predicted three-dimensional coordinates corresponding to each hand key point in each frame image are calculated, and the corresponding second key point position is calculated based on the projection of the predicted three-dimensional coordinates in the image; based on the first key point position, the second key point position and the Jacobian matrix, the pose transformation set is nonlinearly optimized; based on the forward kinematics model and the optimized pose transformation set, the preset hand key point parameters are output. The positions of the hand key points are detected through the image, and the pose transformation set is obtained and the positions of the hand key points are calculated based on the constructed forward kinematics model. The positions obtained by the two methods are used to optimize the pose transformation set. The relevant parameters of the hand in each frame image are finally confirmed based on the pose transformation set. The hand parameters obtained through the forward kinematics model ensure the consistency of the hand shape processing before and after, thereby improving the accuracy of the processing results.
[0109] Figure 6 A structural diagram of a gesture recognition device based on Lie groups provided by an embodiment of the present invention. The gesture recognition device based on Lie groups is used in electronic equipment, referring to Figure 6 The gesture recognition device based on Lie group includes a key point detection unit 210, a coordinate calculation unit 220, a model calculation unit 230, a position calculation unit 240, a nonlinear optimization unit 250 and a parameter output unit 260.
[0110] Among them, the key point detection unit 210 is used to obtain the continuous frame images collected by the camera, detect the hand key points in the continuous frame images, and obtain the first key point position corresponding to each hand key point in each frame image; the coordinate calculation unit 220 is used to calculate the three-dimensional coordinates of each of the hand key points in the corresponding first coordinate system, and the first coordinate system is a three-dimensional coordinate system determined by the adjacent key points of the hand key point away from the fingertips; the model calculation unit 230 is used to input the three-dimensional coordinates into a preset forward kinematics model to obtain the posture transformation set of the hand key point corresponding to each frame image; the forward kinematics model is the origin of multiple three-dimensional coordinate systems connected by rigid bodies, based on the kinematic description of Lie groups; among two adjacent hand key points, the hand away from the fingertips is the first key point of the hand key point. A three-dimensional coordinate system is determined with the key point as the origin, and two adjacent hand key points are connected by a rigid body. The pose transformation set is a set of poses of the three-dimensional coordinate system corresponding to all key points; a position calculation unit 240 is used to calculate the Jacobian matrix corresponding to each frame image and the predicted three-dimensional coordinates corresponding to each hand key point in each frame image according to the pose transformation set, and calculate the corresponding second key point position according to the projection of the predicted three-dimensional coordinates in the image; a nonlinear optimization unit 250 is used to perform nonlinear optimization on the pose transformation set based on the first key point position, the second key point position and the Jacobian matrix; a parameter output unit 260 is used to output preset hand key point parameters according to the forward kinematics model and the optimized pose transformation set.
[0111] Based on the above embodiment, the device further includes:
[0112] The time domain smoothing unit is used to obtain the posture transformation set of a preset number of images before and after the currently processed image with reference to the currently processed image, and perform time domain smoothing on the posture transformation set of the currently processed image based on the obtained posture transformation set.
[0113] Based on the above embodiment, the time domain smoothing process is performed by low-pass filtering.
[0114] Based on the above embodiment, the device further includes:
[0115] Parameter acquisition unit, used to obtain camera model parameters, initial gestures and rigid body connection lengths between hand key points;
[0116] The forward kinematics model is obtained based on the camera model parameters, the initial gesture and the rigid body connection length using a synthesis rule of coordinate transformation of the hand key points.
[0117] Based on the above embodiment, the device further includes:
[0118] A degree of freedom acquisition unit, used to obtain the degree of freedom of each key point of the hand;
[0119] The forward kinematics model is provided with motion state constraints on the degrees of freedom.
[0120] Based on the above embodiment, the initial value of the pose transformation set of the currently processed image is the pose transformation set corresponding to the previous frame image; or, it is the position transformation set initialized based on the currently processed image.
[0121] Based on the above embodiment, the posture transformation set includes the rotation angle and translation amount of each hand key point;
[0122] When initializing the pose transformation set, the rotation angle is initialized to 0. The translation of the hand key point corresponding to the wrist joint is calculated through multi-view geometry and motion recovery structure combined with the camera model parameters of the camera. The translation of other hand key points is calculated based on the translation of the hand key point corresponding to the wrist joint and the rigid body connection length of the hand key point.
[0123] Based on the above embodiment, the nonlinear optimization unit includes:
[0124] a reprojection error calculation module, configured to calculate a reprojection error using the first key point position and the second key point position;
[0125] A temporal error calculation module is used to calculate the temporal error through the pose transformation set of the current processed image and the previous frame image;
[0126] An error vector calculation module, configured to calculate the sum of the reprojection error and the time domain error to obtain an error vector;
[0127] An optimization processing module is used to perform nonlinear optimization on the pose transformation set according to the error vector and the Jacobian matrix.
[0128] Nonlinear optimization unit, the time domain error is obtained by e temporal =||θ t -θ t-1 || 2 Calculate, where e temporal represents the time domain error, θ t Represents the Lie algebra corresponding to the pose transformation set of the currently processed image; θ t-1 Represents the Lie algebra corresponding to the pose transformation set of the previous frame image.
[0129] The nonlinear optimization unit, the optimization processing module, is specifically used to input the error vector and Jacobian matrix into a preset nonlinear optimization model and output an optimized posture transformation set.
[0130] The nonlinear optimization unit presets the hand key point parameters including the 3D coordinates of the key points, the rotation angle and translation amount of the hand key points, and one or more of the posture transformation set of the hand key points.
[0131] The Lie group-based gesture recognition device provided in an embodiment of the present invention is included in an electronic device of a device and can be used to execute any of the Lie group-based gesture recognition methods provided in the above embodiments, and has corresponding functions and beneficial effects.
[0132] It is worth noting that in the above-mentioned embodiment of the gesture recognition device based on Lie groups, the various units and modules included are divided only according to functional logic, but are not limited to the above-mentioned division, as long as they can achieve the corresponding functions; in addition, the specific names of the various functional units are only for the convenience of distinguishing them from each other and are not intended to limit the scope of protection of the present invention.
[0133] Figure 7 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 7 As shown, the electronic device includes a processor 310, a memory 320, an input device 330, an output device 340, and a communication device 350; the number of processors 310 in the electronic device can be one or more. Figure 7 In the figure, a processor 310 is used as an example; the processor 310, memory 320, input device 330, output device 340 and communication device 350 in the electronic device can be connected via a bus or other means. Figure 7 The bus connection is taken as an example.
[0134] The memory 320, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the Lie group-based gesture recognition method in the embodiments of the present invention (e.g., the key point detection unit 210, coordinate calculation unit 220, model calculation unit 230, position calculation unit 240, nonlinear optimization unit 250, and parameter output unit 260 in the Lie group-based gesture recognition device). The processor 310 executes the software programs, instructions, and modules stored in the memory 320 to execute various functional applications and data processing of the electronic device, thereby implementing the aforementioned Lie group-based gesture recognition method.
[0135] The memory 320 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 320 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 320 may further include a memory remotely located relative to the processor 310, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0136] The input device 330 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the electronic device. The output device 340 may include a display device such as a display screen.
[0137] The electronic device includes a Lie group-based gesture recognition device, which can be used to execute any Lie group-based gesture recognition method and has corresponding functions and beneficial effects.
[0138] An embodiment of the present invention also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform relevant operations in the Lie group-based gesture recognition method provided in any embodiment of the present application, and have corresponding functions and beneficial effects.
[0139] Those skilled in the art should understand that the embodiments of the present application may be provided as methods, systems, or computer program products.
[0140] Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0141] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-permanent storage in a computer-readable medium, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0142] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0143] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0144] Note that the above are only preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the concept of the present invention. The scope of the present invention is determined by the scope of the appended claims.
Claims
1. A gesture recognition method based on Lie groups, characterized in that: include: Acquire continuous frame images captured by a camera, detect hand key points in the continuous frame images, and obtain a first key point position corresponding to each hand key point in each frame image; Calculating the three-dimensional coordinates of each of the hand key points in a corresponding first coordinate system, where the first coordinate system is a three-dimensional coordinate system determined by adjacent key points of the hand key point away from the fingertips; The three-dimensional coordinates are input into a preset forward kinematics model to obtain a pose transformation set of the hand key points corresponding to each frame of the image. The forward kinematics model is the origin of multiple three-dimensional coordinate systems connected by rigid bodies, based on the kinematic description of Lie groups. Among two adjacent hand key points, the hand key point farthest from the fingertips is used as the origin to determine a three-dimensional coordinate system. The two adjacent hand key points are connected by a rigid body, and the pose transformation set is the set of poses of the three-dimensional coordinate systems corresponding to all key points. Calculating the Jacobian matrix corresponding to each frame of the image and the predicted three-dimensional coordinates corresponding to each hand key point in each frame of the image based on the pose transformation set, and calculating the corresponding second key point position based on the projection of the predicted three-dimensional coordinates in the image; Based on the first key point position, the second key point position and the Jacobian matrix, nonlinearly optimize the pose transformation set; According to the forward kinematics model and the optimized pose transformation set, preset hand key point parameters are output.
2. The gesture recognition method according to claim 1, characterized in that: After performing nonlinear optimization on the pose transformation set based on the first key point position, the second key point position and the Jacobian matrix, the method further includes: Taking the currently processed image as a reference, obtain the pose transformation sets of a preset number of images before and after, and perform time domain smoothing on the pose transformation set of the currently processed image based on the obtained pose transformation sets.
3. The gesture recognition method according to claim 2, characterized in that: The time domain smoothing process is performed by low-pass filtering.
4. The gesture recognition method according to claim 1, wherein: Before presetting the forward kinematics model, the following steps are also included: Get the camera model parameters, initial gesture and the rigid body connection length between the hand key points; The forward kinematics model is obtained based on the camera model parameters, the initial gesture and the rigid body connection length using a synthesis rule of coordinate transformation of the hand key points.
5. The gesture recognition method according to claim 4, characterized in that: Before presetting the forward kinematics model, the following steps are also included: Get the degrees of freedom of each hand key point; The forward kinematics model is provided with motion state constraints on the degrees of freedom.
6. The gesture recognition method according to claim 1, wherein: The initial value of the pose transformation set of the currently processed image is the pose transformation set corresponding to the previous frame image; or, it is the pose transformation set initialized based on the currently processed image.
7. The gesture recognition method according to claim 6, characterized in that: The pose transformation set includes the rotation angle and translation of each hand key point; When initializing the pose transformation set, the rotation angle is initialized to 0. The translation of the hand key point corresponding to the wrist joint is calculated through multi-view geometry and motion recovery structure combined with the camera model parameters of the camera. The translation of other hand key points is calculated based on the translation of the hand key point corresponding to the wrist joint and the rigid body connection length of the hand key point.
8. The gesture recognition method according to claim 1, wherein: The performing nonlinear optimization on the pose transformation set based on the first key point position, the second key point position and the Jacobian matrix includes: Calculating a reprojection error using the first key point position and the second key point position; Calculate the temporal error by comparing the pose transformation between the current processed image and the previous frame image; Calculating the sum of the reprojection error and the time domain error to obtain an error vector; The pose transformation set is nonlinearly optimized according to the error vector and the Jacobian matrix.
9. The gesture recognition method according to claim 8, characterized in that: The time domain error is expressed by e temporal =||θ t -θ t-1 || 2 Calculate, where e temporal represents the time domain error, θ t Represents the Lie algebra corresponding to the pose transformation set of the currently processed image; θ t-1 Represents the Lie algebra corresponding to the pose transformation set of the previous frame image.
10. The gesture recognition method according to claim 8, characterized in that: The nonlinear optimization of the pose transformation set is performed according to the error vector and the Jacobian matrix, specifically: The error vector and Jacobian matrix are input into a preset nonlinear optimization model, and an optimized pose transformation set is output.
11. The gesture recognition method according to claim 1, wherein: The preset hand key point parameters include one or more of the 3D coordinates of the key points, the rotation angle and translation amount of the hand key points, and the posture transformation set of the hand key points.
12. A gesture recognition device based on Lie group, characterized in that: include: A key point detection unit is used to obtain continuous frame images captured by the camera, detect hand key points in the continuous frame images, and obtain a first key point position corresponding to each hand key point in each frame image; A coordinate calculation unit, configured to calculate the three-dimensional coordinates of each of the hand key points in a corresponding first coordinate system, where the first coordinate system is a three-dimensional coordinate system determined by adjacent key points of the hand key point away from the fingertips; A model calculation unit is configured to input the three-dimensional coordinates into a preset forward kinematics model to obtain a pose transformation set corresponding to each frame of the hand key points; the forward kinematics model is the origin of multiple three-dimensional coordinate systems connected by rigid bodies, and is based on the kinematic description of Lie groups; a three-dimensional coordinate system is determined with the hand key point farther away from the fingertips as the origin, and the two adjacent hand key points are connected by a rigid body, and the pose transformation set is the set of poses of the three-dimensional coordinate systems corresponding to all key points; a position calculation unit, configured to calculate the Jacobian matrix corresponding to each frame of the image and the predicted three-dimensional coordinates corresponding to each hand key point in each frame of the image based on the pose transformation set, and calculate the corresponding second key point position based on the projection of the predicted three-dimensional coordinates in the image; a nonlinear optimization unit, configured to perform nonlinear optimization on the pose transformation set based on the first key point position, the second key point position, and the Jacobian matrix; A parameter output unit is used to output preset hand key point parameters based on the forward kinematics model and the optimized pose transformation set.
13. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the electronic device implements the Lie group-based gesture recognition method according to any one of claims 1 to 11.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the gesture recognition method based on Lie groups as described in any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Three-dimensional coordinate positioning method and device for key point, electronic equipment and storage medium
CN110443154A
Hand key point detection method, gesture recognition method and related devices
CN110991319A