Screw-based gesture recognition method and device, electronic equipment and storage medium
By using a screw-based gesture recognition method, the pose transformation of key hand points is described using a forward kinematics model and screw method, which solves the problem of inconsistent shapes in the gesture recognition results and achieves accuracy and continuity in gesture recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU SHIYUAN ELECTRONICS CO LTD
- Filing Date
- 2021-09-09
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, the shape of the gesture recognition results for the same user is inconsistent, resulting in insufficient continuity and accuracy of gesture interaction.
A screw-based gesture recognition method is adopted. By acquiring continuous frame images captured by the camera, key points of the hand are detected, three-dimensional coordinates are calculated, and the pose transformation of key points of the hand is described by using a forward kinematics model and screw method. Combined with Jacobian matrix and nonlinear optimization techniques, the consistency of hand shape is ensured.
It improves the accuracy and continuity of gesture recognition, ensures consistency in the processing of hand shapes before and after gestures, and enhances the reliability of gesture interaction.
Smart Images

Figure CN115798030B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of image processing technology, and in particular to a gesture recognition method, apparatus, electronic device and storage medium based on spinor. Background Technology
[0002] In the field of human-computer interaction, gesture interaction has become a research hotspot and key technology due to its natural and intuitive advantages. Accurate and rapid hand pose estimation is one of the crucial technologies for gesture interaction. The applications of hand pose estimation are extensive. In the smart home field, users can control home appliances by making corresponding gestures in front of a camera, such as controlling the fast forward, rewind, pause, and play of videos on a display screen or television. In the AR / VR field, the interaction between users' hands and the real or virtual world greatly enhances the immersive experience. In the field of sign language, hand gesture recognition makes communication between sighted and deaf individuals more convenient and accurate.
[0003] The data foundation for hand pose estimation technology is hand motion data. The accuracy of hand motion data acquisition directly determines the smooth implementation of subsequent gesture interactions. In existing solutions, hand motion data can come from various types of input devices, such as data gloves, accelerometers, touchscreens, monocular cameras, multi-view cameras, and depth cameras. Each type of input device has its own advantages and disadvantages in acquiring hand motion data. Monocular and multi-view cameras directly acquire images, analyze the image information of the bare hand from the images to obtain hand motion data, and then identify the type or semantics of the gesture, calculating the hand's position in real three-dimensional space. From a user convenience perspective, this technology does not require additional equipment and is more in line with human interaction habits.
[0004] When performing gesture recognition, the raw data collected by monocular and multi-view cameras is 2D image data. After calculating the three-dimensional spatial position of key hand points from the 2D image data, three-dimensional gesture posture estimation is performed through deep neural networks. However, this estimation is entirely based on the model trained on the image samples. For the recognition results of continuous image processing of the same user, the consistency of the hand shape before and after cannot be guaranteed. Summary of the Invention
[0005] This invention provides a gesture recognition method, device, electronic device, and storage medium based on spinor to solve the technical problem of inconsistent shapes in the gesture recognition results of the same user.
[0006] In a first aspect, embodiments of the present invention provide a gesture recognition method based on spinor, comprising:
[0007] Acquire consecutive frame images captured by the camera, detect key hand points in the consecutive frame images, and obtain the position of the first key point corresponding to each key hand point in each frame image;
[0008] Calculate the three-dimensional coordinates of each of the hand key points in the corresponding first coordinate system, wherein the first coordinate system is a three-dimensional coordinate system determined by the adjacent key points of the hand key points away from the fingertips;
[0009] The three-dimensional coordinates are input into a preset forward kinematics model to obtain the pose transformation set of each frame image corresponding to the hand key points. The forward kinematics model consists of multiple hand key points connected by rigid bodies and is described by screw kinematics. In the forward kinematics model, the pose of the hand key point corresponding to the wrist is described in the camera coordinate system, and the pose of the hand key point corresponding to the metacarpophalangeal joint is represented by the screw method based on the exponential product formula and composition rules of rigid body motion, relative to the hand key point corresponding to the wrist. The pose transformation set is the set of poses corresponding to all hand key points.
[0010] Based on the pose transformation set, calculate the Jacobian matrix and the predicted 3D coordinates of the hand key points corresponding to each frame image, and calculate the position of the corresponding second key point according to the projection of the predicted 3D coordinates in the image.
[0011] Based on the positions of the first key point, the second key point, and the Jacobian matrix, the pose transformation set is nonlinearly optimized.
[0012] Based on the positive kinematics model and the optimized pose transformation set, preset hand key point parameters are output.
[0013] Secondly, embodiments of the present invention also provide a gesture recognition device based on spin, comprising:
[0014] The key point detection unit is used to acquire continuous frame images captured by the camera, detect hand key points in the continuous frame images, and obtain the first key point position corresponding to each hand key point in each frame image.
[0015] The coordinate calculation unit is used to calculate the three-dimensional coordinates of each of the hand key points in the corresponding first coordinate system, wherein the first coordinate system is a three-dimensional coordinate system determined by the adjacent key points of the hand key points away from the fingertips;
[0016] The model calculation unit is used to input the three-dimensional coordinates into a preset forward kinematics model to obtain the pose transformation set of each frame image corresponding to the hand key points. The forward kinematics model consists of multiple hand key points connected by rigid bodies and is described by screw kinematics. In the forward kinematics model, the pose of the hand key point corresponding to the wrist is described in the camera coordinate system, and the pose of the hand key point corresponding to the metacarpophalangeal joint is represented by the screw method based on the exponential product formula and composition rules of rigid body motion, relative to the hand key point corresponding to the wrist. The pose transformation set is the set of poses corresponding to all hand key points.
[0017] The position calculation unit is used to calculate the Jacobian matrix and the predicted three-dimensional coordinates of the hand key points corresponding to each frame image based on the pose transformation set, and to calculate the position of the corresponding second key point according to the projection of the predicted three-dimensional coordinates in the image.
[0018] A nonlinear optimization unit is used to perform nonlinear optimization on the pose transformation set based on the first key point position, the second key point position, and the Jacobian matrix.
[0019] The parameter output unit is used to output preset hand key point parameters based on the positive kinematics model and the optimized pose transformation set.
[0020] Thirdly, embodiments of the present invention also provide an electronic device, comprising:
[0021] One or more processors;
[0022] Memory, used to store one or more programs;
[0023] When the one or more programs are executed by the one or more processors, the electronic device enables the spinor-based gesture recognition method as described in the first aspect.
[0024] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the spinor-based gesture recognition method as described in the first aspect.
[0025] The aforementioned screw-based gesture recognition method, device, electronic device, and storage medium, in this method, acquire consecutive frame images captured by a camera, detect hand keypoints in the consecutive frame images to obtain the position of a first keypoint corresponding to each hand keypoint in each frame image; calculate the three-dimensional coordinates of each hand keypoint in a corresponding first coordinate system, where the first coordinate system is a three-dimensional coordinate system determined by adjacent keypoints of the hand keypoint away from the fingertip; input the three-dimensional coordinates into a preset forward kinematics model to obtain the pose transformation set of the hand keypoints corresponding to each frame image, where the forward kinematics model consists of multiple hand keypoints connected by rigid bodies, and based on the kinematic description of screws, in the forward kinematics model, the wrist corresponds to the hand joint... The poses of keypoints are described in the camera coordinate system. The poses of hand keypoints corresponding to the metacarpophalangeal joints, relative to the hand keypoints corresponding to the wrist, are represented using the screw method based on the exponential product formula and composition rules of rigid body motion. The pose transformation set is the set of poses corresponding to all hand keypoints. Based on the pose transformation set, the Jacobian matrix corresponding to each frame image and the predicted 3D coordinates corresponding to the hand keypoints are calculated, and the corresponding second keypoint position is calculated according to the projection of the predicted 3D coordinates in the image. The pose transformation set is nonlinearly optimized based on the first keypoint position, the second keypoint position, and the Jacobian matrix. Based on the forward kinematics model and the optimized pose transformation set, preset hand keypoint parameters are output. By detecting the position of hand keypoints in the image, and simultaneously obtaining the pose transformation set and calculating the position of hand keypoints based on the constructed forward kinematics model, the positions obtained by the two methods are used to optimize the pose transformation set. Based on the pose transformation set, the relevant parameters of the hand in each frame image are finally confirmed. The hand parameters obtained through the forward kinematics model ensure the consistency of the hand shape processing before and after, thereby improving the accuracy of the processing results. Attached Figure Description
[0026] Figure 1 A flowchart of a gesture recognition method based on spinor provided in an embodiment of the present invention;
[0027] Figure 2 A schematic diagram of the key points and degrees of freedom of the hand;
[0028] Figure 3 This is a schematic diagram of spinor motion;
[0029] Figure 4 A schematic diagram illustrating the recognition result of hand key points in a frame of an image provided in an embodiment of the present invention;
[0030] Figure 5 This is a schematic diagram of the structure of a gesture recognition device based on spinor provided in an embodiment of the present invention;
[0031] Figure 6This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0032] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and not for limiting the invention. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention and not the entire structure.
[0033] It should be noted that, due to space limitations, this application specification does not exhaustively list all possible implementation methods. Those skilled in the art should be able to conceive after reading this application specification that, as long as the technical features do not contradict each other, any combination of technical features can constitute an optional implementation method.
[0034] The embodiments are described in detail below.
[0035] Figure 1 This invention provides a flowchart of a screw-based gesture recognition method for use in electronic devices. As shown in the figure, the screw-based gesture recognition method includes:
[0036] Step S110: Acquire consecutive frame images captured by the camera, detect hand key points in the consecutive frame images, and obtain the first key point position corresponding to each hand key point in each frame image.
[0037] The gesture interaction scenario applied in this solution relies on continuous frame images captured by a camera for gesture recognition. Specifically, this can be a monocular camera or a multi-camera system (e.g., a binocular camera composed of two monocular cameras, or a combination of more monocular cameras). Each camera captures one continuous frame image, and gesture interaction recognition is achieved by analyzing the image information of the bare hand from these continuous frame images.
[0038] Gesture recognition is ultimately achieved by judging the relative positional relationships of key hand points. Generally, key hand points are defined based on the skeletal structure of the hand. The human hand skeleton is mainly composed of carpal bones, metacarpal bones, and phalanges. Phalanges are composed of proximal phalanges, middle phalanges, and distal phalanges. Between the carpal bones, metacarpal bones, and proximal phalanges, and between the middle phalanges and distal phalanges, the following joints are formed: wrist joint (MCP), proximal interphalangeal joint (PIP), and distal interphalangeal joint (DIP). Based on the physiological characteristics of the human hand, different joints have different degrees of freedom (DoF). For example, the MCP joint is generally considered to have 2DoF (flexion / extension and adduction / abduction), while the PIP and DIP joints only have 1DoF (flexion / extension). Key hand points essentially correspond to the location of the joints in the hand skeleton and the location of the fingertips, such as... Figure 2 As shown, key points of the hand are represented by black dots. A normal and complete hand has 21 key points, and different joint types are also marked with DoF.
[0039] Specifically, in this solution, the key points of the hand in the image are first obtained based on content detection. For example, a neural network is pre-trained, and each frame of the image is input into the neural network to obtain the observation values of the key points of the hand. Alternatively, the observation values of the key points of the hand can be obtained through computer vision methods, such as feature point extraction and feature point matching. Key point recognition based on neural networks or computer vision methods has been implemented in the existing technology and will not be described in detail here.
[0040] Step S120: Calculate the three-dimensional coordinates of each of the hand key points in the corresponding first coordinate system, wherein the first coordinate system is a three-dimensional coordinate system determined by the adjacent key points of the hand key points away from the fingertips.
[0041] The calculation process for hand keypoints based on the forward kinematics model requires extracting the 3D coordinates of each keypoint from each frame of the image. However, not all keypoints' 3D coordinates are based on the same 3D coordinate system; instead, each hand keypoint corresponds to a separate first coordinate system. In terms of the connection relationships between hand keypoints, each hand keypoint has one or two adjacent hand keypoints, one of which is farther from the fingertip. The first coordinate system corresponding to each hand keypoint is a 3D coordinate system defined with the hand keypoint farther from the fingertip as its origin. In the specific calculation process, based on multiple preceding frames of images, the structure from motion (SFM) method can be used to obtain the 3D coordinates of each hand keypoint.
[0042] Step S130: Input the three-dimensional coordinates into a preset forward kinematics model to obtain the pose transformation set of each frame image corresponding to the hand key points. The forward kinematics model consists of multiple hand key points connected by rigid bodies and is described by screw kinematics. In the forward kinematics model, the pose of the hand key point corresponding to the wrist is described in the camera coordinate system. The pose of the hand key point corresponding to the metacarpophalangeal joint is represented by the screw method based on the exponential product formula and composition rules of rigid body motion, relative to the hand key point corresponding to the wrist. The pose transformation set is the set of poses corresponding to all hand key points.
[0043] In this scheme, to achieve gesture recognition, the ultimate goal is to obtain a description of the hand's motion process. Based on the analysis of the hand's structure, the hand's motion process is modeled using the exponential product formula method based on screw theory. According to screw theory, the motion of any rigid body can be achieved by a rotation about an axis plus a translation parallel to that axis. The screw method describes the pose of the robot's end effector through the pose of each joint axis. Therefore, at each joint, it is only necessary to find the helical axis of the joint kinematic pair, and the motion of the rigid body can be decomposed into rotation about this axis and translation parallel to this axis.
[0044] Please refer to Figure 3 To illustrate spinor motion using the motion of a three-dimensional point: a three-dimensional point If a rigid body rotates about the ω-axis with a non-zero θ and moves a distance d in a direction parallel to ω, then the rigid body motion g in three-dimensional space can be described as the following kinematic spinor. Index:
[0045]
[0046] v=-ω×q+hω
[0047]
[0048] g is a 4×4 matrix. It is a 4×4 matrix. A unit vector representing the direction of rotation (axis of rotation), Let h be any point on the axis of rotation, and h be the pitch, i.e., h = d / θ; for a revolute joint, In the case of h = 0, ξ ∧ symbols in ∧ This represents transforming a vector into an antisymmetric matrix, abbreviated as: Right now:
[0049]
[0050] Here, The exponent of the spinor can be understood as describing the transformation of a rigid body from its initial position to its final pose, representing the transformation of the rigid body within the same reference frame.
[0051] For a three-dimensional point In other words, the following formula does not represent the transformation of a point between different coordinate systems, but rather the coordinates p(θ) of a point in the same coordinate system after a rigid body rotation from its initial position p(0):
[0052]
[0053] For a pose g, if the T coordinate system is fixed to the rigid body, the initial pose of the rigid body relative to the fixed coordinate system S is denoted as g. ST (0), then after spinor motion, the pose (or configuration) of the T frame relative to the S frame is:
[0054]
[0055] The meaning of this transformation is: g ST (0) represents transforming the coordinates of a point relative to the T-frame to the coordinates relative to the S-frame, with the exponent e of the spinor. θξ∧ This means transforming the point to its final position (still within the S-frame).
[0056] The positive kinematics model in this scheme is established based on the world coordinate system. Therefore, it is necessary to obtain the camera model parameters, initial gesture, and hand configuration. The hand configuration includes the rigid body connection length between hand key points and the position of the hand key points in the camera coordinate system. The positive kinematics model is obtained based on the camera model parameters, initial gesture, and rigid body connection length, using the exponential product formula of the coordinate transformation of the hand key points and the synthesis rules.
[0057] Camera model parameters include camera intrinsic and extrinsic parameters. This patent applies to various cameras besides depth cameras. If one or more monocular cameras are used, the intrinsic parameter calibration of each monocular camera needs to be completed in advance. If there are multiple monocular cameras, the corresponding extrinsic parameter calibration also needs to be completed in advance. That is, for multi-camera systems, the intrinsic and extrinsic parameters have already been calibrated in advance. Camera calibration can use various types of calibration boards, such as the AprilGrid calibration board or the checkerboard calibration board. Obtaining camera model parameters is actually obtaining the intrinsic parameters, or intrinsic and extrinsic parameters, obtained during pre-calibration. The extrinsic parameters are the transformation matrices between each pair of cameras. In this solution, for multi-camera systems, one camera needs to be determined as the reference camera, and its pose in the world coordinate system needs to be obtained accordingly. The initial gesture is the gesture during initialization, usually corresponding to the case where the rotation angles of the hand keypoints are all 0, or the state corresponding to the hand being naturally extended can be set as the reference gesture; the reference gesture can also be flexibly set according to the actual situation, and then the pose transformation of the joints is generated based on this. The initial gesture is described by the corresponding pose transformation set.
[0058] The rigid body connection length between hand keypoints is determined based on the basic dimensions of the hand, such as taking the average length between joints of multiple hands; it can also be set according to the actual dimensions of the current user. The rigid body connection length between hand keypoints is the bone length between hand keypoints, which is part of the hand configuration. The hand configuration can also include the pose of all joints (i.e., hand keypoints other than the fingertips) in the camera coordinate system. If the wrist pose relative to the camera coordinate system is not set initially, the pose can be calculated based on multiple frames of images; multi-view geometry or motion reconstruction methods are used for multi-frame images, combined with the camera model, for calculation.
[0059] When establishing the forward kinematics model, to make it more consistent with the general laws of hand movement, the degrees of freedom (DOF) of each key hand point are further obtained; and motion state constraints related to these DDFs are set in the forward kinematics model. The DDFs of the key hand points are actually the DDFs of the joints corresponding to those key points. Due to the limitations of the human hand's bones and muscles, different joints have different DDFs. Currently, it is generally believed that the human wrist has 6 DDFs, and the human finger joints have no more than 2 DDFs, namely abduction / adduction and flexion / extension. Figure 2The diagram provided is simply a schematic representation of key hand points defined as 26 DoF, which can be used as an example for illustration. Specifically: a) Wrist joint: 6 DoF, i.e., 3 DoF rotation + 3 DoF translation. b) Index, middle, ring, and little fingers: These four fingers have the same defined degrees of freedom, each with 4 DoF: 2 DoF at the MCP joint (abduction / adduction, flexion / extension), 1 DoF at the PIP joint (flexion / extension), and 1 DoF at the DIP joint (flexion / extension). c) Thumb: 1 DoF at the IP joint, 1 DoF at the MCP joint, and 2 DoF at the CMC joint, for a total of 4 DoF for the thumb.
[0060] It should be noted that there is no consensus in the industry regarding the number of DoF (DoF) points on the thumb; the main controversy centers on the two joints of the thumb: MCP (Mean Collateral Pressure) and CMC (Central Collateral Pressure). Figure 2 The 4DoF used in this study is just an exemplary choice; in practice, CMC joints can also have 3DoF and MCP joints can have 2DoF. Furthermore, 26DoF is only an optional solution; when designing a specific forward kinematic model, the maximum DoF can be designed for each joint. Figure 2 The design incorporates a total of 26 DoFs (DoFs). These 26 DoF vectors are called state vectors, which are the quantities to be optimized / estimated. When the values of the state vectors are estimated, the positions of the key points between the fingers can be determined. Of course, the dimensions of the state vectors will change depending on the DoF settings for the hand key points, but the basic calculation process remains the same.
[0061] The initial value of the pose transformation set of the currently processed image is the pose transformation set corresponding to the previous frame image; or, it is the position transformation set initialized based on the currently processed image.
[0062] The pose transformation set includes the rotation and rotation of each hand key point;
[0063] When initializing the pose transformation set, the motion spinor corresponds to the spinor motion of the joint when all other joints except the joint are fixed at a position with a rotation angle of 0. The position of the key hand point corresponding to the wrist joint is calculated by combining the camera model parameters of the camera with multi-view geometry and motion recovery structure.
[0064] In the positive kinematic model of this scheme, the carpal bones, metacarpal bones, proximal phalanges, middle phalanges and distal phalanges of the hand are all considered as rigid links. The carpometacarpal joint (i.e. wrist joint) connecting the carpal bones and metacarpal bones is equivalent to a ball joint, and the other joints are equivalent to rotational joints. Thus, each finger of a person is connected by a set of rigid links through rotational joints.
[0065] When using spinor theory to model rigid body motion, the wrist pose relative to the camera coordinate system can be obtained using the method described above.c g r If the metacarpophalangeal joints only rotate relative to the wrist, then the positive kinematics of the pose transformation set corresponding to the metacarpophalangeal joints, taking a single finger as an example, occurs when the finger joints MCP, PIP, and DIP rotate at angles θ. M θ P θ D During rotation, where θ M Including θ M1 and θ M2 The two parts, expressed using the spinor method and based on the exponential product formula of rigid body motion, are as follows:
[0066]
[0067] r g dip (0) indicates the position of the DIP relative to the wrist during the initial gesture. r g dip (θ M ,θ P ,θ D The expression represents the pose of the DIP relative to the wrist after the movement of each of the aforementioned joints. Furthermore, to obtain the pose transformation set relative to the camera coordinate system, simply left-multiply... c g r If you need to obtain the pose relative to the world coordinate system, simply multiply the camera extrinsic parameters by left.
[0068] Each pose transformation matrix g in the above pose transformation set is determined by the motion screw and the rotation angle θ, where θ is the quantity to be estimated. The motion screw is calculated as follows:
[0069] Each key point of the hand outside the wrist is equivalent to Figure 3 Point q on the central axis, that is, the position of point q is the position of the corresponding hand key point, representing the calculation of the unit vector ω of the rotation axis. If the position of the hand key point has been configured in the hand configuration in advance, the rotation axis of the joint angle is calculated according to the geometric method and the linear algebra method; if it has not been configured, multiple frames of images are acquired and then calculated by combining multi-view geometric methods.
[0070] Step S140: Based on the pose transformation set, calculate the Jacobian matrix corresponding to each frame image and the predicted 3D coordinates corresponding to the hand key points, and calculate the corresponding second key point position according to the projection of the predicted 3D coordinates in the image.
[0071] Based on the construction of a forward kinematic model of the hand at the start of movement, the positions of key hand points are further calculated. Mathematically, during hand movement, the rotation angle and translation are perturbed from the initial gesture, and the values of the pose transformation sets corresponding to each joint change accordingly. Each joint transforms from the initial pose to the target pose along with the pose transformation set.
[0072] The extremities of the bones (i.e., the fingertips) move with the joints, meaning the extremities transform from their initial position to the target position according to the pose transformation set. Therefore, the 3D positions of key points in the hand can be calculated. Taking the fingertip as an example, the 3D positions of key points relative to the camera coordinate system after the pose transformation are calculated from their initial positions:
[0073] c p tip (θ M ,θ P ,θ D )= c g r · r g dip (θ M ,θ P ,θ D )· dip p tip (0)
[0074] in dip p tip (0) represents the pose of the fingertip in the DIP coordinate system at the initial gesture. This yields the predicted 3D coordinates. Based on the camera model parameters, the predicted 3D coordinates are projected onto the corresponding image to obtain the second keypoint pose. Simultaneously, the corresponding Jacobian matrix is calculated based on the current pose transformation set.
[0075] Step S150: Perform nonlinear optimization on the pose transformation set based on the first key point position, the second key point position, and the Jacobian matrix.
[0076] During the calculation, the positions of the first key point (i.e., the hand key points directly identified from the image) and the second key point (i.e., the calculated hand key points) cannot completely coincide, meaning the difference between the two is not zero. In this case, it is necessary to optimize this calculation error to obtain the optimal pose parameters and the coordinates of the three-dimensional points. Specifically, this error is processed through nonlinear optimization.
[0077] In the specific implementation process, step S150 can be implemented through steps S151-S154:
[0078] Step S151: Calculate the reprojection error using the positions of the first key point and the second key point.
[0079] The reprojection error is the L2 norm of the first keypoint position and the second keypoint position.
[0080] Step S152: Calculate the temporal error using the pose transformation set of the currently processed image and the previous frame image.
[0081] The time-domain error is transmitted via e temporal =||θ t -θ t-1 || 2 Calculate, where e temporal Representing the time-domain error, θ t θ represents the rotation and translation amounts corresponding to the pose transformation set of the currently processed image; t-1 The transformation matrix represents the rotation angle and translation amount corresponding to the pose transformation set of the previous frame image. In the specific calculation process, the transformation matrix is calculated based on the screw, rotation angle and initial gesture. The Euler angle and translation amount can be calculated from the transformation matrix.
[0082] Step S153: Calculate the sum of the reprojection error and the time-domain error to obtain the error vector.
[0083] Step S154: Perform nonlinear optimization on the pose transformation set based on the error vector and the Jacobian matrix.
[0084] The error vector and Jacobian matrix are input into a preset nonlinear optimization model, which outputs an optimized pose transformation set. In the specific processing, the pose transformation set for each frame is calculated independently. This calculation method may lead to jitter in the pose calculation results. Therefore, based on the optimized pose transformation sets corresponding to multiple consecutive frames, and using the currently processed image as a reference, a preset number of pose transformation sets for preceding and following frames are obtained. Based on the obtained pose transformation sets, temporal smoothing processing is performed on the pose transformation set of the currently processed image. Specifically, taking the currently processed image as the center, a preset number of pose transformation sets for preceding and following frames are taken, and then an adaptive low-pass filtering algorithm is used to process the rotation angle and translation to alleviate jitter and achieve temporal smoothing processing.
[0085] Specifically, in this scheme, the calculation formulas for the error vector and the Jacobian matrix are nonlinear. Therefore, the corresponding optimization problem can be transformed into a nonlinear least squares problem, which can be solved using a nonlinear optimization algorithm. The commonly used definition of a nonlinear least squares problem is as follows:
[0086] Given a vector function f (containing n unknown parameters and m observations, where m ≥ n), our goal is to find the minimum value of ‖f(x)‖, which is equivalent to finding the minimum value of ‖f(x)‖.
[0087]
[0088] in The vector function f consists of the computational cost and the error between the observations at each hand keypoint, where x is the unknown vector to be optimized, representing the pose of the hand joints. Essentially, x can be understood as the state vector in the robot's state estimation problem. Based on this, commonly used methods for solving nonlinear least squares problems can be employed, such as gradient descent, Newton's method, Gauss-Newton method, and Levenberg-Marquardt algorithm.
[0089] Step S160: Output preset hand key point parameters based on the positive kinematics model and the optimized pose transformation set.
[0090] Based on a defined pose transformation set, and using a forward kinematics model, various preset hand keypoint parameters can be output according to actual needs. These preset hand keypoint parameters may include the 3D coordinates of the keypoints, the rotation angles and translations of the hand keypoints, and one or more parameters from the pose transformation set of the hand keypoints. Calculating these parameters using a forward kinematics model based on a defined pose transformation set is a standard computational process in forward kinematics.
[0091] Please refer to Figure 4 For the currently processed image, there are Figure 4 This is a schematic diagram of the hand keypoints detected in the currently processed image, with 21 keypoints numbered 0-20. The position of each keypoint is the first keypoint position. Assuming the current image is numbered 't' in the entire series of frames, the corresponding first keypoint position is represented as... i represents the hand keypoint number, t represents the t-th frame of the currently calculated image, and (u,v) represents the position coordinates of the hand keypoint in the image. During initialization, as described above, the camera model parameters, initial gesture, and hand configuration are initialized. The hand configuration includes the rigid body connection length between hand keypoints and the position of the hand keypoints in the camera coordinate system. The camera model parameters include the intrinsic parameters of each monocular camera, used to convert the position (x,y,z) in 3D space into (u,v) in the 2D image, and the poses of the left and right cameras relative to the world coordinate system, denoted as... w T cL , w T cR If the hand configuration does not include the wrist pose relative to the camera coordinate system, multiple frames of images are acquired. Based on these multiple frames, a multi-view set or motion recovery structure is used, combined with the camera model and the existing hand configuration, to calculate the wrist pose.
[0092] Specifically, the forward kinematics model of the hand at the start of movement assumes that during hand activity, the rotation angle and translation are perturbed based on the initial gesture, and the values of the pose transformation set change accordingly. The key points of the hand corresponding to each joint transform from the initial pose to the target pose along with the pose transformation set. The skeletal ends (i.e., fingertips) move with the joint transformations, thus allowing the calculation of the three-dimensional coordinates of the key points of the hand. The positions of each key point of the hand (including fingertips and joints) after pose transformation from the initial position are shown below, taking the fingertip as an example:
[0093] c p tip (θ M ,θ P ,θ D )= c g r · r g dip (θ M ,θ P ,θ D )· dip p tip (0)
[0094] in, dip p tip (0) indicates the position of the fingertip in the DIP coordinate system at the initial gesture.
[0095] Calculate the 2D position of the fingertip on the image plane based on the camera projection transformation (denoted by Proj()):
[0096] uv p tip =Proj( c p tip )
[0097] If there are multiple cameras, the projection transformation of the corresponding camera can be used. The positions of other hand key points are calculated in the same way.
[0098] Calculating the Jacobian matrix of the pose transformation set is an existing implementation. Specifically, in this scheme, for each finger, based on the aforementioned calculation of the fingertip after pose transformation, the Jacobian matrix needs to be further calculated in the following way (i.e., taking the fingertip as an example):
[0099] Taking the partial derivative of the perturbation generated at each joint, and using the perturbation of the DIP joint as an example: when the DIP joint generates an angle change dθ... D At that time, the pose changes, and the three-dimensional coordinates of the key points become:
[0100]
[0101] Where, (θ Mθ P θ D Abbreviated as θ one Its perturbation is abbreviated as dθ one Based on the above equation, the disturbance quantity dθ D Find the partial derivative, using The sign for partial derivatives is denoted as:
[0102]
[0103] Similarly, it can be calculated
[0104] Since the projection is onto a 2D image, the Jacobian of the projection transformation also needs to be calculated:
[0105] Based on the obtained positions of the key hand points and the Jacobian matrix, further error processing is performed. The error function is defined as the sum of the reprojection error and the time-domain error, E = e 2D +e temporal Reprojection error: e 2D This represents the updated error vector obtained by comparing the calculated updated hand keypoint positions (i.e., the second keypoint positions) with the detected hand keypoint positions (i.e., the first keypoint positions). The 2D coordinates of the i-th keypoint in the image space of the j-th camera are obtained through detection; uv p i The calculation is based on camera projection transformation, as mentioned above, where Proj j () represents a function that projects a point in 3D space onto the image space of the j-th camera. Temporal error: e temporal =||θ t -θ t-1 || 2 Here, the state variable θ includes the 6 degrees of freedom of the wrist joint and the 20 rotation angles of the other finger joints (excluding the fingertips), making it a vector with a total of 26 dimensions.
[0106] Based on the Jacobian matrix and error vector calculated above, the state vector of the hand key points is optimized: specifically, the error and Jacobian function are used as inputs into a general optimization algorithm (such as the nonlinear optimization algorithm Levenberg-Marquardt algorithm or Gauss-Newton algorithm), and the iteration termination condition is set to obtain the optimized and updated pose transformation set.
[0107] The optimized pose transformation set is obtained by independently processing single-frame images, which may lead to jitter in the pose description between multiple images. Therefore, this scheme sets a temporal window of length k frames. After optimization calculation of the t-th frame image, a total of k frames before and after it are taken (e.g., the optimized pose transformation set of frame t-1 and the optimized pose transformation set of frame t). Then, an adaptive low-pass filtering algorithm is used to process the rotation and translation amounts to obtain a smoothed pose transformation set to alleviate jitter. Here, a first-order low-pass filtering is used as an example, and the following formula is applied to each component of the 26 DoFs:
[0108]
[0109] The value of α ranges from 0.2 to 0.9, and is determined based on the proportion of pixels occupied by the key points of the hand in the entire 2D image.
[0110] Based on the smoothing process, various numerical values can be output, including but not limited to (the following example only uses the middle finger):
[0111] Coordinates of the fingertip: c p tip (θ one +dθ one ).
[0112] The same applies to the distal interphalangeal joint, proximal interphalangeal joint, metacarpophalangeal joint, and wrist joint.
[0113] The location of key hand points on the 2D image: simply perform a projection transformation on the 3D points obtained above.
[0114] When calculating the optimized estimate for each corner of the pre-set set of joints, the updated corner method can be obtained.
[0115] The updated pose transformation set represents the hand's transformation from the initial pose to the updated pose.
[0116] This solution eliminates the need for model training based on extensive sample data collection. Accurate recognition results can be obtained using a hand motion model constructed based on forward kinematics, improving processing efficiency and accuracy. Building upon computer vision-based keypoint detection, a kinematic model in 3D space is added to provide accurate motion constraints, outputting scale information and various 3D spatial information, and in particular, preventing the generation of impossible gestures.
[0117] The above method acquires consecutive frame images captured by a camera, detects hand keypoints in the consecutive frame images to obtain the position of the first keypoint corresponding to each hand keypoint in each frame image; calculates the three-dimensional coordinates of each hand keypoint in the corresponding first coordinate system, where the first coordinate system is a three-dimensional coordinate system determined by the adjacent keypoints of the hand keypoint away from the fingertip; inputs the three-dimensional coordinates into a preset forward kinematics model to obtain the pose transformation set of the hand keypoints corresponding to each frame image, where the forward kinematics model consists of multiple hand keypoints connected by rigid bodies, and is based on the kinematic description of screws. In the forward kinematics model, the pose of the hand keypoint corresponding to the wrist is transformed in the camera coordinate system. The description describes the pose of key hand points corresponding to the metacarpophalangeal joints, relative to the key hand points corresponding to the wrist. This pose is represented using the screw method based on the exponential product formula and composition rules of rigid body motion. The pose transformation set is the set of poses corresponding to all key hand points. Based on the pose transformation set, the Jacobian matrix corresponding to each frame image and the predicted 3D coordinates corresponding to the key hand points are calculated. The position of the corresponding second key point is calculated based on the projection of the predicted 3D coordinates onto the image. The pose transformation set is then nonlinearly optimized based on the first key point position, the second key point position, and the Jacobian matrix. Preset key hand point parameters are output based on the forward kinematics model and the optimized pose transformation set. By detecting the position of key hand points in images and simultaneously obtaining the pose transformation set and calculating the position of key hand points based on the constructed forward kinematics model, the positions obtained in both methods are used to optimize the pose transformation set. Based on the pose transformation set, the relevant parameters of the hand in each frame image are finally confirmed. The acquisition of hand parameters through the forward kinematics model ensures the consistency of hand shape processing before and after, thereby improving the accuracy of the processing results.
[0118] Figure 5 This is a schematic diagram of a screw-based gesture recognition device provided in an embodiment of the present invention. This screw-based gesture recognition device is used in electronic devices. (Refer to...) Figure 5 The spinor-based gesture recognition device includes a key point detection unit 210, a coordinate calculation unit 220, a model calculation unit 230, a position calculation unit 240, a nonlinear optimization unit 250, and a parameter output unit 260.
[0119] The keypoint detection unit 210 is used to acquire continuous frame images captured by the camera and detect hand keypoints in the continuous frame images to obtain the position of the first keypoint corresponding to each hand keypoint in each frame image; the coordinate calculation unit 220 is used to calculate the three-dimensional coordinates of each hand keypoint in the corresponding first coordinate system, where the first coordinate system is a three-dimensional coordinate system determined by the adjacent keypoints of the hand keypoint away from the fingertip; the model calculation unit 230 is used to input the three-dimensional coordinates into a preset positive kinematics model to obtain the pose transformation set of the hand keypoints corresponding to each frame image; the positive kinematics model consists of multiple hand keypoints connected by rigid bodies, and is based on the kinematic description of screws. In the positive kinematics model, the pose of the hand keypoints corresponding to the wrist is calculated in the camera coordinate system. The description defines the pose of the hand key points corresponding to the metacarpophalangeal joints, relative to the hand key points corresponding to the wrist, using the screw method based on the exponential product formula and composition rules of rigid body transformation. The pose transformation set is the set of poses corresponding to all hand key points. The position calculation unit 240 is used to calculate the Jacobian matrix corresponding to each frame image and the predicted three-dimensional coordinates corresponding to the hand key points based on the pose transformation set, and calculate the corresponding second key point position according to the projection of the predicted three-dimensional coordinates in the image. The nonlinear optimization unit 250 is used to perform nonlinear optimization on the pose transformation set based on the first key point position, the second key point position, and the Jacobian matrix. The parameter output unit 260 is used to output preset hand key point parameters according to the positive kinematics model and the optimized pose transformation set.
[0120] Based on the above embodiments, the device further includes:
[0121] The temporal smoothing unit is used to obtain a set of pose transformations of a preset number of images before and after the currently processed image as a reference, and to perform temporal smoothing on the pose transformation set of the currently processed image based on the obtained pose transformation set.
[0122] Based on the above embodiments, the time-domain smoothing process is performed using low-pass filtering.
[0123] Based on the above embodiments, the device further includes:
[0124] The parameter acquisition unit is used to acquire camera model parameters, initial gestures, and hand configurations. The hand configurations include the rigid body connection lengths between hand keypoints and the positions of the hand keypoints in the camera coordinate system.
[0125] The positive kinematics model is obtained based on the camera model parameters, initial gesture, and rigid body connection length, using the exponential product formula of the rigid body transformation of hand key points and the synthesis rules.
[0126] Based on the above embodiments, the device further includes:
[0127] The degree-of-freedom acquisition unit is used to acquire the degree of freedom of each key hand point;
[0128] The forward kinematics model includes motion state constraints related to the degrees of freedom.
[0129] Based on the above embodiments, the initial value of the pose transformation set of the currently processed image is the pose transformation set corresponding to the previous frame image; or, it is the position transformation set initialized based on the currently processed image.
[0130] Based on the above embodiments, the pose transformation set includes the rotation and rotation angle of each hand key point;
[0131] When initializing the pose transformation set, the motion spinor corresponds to the spinor motion of the joint when all other joints except the joint are fixed at a position with a rotation angle of 0. The pose of the key hand points corresponding to the wrist joint is calculated by combining the camera model parameters of the camera with multi-view geometry and motion recovery structure.
[0132] Based on the above embodiments, the nonlinear optimization unit includes:
[0133] The reprojection error calculation module is used to calculate the reprojection error based on the positions of the first key point and the second key point.
[0134] The temporal error calculation module is used to calculate the temporal error using the pose transformation set between the currently processed image and the previous frame image.
[0135] The error vector calculation module is used to calculate the sum of the reprojection error and the time-domain error to obtain the error vector;
[0136] An optimization processing module is used to perform nonlinear optimization on the pose transformation set based on the error vector and the Jacobian matrix.
[0137] The nonlinear optimization unit, wherein the time-domain error is transmitted through e temporal =||θ t -θ t-1 || 2 Calculate, where e temporal Representing the time-domain error, θ t θ represents the rotation and translation amounts corresponding to the pose transformation set of the currently processed image; t-1 This represents the rotation and translation amounts corresponding to the pose transformation set of the previous frame image.
[0138] The nonlinear optimization unit, specifically the optimization processing module, is used to input the error vector and Jacobian matrix into a preset nonlinear optimization model and output the optimized pose transformation set.
[0139] The nonlinear optimization unit has preset hand key point parameters including the 3D coordinates of the key points, the rotation angle and translation of the hand key points, and one or more of the pose transformation set of the hand key points.
[0140] The screw-based gesture recognition device provided in this embodiment of the invention is included in the electronic device of the device and can be used to execute any of the screw-based gesture recognition methods provided in the above embodiments, and has corresponding functions and beneficial effects.
[0141] It is worth noting that in the above embodiments of the gesture recognition device based on spin, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0142] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Figure 6 As shown, the electronic device includes a processor 310, a memory 320, an input device 330, an output device 340, and a communication device 350; the number of processors 310 in the electronic device can be one or more. Figure 6 Taking a processor 310 as an example; the processor 310, memory 320, input device 330, output device 340, and communication device 350 in the electronic device can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.
[0143] The memory 320, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the screw-based gesture recognition method in this embodiment of the invention (e.g., the key point detection unit 210, coordinate calculation unit 220, model calculation unit 230, position calculation unit 240, nonlinear optimization unit 250, and parameter output unit 260 in the screw-based gesture recognition device). The processor 310 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 320, thereby realizing the aforementioned screw-based gesture recognition method.
[0144] The memory 320 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 320 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 320 may further include memory remotely located relative to the processor 310, which can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0145] Input device 330 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the electronic device. Output device 340 may include display devices such as a display screen.
[0146] The aforementioned electronic device includes a spinor-based gesture recognition device, which can be used to execute any spinor-based gesture recognition method and has corresponding functions and beneficial effects.
[0147] This invention also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform relevant operations in the spinor-based gesture recognition method provided in any embodiment of this application, and have corresponding functions and beneficial effects.
[0148] Those skilled in the art will understand that embodiments of this application may be provided as methods, systems, or computer program products.
[0149] Therefore, this application may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce implementations of the flowchart... Figure 1 One or more processes and / or boxes Figure 1The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0150] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0151] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0152] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0153] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A gesture recognition method based on spinor, characterized in that, include: Acquire consecutive frame images captured by the camera, detect key hand points in the consecutive frame images, and obtain the position of the first key point corresponding to each key hand point in each frame image; Calculate the three-dimensional coordinates of each of the hand key points in the corresponding first coordinate system, wherein the first coordinate system is a three-dimensional coordinate system determined by the adjacent key points of the hand key points away from the fingertips; The three-dimensional coordinates are input into a preset forward kinematics model to obtain the pose transformation set of each frame image corresponding to the hand key points. The forward kinematics model consists of multiple hand key points connected by rigid bodies and is described by screw kinematics. In the forward kinematics model, the pose of the hand key point corresponding to the wrist is described in the camera coordinate system, and the pose of the hand key point corresponding to the metacarpophalangeal joint is represented by the screw method based on the exponential product formula and composition rules of rigid body motion, relative to the hand key point corresponding to the wrist. The pose transformation set is the set of poses corresponding to all hand key points. Based on the pose transformation set, calculate the Jacobian matrix and the predicted 3D coordinates of the hand key points corresponding to each frame image, and calculate the position of the corresponding second key point according to the projection of the predicted 3D coordinates in the image. Based on the positions of the first key point, the second key point, and the Jacobian matrix, the pose transformation set is nonlinearly optimized. Based on the positive kinematics model and the optimized pose transformation set, preset hand key point parameters are output.
2. The gesture recognition method according to claim 1, characterized in that, After performing nonlinear optimization on the pose transformation set based on the first key point position, the second key point position, and the Jacobian matrix, the method further includes: Using the currently processed image as a reference, obtain a set of pose transformations for a preset number of images before and after the current image. Based on the obtained set of pose transformations, perform temporal smoothing on the pose transformation set of the currently processed image.
3. The gesture recognition method according to claim 2, characterized in that, The time-domain smoothing process is performed using a low-pass filter.
4. The gesture recognition method according to claim 1, characterized in that, Before presetting the positive kinematics model, it also includes: Obtain camera model parameters, initial gestures, and hand configuration, wherein the hand configuration includes the rigid body connection length between hand key points and the position of the hand key points in the camera coordinate system; The positive kinematics model is obtained based on the camera model parameters, initial gesture, and rigid body connection length, using the exponential product formula of the rigid body transformation of hand key points and the synthesis rules.
5. The gesture recognition method according to claim 4, characterized in that, Before presetting the positive kinematics model, it also includes: Obtain the degrees of freedom for each key hand point; The forward kinematics model includes motion state constraints related to the degrees of freedom.
6. The gesture recognition method according to claim 1, characterized in that, The initial value of the pose transformation set of the currently processed image is the pose transformation set corresponding to the previous frame image; or, it is the pose transformation set initialized based on the currently processed image.
7. The gesture recognition method according to claim 6, characterized in that, The pose transformation set includes the rotation and rotation of the joints corresponding to each key hand point; When initializing the pose transformation set, the motion spinor corresponds to the spinor motion of the joint when all other joints except the joint are fixed at a position with a rotation angle of 0. The pose of the key hand points corresponding to the wrist joint is calculated by combining the camera model parameters of the camera with multi-view geometry and motion recovery structure.
8. The gesture recognition method according to claim 1, characterized in that, The step of performing nonlinear optimization on the pose transformation set based on the first key point position, the second key point position, and the Jacobian matrix includes: The reprojection error is calculated using the positions of the first and second key points. The temporal error is calculated using the pose transformation set of the currently processed image and the previous frame image; The sum of the reprojection error and the time-domain error is calculated to obtain the error vector; The pose transformation set is nonlinearly optimized based on the error vector and the Jacobian matrix.
9. The gesture recognition method according to claim 8, characterized in that, The time-domain error is passed through Calculation, where Indicates time-domain error, This represents the rotation and translation amounts corresponding to the pose transformation set of the currently processed image; This represents the rotation and translation amounts corresponding to the pose transformation set of the previous frame image; It represents the square of the L2 norm of a vector.
10. The gesture recognition method according to claim 8, characterized in that, The nonlinear optimization of the pose transformation set based on the error vector and the Jacobian matrix specifically involves: The error vector and Jacobian matrix are input into a preset nonlinear optimization model, and the optimized pose transformation set is output.
11. The gesture recognition method according to claim 1, characterized in that, The preset hand key point parameters include the 3D coordinates of the hand key points, the rotation angle and translation of the hand key points, and one or more of the pose transformation set of the hand key points.
12. A gesture recognition device based on spinor, characterized in that, include: The key point detection unit is used to acquire continuous frame images captured by the camera, detect hand key points in the continuous frame images, and obtain the first key point position corresponding to each hand key point in each frame image. The coordinate calculation unit is used to calculate the three-dimensional coordinates of each of the hand key points in the corresponding first coordinate system, wherein the first coordinate system is a three-dimensional coordinate system determined by the adjacent key points of the hand key points away from the fingertips; The model calculation unit is used to input the three-dimensional coordinates into a preset forward kinematics model to obtain the pose transformation set of each frame image corresponding to the hand key points. The forward kinematics model consists of multiple hand key points connected by rigid bodies and is described by screw kinematics. In the forward kinematics model, the pose of the hand key point corresponding to the wrist is described in the camera coordinate system, and the pose of the hand key point corresponding to the metacarpophalangeal joint is represented by the screw method based on the exponential product formula and composition rules of rigid body motion, relative to the hand key point corresponding to the wrist. The pose transformation set is the set of poses corresponding to all hand key points. The position calculation unit is used to calculate the Jacobian matrix and the predicted three-dimensional coordinates of the hand key points corresponding to each frame image based on the pose transformation set, and to calculate the position of the corresponding second key point according to the projection of the predicted three-dimensional coordinates in the image. A nonlinear optimization unit is used to perform nonlinear optimization on the pose transformation set based on the first key point position, the second key point position, and the Jacobian matrix. The parameter output unit is used to output preset hand key point parameters based on the positive kinematics model and the optimized pose transformation set.
13. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the electronic device implements the spinor-based gesture recognition method as described in any one of claims 1-11.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the spinor-based gesture recognition method as described in any one of claims 1-11.