Sample data generation method and apparatus, and electronic device

By constructing a hand kinematics model and generating sample data using user-defined standard gestures, the problem of difficulty in obtaining sample data in existing technologies is solved, and the prediction accuracy of deep learning models and the diversity of sample data are improved.

CN113705378BActive Publication Date: 2026-02-03GUANGZHOU HUYA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110919681.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-11
Publication Date
2026-02-03
Estimated Expiration
2041-08-11

AI Technical Summary

Technical Problem

Existing technologies face difficulties in obtaining sample data of hand key points and pose information when training deep learning models, resulting in limited sample data and affecting the model's prediction accuracy, especially the low accuracy in recognizing specific gestures.

Method used

By constructing a hand kinematic model, multiple sets of sample data are generated using user-defined standard gestures and hand posture information, including key hand point location information and posture information, reducing manual annotation steps and improving the diversity and accuracy of sample data.

Benefits of technology

It enables the rapid and convenient acquisition of large amounts of diverse sample data, improving the prediction accuracy of deep learning models, especially the recognition accuracy of standard gestures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113705378B_ABST
    Figure CN113705378B_ABST
Patent Text Reader

Abstract

Embodiments of the present specification provide a sample data generation method and device and electronic equipment. When constructing sample data for training a deep learning model, a hand kinematics model can be constructed, which is used to describe the influence of a hand posture and a hand bone length on the position of a hand key point, and then a plurality of sets of sample data are generated according to hand posture information of a standard gesture predefined by a user and the hand kinematics model. Through this method, a large amount of sample data can be quickly and conveniently obtained without manual labeling of hand key points by the user, the difficulty of obtaining sample data is reduced, and the sample data constructed can be more diverse and have higher accuracy, thereby improving the prediction accuracy of the trained deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to a method, apparatus and electronic device for generating sample data. Background Technology

[0002] Gesture recognition has wide applications in many fields; for example, it can control smart devices. Accurately recognizing user gestures is a prerequisite for gesture control. Currently, when recognizing user gestures, 2D or 3D positional information of hand key points can be obtained from RGB or depth images containing the hand. This 2D or 3D positional information is then input into a pre-trained deep learning model to predict the hand pose in the image. However, this method is cumbersome when training the deep learning model because it requires labeling hand key points and hand poses in the image, and the limited sample data can affect the prediction accuracy of the trained deep learning model. Summary of the Invention

[0003] Based on this, this specification provides a method, apparatus, and electronic device for generating sample data.

[0004] According to a first aspect of the embodiments of this specification, a method for generating sample data is provided, the sample data being used to train a deep learning model, the deep learning model being used to predict hand pose information based on hand key point location information, the method comprising:

[0005] Acquire hand gesture information for a standard gesture, which is preset by the user.

[0006] Based on the hand posture information of the standard gesture and the preset hand kinematics model, multiple sets of sample data are generated. Each set of sample data includes the hand key point position information and the hand posture information corresponding to the hand key point position information.

[0007] The hand kinematic model is used to describe the influence of hand posture and hand bone length on the location of key hand points.

[0008] According to a second aspect of the embodiments of this specification, a sample data generation apparatus is provided, the sample data being used to train a deep learning model, the deep learning model being used to predict hand pose information based on hand key point location information, the apparatus comprising:

[0009] The acquisition module is used to acquire hand posture information of standard gestures, which is preset by the user.

[0010] The sample data generation module is used to generate multiple sets of sample data based on the hand posture information of the standard gesture and the preset hand kinematic model. Each set of sample data includes hand key point position information and hand posture information corresponding to the hand key point position information.

[0011] The hand kinematic model is used to describe the influence of hand posture and hand bone length on the location of key hand points.

[0012] According to a third aspect of the embodiments of this specification, an electronic device is provided, the electronic device including a processor, a memory, and a computer program stored in the memory that is executable by the processor, wherein the processor executes the computer program to implement the method mentioned in the first aspect above.

[0013] By applying the scheme of the embodiments in this specification, when constructing sample data for training a deep learning model, a hand kinematics model can be built. This hand kinematics model describes the influence of hand posture and hand bone length on the location of key hand points. Then, based on the hand posture information of a user-defined standard gesture and the hand kinematics model, multiple sets of sample data are generated. The sample data includes key hand point location information and corresponding hand posture information. The key hand point location information is used as the sample input to the deep learning model, and the hand posture information is used as the label for the sample input. This method allows for the rapid and convenient acquisition of a large amount of sample data without requiring users to manually label hand key points, reducing the difficulty of obtaining sample data. Furthermore, the constructed sample data can be more diverse and accurate, thereby improving the prediction accuracy of the trained deep learning model.

[0014] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this specification and, together with the description, serve to explain the principles of this specification.

[0016] Figure 1 This is a flowchart of a sample data generation method according to one embodiment of this specification.

[0017] Figure 2 This is a schematic diagram of a hand model according to one embodiment of this specification.

[0018] Figure 3 This is a schematic diagram of a standard gesture according to one embodiment of this specification.

[0019] Figure 4(a) is a schematic diagram of an interactive interface for setting standard gestures by a user according to an embodiment of this specification.

[0020] Figure 4(b) is a schematic diagram of a specified gesture according to an embodiment of this specification.

[0021] Figure 5(a) and 5(b) This is a schematic diagram of a uniform sampling of the wrist joint angle on a spherical surface according to one embodiment of this specification.

[0022] Figure 6 This is a schematic diagram of a process for generating sample data according to one embodiment of this specification.

[0023] Figure 7 This is a logical structure block diagram of a sample generation device according to one embodiment of this specification.

[0024] Figure 8 This is a logical structure block diagram of an electronic device according to one embodiment of this specification. Detailed Implementation

[0025] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.

[0026] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0027] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0028] Gesture recognition has wide applications in many fields; for example, it allows for the control of smart devices through different gestures. Accurately recognizing user gestures is a prerequisite for gesture control. Currently, there are two main methods for recognizing user gestures. One method involves directly inputting an RGB or depth image containing the hand into a pre-trained deep learning model, which then directly outputs a prediction of the hand's pose in the image. However, this method lacks a criterion for determining the correctness of the prediction during the model inference stage, resulting in poor interpretability.

[0029] Another approach is to first obtain 2D or 3D positional information of hand keypoints from RGB or depth images containing the hand, and then input this information into a deep learning model. The deep learning model then outputs a prediction of the hand pose in the image. This method allows for secondary verification of the hand pose during the model inference phase using the positional information of the hand keypoints combined with reprojection errors, ensuring the reasonableness of the output hand pose. However, this method of predicting hand pose presents significant challenges in obtaining sample datasets when training the deep learning model. For example, it requires first acquiring RGB or depth images containing the hand, then labeling the hand keypoints in the RGB or depth images, and determining the hand pose in the image. Expanding the sample dataset using this method is difficult, limiting its size and resulting in fewer sample types that cannot cover the entire sample space. Consequently, the hand prediction results of the deep learning model trained on this sample dataset are often inaccurate. Furthermore, this method often yields limited sample data for specific gestures of interest, leading to lower accuracy in recognizing these specific gestures.

[0030] Based on this, embodiments of this application provide a method for generating sample data. This sample data can be used to train a deep learning model. This deep learning model can be used to predict hand pose information based on the position information of hand key points. Therefore, each set of sample data used to train the deep learning model may include the position information of hand key points as input samples of the model, and may include the hand pose information corresponding to the position information of hand key points as labels of input samples.

[0031] This application embodiment can generate a large amount of sample data based on the hand posture information of the user's predefined standard gesture and the pre-built hand kinematic model. Since the user does not need to manually annotate the key points of the hand, the difficulty of obtaining sample data is reduced, and a large amount of sample data can be obtained quickly and conveniently. Moreover, the constructed sample data can be more diverse and have higher accuracy, thereby improving the prediction accuracy of the trained deep learning model.

[0032] The sample data generation method provided in this application embodiment can be executed by a terminal device with the function of generating sample data. The function of generating sample data can exist in the terminal device in the form of an application, toolkit, component or service.

[0033] Specifically, such as Figure 1 As shown, the method for generating sample data provided in this application includes the following steps:

[0034] S102. Obtain hand posture information of standard gestures, wherein the hand posture information of standard gestures is preset by the user;

[0035] In this embodiment, the hand posture information can be the angle values ​​of the rotation angles of each joint of the hand. Typically, such as... Figure 2 As shown, the structure of the hand can be described using 21 key points (i.e., skeletal points). Figure 2 The three images, left and right, respectively illustrate the distribution of the hand's bones, the distribution of 21 key points in the hand, and the rotation angles of the hand's joints. The hand joint rotation angles include the wrist joint angle and the rotation angles of the 15 finger joints. From... Figure 2 As can be seen, hand posture is determined by the rotation angle of the hand joints. That is, when the size of the wrist joint rotation angle and the rotation angle of the 15 finger joints in the hand are determined, the corresponding hand posture can also be determined.

[0036] The standard gestures in this application embodiment can be artificially defined gestures with specific semantics, such as the "OK gesture," the "heart gesture," and the "V-sign," etc. Figure 3 The diagram illustrates different standard hand gestures. Since gesture recognition typically identifies gestures containing specific semantics (i.e., standard gestures), when generating sample data, one or more user-defined standard gestures can be obtained. Then, a large amount of sample data can be generated based on these standard gestures to improve the prediction accuracy of the trained deep learning model for standard gestures. The hand posture information (i.e., the angle values ​​of each hand joint rotation) for each standard gesture can be preset by the user. For example, for the standard gesture "OK," the user can preset the angle values ​​of each hand joint rotation in this gesture as the hand posture information for that standard gesture.

[0037] S202. Based on the hand posture information of the standard gesture and the preset hand kinematics model, generate multiple sets of sample data. Each set of sample data includes hand key point position information and hand posture information corresponding to the hand key point position information. The hand kinematics model is used to describe the influence of hand posture and hand bone length on the position of hand key points.

[0038] The hand kinematic model in this embodiment can be obtained by kinematic modeling the topological relationships of the hand, such as... Figure 2 As shown, the hand structure can be described using 21 hand keypoints (skeletal points). The bone length between any two bone points is denoted as Li, representing the length of the i-th bone. Bone lengths can be obtained empirically. Hand pose information typically includes the angle values ​​of hand joint rotations, including wrist joint rotations and 15 finger joint rotations. The wrist key rotation angle can be represented by R, and it can rotate arbitrarily in three-dimensional space. The wrist joint rotation angle can be represented using 6D to ensure the continuity of wrist rotation in space. 6D represents the rotation of the wrist joint in space using six parameters. The first three parameters have a magnitude of 1, the last three parameters have a magnitude of 1, and the vector formed by the first three parameters is orthogonal to the vector formed by the last three parameters.

[0039] Because finger joints have obvious motion constraints (e.g., fingers cannot bend outwards), Euler angles (Theta) can be used to represent them. The rotation of each finger joint in space can be represented by three angles, such as the angles of rotation around the X, Y, and Z axes. Furthermore, upper and lower bounds li and ui can be set for the Euler angles, where li and ui represent the constraints on the rotation angle of the i-th finger joint.

[0040] Clearly, the size of the joint angles and the length of the hand bones determine the location of the key points in the hand. Given a fixed joint angle and bone length, the locations of these joints are also determined. Therefore, a hand kinematic model can describe the influence of hand posture and bone length on the location of key hand points.

[0041] The kinematic model of the hand can be represented by the following formula (1):

[0042] P=Φ({Li},R,Theta) Formula (1)

[0043] Where P represents the key hand position information, which can be three-dimensional or two-dimensional; Φ is a function used to describe the influence of bone length, wrist joint angle, and finger joint angle on the key hand position information; {Li} represents the set of all joint bone lengths; R represents the wrist joint angle; and Theta represents the finger joint angle.

[0044] Based on the above hand kinematics model, it can be seen that after determining the gesture posture information and the length of the hand bones, the position information of the corresponding key points of the hand can be determined using the hand kinematics model, thereby generating multiple sets of sample data.

[0045] Therefore, after determining the hand posture information of the standard gesture, multiple sets of sample data can be generated using the standard gesture's hand posture information as a benchmark and a pre-defined hand kinematics model. Each set of sample data includes the position information of key hand points and the corresponding hand posture information. Then, the position information of the key hand points in the sample data can be used as the sample input of a deep learning model, and the hand posture information can be used as the label of the sample input for training the deep learning model.

[0046] The location information of hand key points can be two-dimensional or three-dimensional. For example, in some embodiments, the location information of hand key points can be the pixel coordinates of the hand key points in the image, and the deep learning model can predict hand pose information based on the pixel coordinates of the hand key points. In some embodiments, the location information of hand key points can also be three-dimensional coordinates in three-dimensional space, and the deep learning model can predict hand pose information based on the three-dimensional coordinates of the hand key points.

[0047] This application embodiment pre-constructs a hand kinematics model, and automatically generates a large amount of sample data based on the hand posture information of the user-defined standard gesture and the hand kinematics model. This data is used to train a deep learning model without requiring the user to manually annotate key hand points, reducing the difficulty of obtaining sample data. A large amount of sample data can be obtained quickly and conveniently, and the constructed sample data can be more diverse and accurate, thereby improving the prediction accuracy of the trained deep learning model.

[0048] To facilitate users in setting hand gesture information for standard hand gestures, a user interface can be provided. Users can input hand gesture information for standard hand gestures through this interface. Since standard hand gestures are generally only related to the angles of the finger joints, and the wrist joint angle does not affect the gesture—for example, the "OK" gesture remains the same even if the wrist joint is rotated—users only need to set the angle values ​​of 15 finger joints when setting the hand gesture information for standard hand gestures. In some embodiments, to help users determine whether the currently input finger joint angles are appropriate, a gesture model corresponding to the user's input command can be displayed in the interface, allowing users to determine the appropriateness of the currently input finger joint angles based on the gesture model.

[0049] When setting a standard gesture, users can directly input the specific angle values ​​of each finger joint. For example, recommended angle values ​​for each finger joint of a standard gesture can be loaded into the interactive interface. Users can adjust the angle values ​​of each joint based on the recommended values ​​to set the hand posture information of the standard gesture. Alternatively, users can directly input angle adjustment commands through the interactive interface. For instance, an angle adjustment control can be set for each joint angle in the interactive interface. When the user clicks the control, the angle value of that joint angle can be increased or decreased in steps. After receiving the user's angle adjustment command, the corresponding gesture model can be displayed on the interactive interface. When the user determines that the current gesture has been adjusted to the desired standard gesture (e.g., the OK gesture) based on the displayed gesture model, they can input the command to end the angle adjustment. At this point, the gesture posture information corresponding to the gesture model currently displayed on the interactive interface can be used as the hand posture information of the standard gesture and stored.

[0050] Figure 4(a) illustrates a user interface for determining a standard gesture in one embodiment of this application. When setting the finger joint angle of a standard gesture, the user can select the finger to be operated, such as the thumb or index finger, and then select the joint to be operated. For example, if a finger contains different joints, the user can select the joint to be operated. Then, the user can select the corresponding rotation axis (X, Y, Z axes) and adjust the angle relative to each rotation axis. During the adjustment process, the interface displays the adjusted gesture posture model so that the user can determine whether the standard gesture adjustment is complete based on the model. After adjustment, the user can save the adjusted gesture as a standard gesture and store it in the gesture library. Simultaneously, the user can also load standard gestures from the gesture library for viewing.

[0051] In some embodiments, when generating multiple sets of sample data based on the hand posture information of a standard gesture and a preset hand kinematic model, the hand posture information of the standard gesture can be changed to obtain multiple sets of new hand posture information. Then, for each set of new hand posture information, the hand key point position information corresponding to that set of new hand posture information can be determined based on the hand kinematic model, thereby obtaining multiple sets of sample data.

[0052] Based on the hand kinematics model, it is known that changes in parameters such as finger joint angles, wrist joint angles, and hand bone lengths will alter the corresponding key hand position information. Therefore, by altering these parameters in the standard hand gesture posture information, multiple sets of new hand posture information can be obtained. Then, based on the hand kinematics model, the corresponding key hand position information can be determined, resulting in multiple sets of sample data. This method can generate a large amount of sample data in the sample space, ensuring the diversity and comprehensiveness of the sample data.

[0053] Because a single standard gesture can correspond to many sets of finger joint angles—for example, taking the "OK" gesture—even if the finger joint angles change, as long as the finger shape remains "OK," it can still be recognized as the "OK" gesture. Therefore, for each standard gesture, the angle values ​​of each finger joint can be continuously changed within a small range, ensuring that the changed gesture still has a high similarity to the standard gesture, meaning it can still be recognized as that standard gesture. This enriches the sample data corresponding to each standard gesture, making the sample data more diverse and improving the prediction accuracy of that standard gesture.

[0054] Therefore, in some embodiments, when changing the hand posture information of a standard gesture to obtain multiple sets of new hand posture information, multiple first angle changes can be randomly generated within a first angle range for each finger joint angle of the standard gesture. Then, the angle value of each finger joint angle is changed based on the multiple first angle changes to obtain multiple sets of new hand posture information. The similarity between the multiple sets of new hand posture information obtained based on the first angle changes and the posture information of the standard gesture is greater than a first preset threshold. The first angle range can be a pre-set angle range that ensures the changed finger joint angles are within the constraints of the finger joints, and that the shape of the gesture formed after the finger joint angle change does not change significantly, i.e., it has a high similarity to the predefined standard gesture.

[0055] Besides standard gestures, there are also non-standard gestures without specific semantics. Therefore, when generating sample data, some gestures without specific semantics can be automatically and randomly generated within the constraint range of joint angles to fill the sample space between standard gestures, avoiding the deep learning model being trained to converge only to the vicinity of standard gestures. Thus, in some embodiments, when generating gestures without specific semantics, a specific gesture can be preset. This specified gesture can be a standard gesture or a gesture where the angle of each finger joint is 0, as shown in Figure 4(b). For each finger joint angle of the specified gesture, multiple second angle changes can be randomly generated within a second angle range. Then, based on these multiple second angle changes, the angle value of the finger joint angle can be changed to obtain multiple sets of new hand posture information. The similarity between the multiple sets of new hand posture information obtained based on the second angle changes and the posture information of the standard gesture is less than a second preset threshold, and the second preset threshold is less than a first preset threshold. When the specified gesture is a standard gesture, its second angle range can be a pre-set angle range that is different from the first angle range. The angle value in this angle range is greater than the angle value in the first angle range. This angle range can ensure that the changed finger joint angle is within the constraint range of the finger joint, and the shape of the gesture formed after the change of the finger joint angle changes significantly, that is, the similarity with the predefined standard gesture is low.

[0056] Of course, the length of the hand bones varies from person to person. In order to ensure the diversity and comprehensiveness of the samples, in some embodiments, when changing the hand posture information of the standard gesture to obtain multiple sets of new hand posture information, multiple scaling factors can be randomly generated within a preset range first, and then the length of the hand bones of the standard gesture can be scaled using these multiple scaling factors to obtain multiple sets of new hand posture information.

[0057] Unlike finger joints, wrist joints rotate unrestricted in space. Therefore, it is necessary to ensure that the distribution of wrist joint angles in the rotation space is dense and uniform to ensure that as much of the sample space as possible is covered. Thus, in some embodiments, when changing the hand posture information of a standard gesture to obtain multiple sets of new hand posture information, the wrist joint angles of the standard gesture can be uniformly sampled in the rotation space to obtain multiple sets of wrist joint angles, thereby generating multiple sets of new hand posture information.

[0058] Since the first three parameters in the 6D representation have a magnitude of 1, and the last three parameters have a magnitude of 1, and the vector formed by the first three parameters is orthogonal to the vector formed by the last three parameters, in some embodiments, when uniformly sampling the wrist joint angle of a standard gesture in rotational space to obtain multiple sets of wrist joint angles, it is possible to first uniformly sample on a sphere with a radius of 1 to obtain multiple first sampling points. The spatial coordinates of each first sampling point are used as the first three parameters in the 6D representation of the wrist joint angle. Then, for each first sampling point, the line connecting the first sampling point and the center of the sphere is determined, and a circle with a radius of 1 perpendicular to the line is determined. Uniform sampling is performed on this circle, for example, by sampling a second sampling point at predetermined angle intervals to obtain multiple second sampling points. The spatial coordinates of each second sampling point are used as the last three parameters in the 6D representation of the wrist joint angle, thereby obtaining multiple sets of 6D representations of the wrist joint angle. In some embodiments, the spatial coordinates of the first sampling point can be coordinates with the center of the sphere as the origin, and the spatial coordinates of the second sampling point can be coordinates with the center of the circle as the origin.

[0059] The wrist joint angle can be represented using 6D, for example, as [x,y,z,a,b,c]. According to the definition of 6D representation, x, y, z and a, b, c are taken from the first two columns of the rotation matrix corresponding to the 6D representation. Therefore, vectors [x,y,z] and [a,b,c] are orthogonal to each other, and their magnitudes are both 1. Thus, the points [x,y,z] are distributed on a sphere with a radius of 1. Let the point be X. To ensure the uniform distribution of R, we first need to consider the uniform distribution of point X on the sphere. Here, we introduce a Fibonacci grid, which can ensure the uniformity and density of point X, as shown in Figure 5(a). The spatial coordinates of point X are calculated by formula (2):

[0060]

[0061] Where P represents the total number of sampling points, a constant. Using the golden ratio, [xp,yp,zp] represents the spatial coordinates of the p-th sampling point Xp, as shown in Figure 5(b) outside the circle a. Let the vector xyz = [xp,yp,zp] represent the straight line in Figure 5(b); let the vector abc = [a,b,c]. Since xyz is orthogonal to abc, the point [a,b,c] is located on the circle with a radius of 1 that is perpendicular to the vector xyz, as shown in Figure 5(b) on the circle a. Let the number of sampling points be Q, then the point coordinates [aq,bq,cq] can be obtained directly by uniform sampling of the angle. Therefore, the final 6D is represented as [xp,yp,zp,aq,bq,cq].

[0062] As can be seen from the above embodiments, when changing the hand posture information of a standard gesture to generate multiple sets of new hand posture information, different change methods can be adopted. For example, various change methods can be used, such as changing the finger joint angle, changing the wrist joint angle, and changing the length of the hand bones. For different change methods, the user can pre-set the sampling quantity corresponding to each change method according to their needs. For example, taking changing the finger joint angle as an example, multiple hand postures similar to the standard gesture can be generated by generating multiple first angle change quantities, or multiple non-standard hand postures that differ significantly from the standard gesture can be generated by generating multiple second angle change quantities. For each change method, the corresponding sampling data can be pre-set. For example, the number of first angle change quantities and second angle change quantities can be pre-set by the user. Furthermore, the number of sample data generated based on different standard gestures can also be different.

[0063] Therefore, in some embodiments, before changing the hand posture information of the standard gesture to obtain multiple sets of new hand posture information, configuration information input by the user through the interactive interface can be obtained first. The configuration information is used to indicate the number of samples corresponding to each change method when changing the hand posture information of the standard gesture using different change methods. Then, the hand posture information of the standard gesture is changed based on the number of samples corresponding to each change method to obtain multiple sets of new hand posture information.

[0064] After training a deep learning model using generated sample data, if the predicted accuracy of the trained deep learning model for a specific gesture is low, more samples can be generated for that specific gesture. For example, the finger joint angles, wrist joint angles, or bone lengths of the standard gesture can be changed with finer granularity to obtain more sample data for that specific gesture. This data is then used to train the deep learning model, improving its prediction accuracy for that standard gesture. Therefore, in some embodiments, when the predicted accuracy of the trained deep learning model for the target gesture is lower than a preset accuracy, the hand posture information of the target gesture is changed to obtain multiple sets of new hand posture information. Based on the obtained new hand posture information and the hand kinematics model, the key point position information corresponding to these new hand posture information is determined to generate multiple sets of sample data for training the deep learning model. In this way, the structure of the sample data can be adjusted in a targeted manner, enabling the trained model to have high prediction accuracy for various gestures. The target gesture can be a standard gesture or a non-standard gesture.

[0065] Since the input to a deep learning model can be either the 3D position information of hand keypoints (e.g., 3D coordinates in 3D space) or the 2D position information of hand keypoints in an image (e.g., pixel coordinates in an image), in some embodiments, if the input to the deep learning model is the pixel coordinates of hand keypoints in an image, that is, the position information of hand keypoints in the sample data can be the pixel coordinates of hand keypoints in an image. If the position information of hand keypoints determined based on the kinematic model is the 3D position information of hand keypoints in 3D space, when generating multiple sets of sample data based on the hand posture information of standard gestures and the preset hand kinematic model, multiple sets of 3D position information of hand keypoints in 3D space, as well as hand posture information corresponding to the 3D position information of hand keypoints, can be generated based on the hand posture information of standard gestures and the preset hand kinematic model. Then, the 3D position information of hand keypoints is orthogonally projected to obtain the pixel coordinates of hand keypoints in the image.

[0066] To further explain the sample data generation method provided in the embodiments of this application, the following explanation is based on a specific embodiment.

[0067] When training a deep learning model for predicting hand pose based on hand keypoint location information, the difficulty in expanding the sample data is high because it is necessary to annotate the hand keypoints and hand poses in the image when constructing the sample data, and it is impossible to guarantee the diversity of the samples, resulting in low prediction accuracy of the trained deep learning model.

[0068] In order to quickly obtain a large amount of high-precision sample data, this embodiment can generate a large amount of sample data for training deep learning models based on hand kinematics models and user-defined standard gestures. This can ensure the diversity and comprehensiveness of sample data and improve the prediction accuracy of the trained model.

[0069] The process of constructing sample data can be referenced. Figure 6 Specifically, it includes the following steps:

[0070] 1. Construct a kinematic model of the hand:

[0071] Perform kinematic modeling of the topological relationships of the hand, such as Figure 2The structure of the hand can be described using 21 skeletal points (keypoints). The bone length between any two skeletal points is denoted as Li, representing the length of the i-th bone. The bone lengths are statistical values ​​from relevant literature. Hand joint angles include wrist joint angle R (global rotation) and 15 finger joint angles Theta (local rotation). The wrist joint angle R uses a 6D representation, employing 6 parameters to describe spatial rotation, ensuring the continuity of wrist rotation in space. Finger joints, due to significant motion constraints (e.g., fingers cannot bend outward), are represented using Euler angles, with upper and lower bounds li and ui set. li and ui represent the constraints of the i-th finger joint angle. Let FK be the function describing the forward kinematics of the hand. Then, the coordinates of the 3D keypoints of the hand can be obtained as P3D = FK({Li}, R, Theta), where {Li} represents the set of all joint bone lengths.

[0072] After establishing the kinematic model of the hand, it can be seen that in order to ensure the diversity of 3D key points, it is necessary to change three types of parameters: {Li}, R, and Theta.

[0073] (2) User-defined N types of standard gestures and the number of samples for each variation.

[0074] Since gesture estimation usually requires the recognition of gestures containing specific semantics, users can predefine some standard gestures and then generate sample data based on the standard gestures, which can improve the prediction accuracy of the trained model for standard gestures.

[0075] Standard gestures can be generated by the user manually adjusting the Theta parameter (the R parameter remains unchanged for now). When defining standard gestures, an interactive interface can be provided. Users can click on controls on the interface to increase or decrease the Theta angle of each finger joint. Simultaneously, the interface displays the corresponding gesture model after the user adjusts the finger joint angles, allowing the user to easily observe whether the desired standard gesture has been achieved. All standard gestures containing semantics can be predefined through an exhaustive approach to ensure sample diversity. Through the manual adjustment method, adjusting each standard gesture takes approximately <1 minute.

[0076] Since users can change the three types of parameters {Li}, R, and Theta to generate sample data, users can define the sampling quantity corresponding to each change method according to their needs. For example, users can set the sampling quantity corresponding to each change method through the interactive interface to ensure that the samples meet the actual needs.

[0077] (3) Change the finger joint angle of the standard hand gesture to generate a new hand posture.

[0078] After predefining standard gestures, additional movements can be superimposed on the finger angles corresponding to these gestures to further enhance the diversity of the gestures. Since finger joint angles have certain constraints, changes to finger joint angles must be made within these constraints.

[0079] Since a standard gesture can have many sets of finger joint angles—for example, a slight change in the finger joint angle does not change the meaning of the gesture—the finger joint angle can be varied within a small range to make the samples of each standard gesture more diverse. Specifically, a variable delta is introduced for the joint angle of each finger, uniformly distributed and randomly generated between -15 degrees and 15 degrees. Delta is in degrees. The superimposed finger joint angle is Theta + delta. The superimposed joint angle still needs to satisfy the upper and lower bounds li and ui, hence the 3D keypoint P3D = FK({Li},R,Theta + delta). Assuming there are N types of standard gestures, and each standard gesture generates M delta variations, the number of samples obtained is currently N*M. Typically, N is on the order of 10, and M is on the order of 100.

[0080] In addition to standard gestures, some gestures without specific semantics need to be automatically and randomly generated within the constraints to fill the sample space between standard gestures, so as to avoid the trained model only converging to the vicinity of standard gestures. The number of samples without semantics is large, with a number of O, on the order of 10,000. Therefore, a total of N*M+O samples are obtained in this stage.

[0081] In addition, if the trained model has low prediction accuracy for certain standard gestures, the sample data for those standard gestures can be appropriately increased.

[0082] (4) Changing the wrist joint angle to generate hand posture

[0083] The wrist joint rotation angle R is uniformly sampled in the rotation space. Unlike the finger joints, the wrist rotation is unconstrained in the rotation space. Therefore, it is necessary to ensure that the distribution of R in the space is dense and uniform in order to ensure that all sample spaces are covered as much as possible. The wrist joint rotation angle can be represented by 6D, which can be represented as [x,y,z,a,b,c]. According to the definition of 6D representation, x,y,z and a,b,c are taken from the first two columns of the rotation matrix corresponding to the 6D representation. Therefore, the vectors [x,y,z] and [a,b,c] are orthogonal to each other and both have a magnitude of 1. It can be seen that the point [x,y,z] is distributed on a sphere with a radius of 1. Let the point be X. To ensure the uniform distribution of R, we first need to consider the uniform distribution of point X on the sphere. Here, we introduce a Fibonacci grid, which can ensure the uniformity and density of point X. The sampling effect is shown in Figure 5(a). The spatial coordinates of point X (with the center of the sphere as the origin) are calculated by formula (2) as follows:

[0084]

[0085] Where P represents the total number of sampling points, a constant. Using the golden ratio, [xp,yp,zp] represents the spatial coordinates of the p-th sampling point Xp, as shown in Figure 5(b) outside the circle a. Let the vector xyz = [xp,yp,zp] represent the straight line in Figure 5(b); let the vector abc = [a,b,c]. Since xyz is orthogonal to abc, the point [a,b,c] is located on the circle with a radius of 1, which is perpendicular to the vector xyz, as shown in Figure 5(b) on the circle a. Let the number of sampling points be Q, then the point coordinates [aq,bq,cq] (with the center of the circle as the origin) can be obtained directly by uniform sampling of the angle. Therefore, the final 6D is represented as [xp,yp,zp,aq,bq,cq]. Given 3D keypoints P3D=FK({Li},[xp,yp,zp,aq,bq,cq],Theta+delta), the total number of samples collected is P*Q*(N*M+O), where P and Q are on the order of 100 and 10, respectively.

[0086] (5) Change the length of the bones to generate hand gestures

[0087] The bone length {Li} is randomly modified within a certain range. Considering that the length of the hand joint bones varies from person to person, a scaling factor is needed to scale the bone length of each joint. Let the scaling factor be αi, which represents the scaling factor of the i-th bone. It is uniformly distributed and randomly generated between 0.75 and 1.25. Therefore, the 3D key point P3D = FK({Li*αi},[xp,yp,zp,aq,bq,cq],Theta+delta).

[0088] (6) After generating a large number of new hand poses, the position information of the key points of the hand corresponding to the hand poses is generated based on the kinematic model. In this way, a large-scale dataset with uniform coverage of the sample space can be obtained. In this dataset, the 3D key point information of the hand is used as the input signal of the sample, and the bone length {Li*αi}, wrist rotation R, and finger joint rotation angle theta are used as the annotation of the sample.

[0089] In addition, if the input of the model is the 2D position information of the key points of the hand, the 2D key points can be obtained by setting the orthogonal projection parameters s and t and then performing a projection transformation. These can also be saved as the input signal of the sample. Here, s represents scaling and t represents translation. Just make sure that the 2D key points after projection are located within the image.

[0090] Corresponding to the above method, embodiments of this application also provide a sample data generation apparatus, such as... Figure 7 As shown, the device includes:

[0091] The sample data is used to train a deep learning model, which is used to predict hand pose information based on the hand key point location information. The device includes:

[0092] The acquisition module 71 is used to acquire hand posture information of a standard gesture, which is preset by the user.

[0093] The sample data generation module 72 is used to generate multiple sets of sample data based on the hand posture information of the standard gesture and the preset hand kinematic model. Each set of sample data includes hand key point position information and hand posture information corresponding to the hand key point position information.

[0094] The hand kinematic model is used to describe the influence of hand posture and hand bone length on the location of key hand points.

[0095] In some embodiments, the hand gesture information of the standard gesture is determined based on the following method:

[0096] In response to an angle adjustment command input by a user through an interactive interface, a gesture model corresponding to the angle adjustment command is presented on the interactive interface, wherein the angle adjustment command is used to adjust the angle value of each finger joint of the standard gesture.

[0097] In response to the user's input command to adjust the end angle, the hand posture information of the gesture model presented on the interactive interface at the current moment is used as the hand posture information of the standard gesture.

[0098] In some embodiments, when the sample data generation module generates multiple sets of sample data based on the hand posture information of the standard gesture and a preset hand kinematic model, it is specifically used for:

[0099] By changing the hand posture information of the standard gesture, multiple sets of new hand posture information are obtained;

[0100] For each new set of hand posture information, the hand key point position information corresponding to the new hand posture information is determined based on the hand kinematic model to obtain the multiple sets of sample data.

[0101] In some implementations, when the sample data generation module is used to change the hand posture information of the standard gesture to obtain multiple sets of new hand posture information, it is specifically used for:

[0102] For each finger joint rotation angle of the standard gesture, multiple first angle changes are randomly generated within the first angle range;

[0103] The angle value of the finger joint rotation is changed based on the multiple first angle changes to obtain multiple sets of new hand posture information, wherein the similarity between the obtained multiple sets of new hand posture information and the posture information of the standard gesture is greater than a first preset threshold.

[0104] In some embodiments, when the sample data generation module is used to change the hand posture information of the standard gesture to obtain multiple sets of new hand posture information, it is specifically used for:

[0105] Randomly generate multiple scaling factors within a preset range;

[0106] The length of the hand bones of the standard gesture is scaled using the multiple scaling factors to obtain multiple sets of new hand posture information.

[0107] In some embodiments, when the sample data generation module is used to change the hand posture information of the standard gesture to obtain multiple sets of new hand posture information, it is specifically used for:

[0108] The wrist joint angle of the standard gesture is uniformly sampled in the rotation space to obtain multiple sets of wrist joint angles, thereby generating multiple sets of new hand posture information.

[0109] In some embodiments, when the sample data generation module is used to uniformly sample the wrist joint angle of the standard gesture in rotation space to obtain multiple sets of wrist joint angles, it is specifically used for:

[0110] Uniform sampling is performed on a sphere with a radius of 1 to obtain multiple first sampling points. The spatial coordinates of each first sampling point are used as the first three parameters in the 6D representation of the wrist joint rotation angle.

[0111] For each first sampling point, determine the line connecting the first sampling point to the center of the sphere;

[0112] Uniform sampling is performed on a circle with a radius of 1 perpendicular to the line connecting the two points to obtain multiple second sampling points. The spatial coordinates of each second sampling point are used as the last three parameters of the 6D representation of the wrist joint angle to obtain multiple sets of 6D representations of the wrist joint angle.

[0113] In some embodiments, the device is further configured to:

[0114] For each finger joint rotation angle of a specified gesture, multiple second angle variations are randomly generated within a second angle range, wherein the specified gesture includes the standard gesture and / or a gesture where each finger joint has a key rotation angle of 0°;

[0115] The angle value of the finger joint rotation of the specified gesture is changed based on the multiple second angle changes to obtain multiple sets of new hand posture information, wherein the similarity between the obtained multiple sets of new hand posture information and the posture information of the standard gesture is less than a second preset threshold.

[0116] Multiple sets of sample data are generated based on the obtained new hand posture information and the hand kinematic model.

[0117] In some embodiments, the device is further configured to:

[0118] The configuration information input by the user through the interactive interface is obtained. The configuration information is used to indicate the number of samples corresponding to each change method when the hand posture information of the standard gesture is changed in different ways.

[0119] The hand posture information of the standard gesture is changed based on the number of samples corresponding to each change method to obtain the multiple sets of new hand posture information.

[0120] In some embodiments, after training the deep learning model using the generated sample data, the apparatus is further configured to:

[0121] If the prediction accuracy of the deep learning model for the target gesture is lower than the preset accuracy after training, the gesture posture information of the target gesture is changed to obtain multiple sets of new gesture posture information. Based on the obtained new gesture posture information and the hand kinematics model, multiple sets of sample data are generated for training the deep learning model.

[0122] In some embodiments, the hand key point location information in the multiple sets of sample data includes the pixel coordinates of the hand key points in the image. When the sample data generation module generates multiple sets of sample data based on the hand posture information of the standard gesture and a preset hand kinematic model, it is specifically used for:

[0123] Based on the hand posture information of the standard gesture and the preset hand kinematics model, multiple sets of three-dimensional position information of hand key points in three-dimensional space are generated, as well as hand posture information corresponding to the three-dimensional position information of the hand key points.

[0124] The three-dimensional position information of the key hand points is orthogonally projected to obtain the pixel coordinates of the key hand points in the image.

[0125] Furthermore, embodiments of this application also provide an electronic device, such as... Figure 8 As shown, the electronic device includes a processor 81, a memory 82, and a computer program stored in the memory 82 that can be executed by the processor 81. When the processor 81 executes the computer program, it implements the method described in any of the above embodiments.

[0126] Accordingly, this application also provides a computer storage medium storing a program, which, when executed by a processor, implements the methods described in any of the above embodiments.

[0127] The embodiments of this specification may take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing program code. Computer-usable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to: phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0128] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0129] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the description disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0130] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

[0131] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for generating sample data, characterized in that, The sample data is used to train a deep learning model, which is used to predict hand pose information based on the location information of hand key points. The method includes: Acquire hand gesture information for a standard gesture, which is preset by the user. The hand posture information of the standard gesture is changed to obtain multiple sets of new hand posture information; for each set of new hand posture information, the hand key point position information corresponding to the new hand posture information is determined based on the hand kinematics model to obtain multiple sets of sample data. Each set of sample data includes the hand key point position information and the hand posture information corresponding to the hand key point position information. The hand kinematic model is used to describe the influence of hand posture and hand bone length on the location of key hand points. By changing the hand posture information of the standard gesture, multiple sets of new hand posture information are obtained, including: For each finger joint rotation angle of the standard gesture, multiple first angle changes are randomly generated within the first angle range; The angle value of the finger joint rotation is changed based on the multiple first angle changes to obtain multiple sets of new hand posture information, wherein the similarity between the obtained multiple sets of new hand posture information and the posture information of the standard gesture is greater than a first preset threshold.

2. The method according to claim 1, characterized in that, The hand gesture information of the standard hand gesture is determined based on the following method: In response to an angle adjustment command input by a user through an interactive interface, a gesture model corresponding to the angle adjustment command is presented on the interactive interface, wherein the angle adjustment command is used to adjust the angle value of each finger joint of the standard gesture. In response to the user's input command to adjust the end angle, the hand posture information of the gesture model presented on the interactive interface at the current moment is used as the hand posture information of the standard gesture.

3. The method according to claim 1, characterized in that, By changing the hand posture information of the standard gesture, multiple sets of new hand posture information are obtained, including: Randomly generate multiple scaling factors within a preset range; The length of the hand bones of the standard gesture is scaled using the multiple scaling factors to obtain multiple sets of new hand posture information.

4. The method according to claim 1, characterized in that, By changing the hand posture information of the standard gesture, multiple sets of new hand posture information are obtained, including: The wrist joint angle of the standard gesture is uniformly sampled in the rotation space to obtain multiple sets of wrist joint angles, thereby generating multiple sets of new hand posture information.

5. The method according to claim 4, characterized in that, The wrist joint angle of the standard gesture is uniformly sampled in rotation space to obtain multiple sets of wrist joint angles, including: Uniform sampling is performed on a sphere with a radius of 1 to obtain multiple first sampling points. The spatial coordinates of each first sampling point are used as the first three parameters in the 6D representation of the wrist joint rotation angle. For each first sampling point, determine the line connecting the first sampling point to the center of the sphere; Uniform sampling is performed on a circle with a radius of 1 perpendicular to the line connecting the two points to obtain multiple second sampling points. The spatial coordinates of each second sampling point are used as the last three parameters of the 6D representation of the wrist joint angle to obtain multiple sets of 6D representations of the wrist joint angle.

6. The method according to claim 1, characterized in that, The method further includes: For each finger joint rotation angle of a specified gesture, multiple second angle variations are randomly generated within a second angle range, wherein the specified gesture includes the standard gesture and / or a gesture where each finger joint has a key rotation angle of 0°; The angle value of the finger joint rotation of the specified gesture is changed based on the multiple second angle changes to obtain multiple sets of new hand posture information, wherein the similarity between the obtained multiple sets of new hand posture information and the posture information of the standard gesture is less than a second preset threshold. Multiple sets of sample data are generated based on the obtained new hand posture information and the hand kinematic model.

7. The method according to claim 1, characterized in that, The method further includes: The configuration information input by the user through the interactive interface is obtained. The configuration information is used to indicate the number of samples corresponding to each change method when the hand posture information of the standard gesture is changed in different ways. The hand posture information of the standard gesture is changed based on the number of samples corresponding to each change method to obtain the multiple sets of new hand posture information.

8. The method according to claim 1, characterized in that, After training the deep learning model using the generated sample data, the method further includes: If the prediction accuracy of the deep learning model for the target gesture is lower than the preset accuracy after training, the gesture posture information of the target gesture is changed to obtain multiple sets of new gesture posture information. Based on the obtained new gesture posture information and the hand kinematics model, multiple sets of sample data are generated for training the deep learning model.

9. The method according to claim 1, characterized in that, The hand key point location information in the multiple sets of sample data includes the pixel coordinates of the hand key points in the image. Multiple sets of sample data are generated based on the hand posture information of the standard gesture and a preset hand kinematic model, including: Based on the hand posture information of the standard gesture and the preset hand kinematics model, multiple sets of three-dimensional position information of hand key points in three-dimensional space are generated, as well as hand posture information corresponding to the three-dimensional position information of the hand key points. The three-dimensional position information of the key hand points is orthogonally projected to obtain the pixel coordinates of the key hand points in the image.

10. A sample data generation apparatus, characterized in that, The sample data is used to train a deep learning model, which is used to predict hand pose information based on the hand key point location information. The device includes: The acquisition module is used to acquire hand posture information of standard gestures, which is preset by the user. The sample data generation module is used to change the hand posture information of the standard gesture to obtain multiple sets of new hand posture information; for each set of new hand posture information, the key hand point position information corresponding to the new hand posture information is determined based on the hand kinematics model to obtain multiple sets of sample data. Each set of sample data includes the key hand point position information and the hand posture information corresponding to the key hand point position information. The hand kinematic model is used to describe the influence of hand posture and hand bone length on the location of key hand points. The sample data generation module is used to change the hand posture information of the standard gesture to obtain multiple sets of new hand posture information, specifically for: For each finger joint rotation angle of the standard gesture, multiple first angle changes are randomly generated within the first angle range; The angle value of the finger joint rotation is changed based on the multiple first angle changes to obtain multiple sets of new hand posture information, wherein the similarity between the obtained multiple sets of new hand posture information and the posture information of the standard gesture is greater than a first preset threshold.

11. An electronic device, characterized in that, The electronic device includes a processor, a memory, and a computer program stored in the memory that can be executed by the processor, wherein the processor executes the computer program to implement the method according to any one of claims 1-9.