Sample generation method and hand action recognition model training method
By acquiring and combining the freedom data of hand videos, training samples that are more in line with real hand movements are generated, which solves the problem of insufficient training samples in the prior art and improves the robustness and recognition accuracy of the hand movement recognition model.
Patent Information
- Application Number
- CN202311828893.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
Due to insufficient training samples, existing hand motion recognition models cannot be trained to obtain a robust hand motion recognition model.
By acquiring the first video and the second video, including multiple frames of continuous images of the first hand and the second hand, the degree of freedom data of the wrist and finger are determined, combined, and rendered into the hand video as a training sample for training the hand motion recognition model.
The generated training samples are more in line with real hand movements, improving the robustness of the hand movement recognition model, and allowing the model to more accurately identify the user's hand movements.
Smart Images

Figure CN120219867A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to data processing methods, and in particular, to a sample generation method and a hand gesture recognition model training method. Background Art
[0002] In the field of VR (Virtual Reality), users can interact with VR devices through hand gestures. High accuracy in hand gesture recognition can improve the user experience of VR devices.
[0003] Currently, a hand gesture recognition model is usually used to recognize users' hand gestures. Among them, the hand gesture recognition model is pre-trained, and insufficient training samples will result in an inability to train a hand gesture recognition model with strong robustness. Summary of the Invention
[0004] Embodiments of the present disclosure provide a sample generation method and a hand gesture recognition model training method, which can generate training samples for training a hand gesture recognition model with strong robustness.
[0005] In a first aspect, embodiments of the present disclosure provide a sample generation method, including: obtaining a first video and a second video, where the first video includes multiple consecutive first images of a first hand, and the second video includes multiple consecutive second images of a second hand; determining the wrist degree-of-freedom data of the first images and the finger degree-of-freedom data of the second images; combining the wrist degree-of-freedom data and the finger degree-of-freedom data to obtain multiple hand degree-of-freedom data; rendering a set of hand degree-of-freedom data to obtain a hand video, where the set of hand degree-of-freedom data includes some or all of the multiple hand degree-of-freedom data.
[0006] In a second aspect, embodiments of the present disclosure provide a hand gesture recognition model training method, including: obtaining training samples, where the training samples are obtained according to the sample generation method in the first aspect; obtaining label data of the training samples, where the label data is the set of hand degree-of-freedom data for generating the training samples; training a hand gesture recognition model according to the training samples and the label data to obtain a trained hand gesture recognition model.
[0007] In a third aspect, embodiments of the present disclosure provide a hand gesture recognition method, including: obtaining a hand video of a user; inputting the hand video into a hand gesture recognition model for recognition to obtain a set of hand degree-of-freedom data, where the set of hand degree-of-freedom data is used to represent the hand gesture of the user, and the hand gesture recognition model is obtained according to the hand gesture recognition model training method in the second aspect.
[0008] In a fourth aspect, embodiments of the present disclosure provide a sample generation device, including:
[0009] An acquisition unit for acquiring a first video and a second video, where the first video includes multiple consecutive first images of a first hand, and the second video includes multiple consecutive second images of a second hand;
[0010] A determination unit for determining the wrist degree - of - freedom data of the first images and the finger degree - of - freedom data of the second images;
[0011] A combination unit for combining the wrist degree - of - freedom data and the finger degree - of - freedom data to obtain multiple hand degree - of - freedom data;
[0012] A rendering unit for rendering a set of hand degree - of - freedom data to obtain a hand video, where the set of hand degree - of - freedom data includes some or all of the multiple hand degree - of - freedom data.
[0013] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, including: at least one processor and a memory;
[0014] The memory stores computer - executable instructions;
[0015] At least one processor executes the computer - executable instructions stored in the memory, so that at least one processor executes the methods provided in the above - mentioned first aspect, second aspect, and / or third aspect.
[0016] In a sixth aspect, an embodiment of the present disclosure provides a computer - readable storage medium, in which computer - executable instructions are stored. When a processor executes the computer - executable instructions, the methods provided in the above - mentioned first aspect, second aspect, and / or third aspect are implemented.
[0017] In a seventh aspect, according to one or more embodiments of the present disclosure, a computer program product is provided. The computer program product includes computer - executable instructions. When a processor executes the computer - executable instructions, the methods provided in the above - mentioned first aspect, second aspect, and / or third aspect are implemented.
[0018] The sample generation method provided in this embodiment, by acquiring a first video and a second video, where the first video includes multiple consecutive first images of a first hand, and the second video includes multiple consecutive second images of a second hand; determining the wrist degree - of - freedom data of the first images and the finger degree - of - freedom data of the second images; combining the wrist degree - of - freedom data and the finger degree - of - freedom data to obtain multiple hand degree - of - freedom data; rendering a set of hand degree - of - freedom data to obtain a hand video, where the set of hand degree - of - freedom data includes some or all of the multiple hand degree - of - freedom data, can obtain a hand video that more conforms to real hand movements as a training sample, and using this training sample can train a hand - motion recognition model with better robustness. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is a flowchart of the steps of a sample generation method provided by an embodiment of the present disclosure;
[0021] Figure 2 It is a schematic diagram of generating hand degree-of-freedom data provided by an embodiment of the present disclosure;
[0022] Figure 3 It is a flowchart of the steps of a hand gesture recognition model training method provided by an embodiment of the present disclosure;
[0023] Figure 4 It is a flowchart of the steps of a hand gesture recognition method provided by an embodiment of the present disclosure;
[0024] Figure 5 It is a block diagram of the structure of a sample generation device provided by an embodiment of the present disclosure;
[0025] Figure 6 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.
[0027] In the related art, gestures are sampled from a gesture set, the gesture is rendered to obtain an image containing a hand, and then the image is used to train a hand gesture recognition model. Among them, the gesture set contains multiple gestures, and these gestures are represented by 26 degrees-of-freedom data. These gestures are determined by a computer or collected for a single hand. The gestures in the gesture set are not continuous. Therefore, the hand gestures in the images obtained using these gestures are also not continuous. Furthermore, using these images to train a hand gesture recognition model enables the hand gesture recognition model to only learn the gestures of a single image and cannot learn the continuous actions of the gestures between images. Therefore, the trained hand gesture recognition model cannot accurately recognize hand gestures and has low robustness.
[0028] Based on the above problems, the first video and the second video adopted by the present disclosure are collected for real hand movements and have authenticity. In addition, by combining the separately collected wrist degree-of-freedom data and the separately collected finger degree-of-freedom data, the coverage of the hand degree-of-freedom data can be improved. Rendering the hand video based on this hand degree-of-freedom data makes the obtained hand video have authenticity, continuity, and high coverage. Using this hand video to train the hand movement recognition model can improve the robustness of the hand movement recognition model.
[0029] One application scenario of the present disclosure is the recognition of hand movements of VR (virtual reality) devices. Among them, when a user uses a VR device, the hand will have a series of continuous movements (such as clicking, swiping, etc.). After the camera of the VR device captures these movements, the VR device will recognize these movements and control the content displayed by the VR device according to the recognized movements. Further, the VR device recognizes the captured movements through a hand movement recognition model, and this hand movement recognition model is pre-trained according to training samples. When the robustness of the hand movement recognition model is higher, it can more accurately recognize various hand movements of the user, and thus the interaction between the user and the VR device can be realized, improving the user experience of using the VR device.
[0030] Reference Figure 1 , is a schematic flowchart of the sample generation method provided by the embodiment of the present disclosure. As Figure 1 shown, the sample generation method specifically includes the following steps:
[0031] S101. Obtain a first video and a second video.
[0032] Among them, the first video includes multiple consecutive first images of a first hand, and the second video includes multiple consecutive second images of a second hand.
[0033] Specifically, multiple consecutive first images can be collected when the wrist of the first hand is moving, and multiple consecutive second images can be collected when the fingers of the second hand are moving.
[0034] In the present disclosure, the first hand and the second hand can be the same or different hands. During the video collection process, it is possible to control all the fingers of the first hand to be straightened and in a palm-spread posture, and the postures of the wrist in various degrees of freedom (6 degrees of freedom) can be collected. During the video collection process, it is also possible to relatively fix the orientation of the first hand (i.e., fix the wrist degree-of-freedom data), and then various finger postures (i.e., finger postures with different finger degree-of-freedom data) can be collected. It can be understood that during the video collection process, it is possible to control the change or non-change of the hand orientation, and it is also possible to control the change or non-change of the postures of each finger.
[0035] Further, when the wrist of the first hand is in continuous movement, a camera is used to collect a first video, which includes multiple consecutive first images. When the fingers of the second hand are in movement, a camera is used to collect a second video, which includes multiple consecutive second images.
[0036] In the present disclosure, both the first video and the second video are collected for real hand movements. Moreover, the first images in the obtained first video are consecutive, and the second images in the second video are also consecutive.
[0037] Further, in the present disclosure, the wrist degree-of-freedom data and the finger degree-of-freedom data are collected separately, which can avoid the interference with the fingers due to special hand postures (such as making a fist) during simultaneous collection, thus making it impossible to collect the finger data.
[0038] S102. Determine the wrist degree-of-freedom data of the first image and the finger degree-of-freedom data of the second image.
[0039] In the present disclosure, each wrist in the first image corresponds to wrist degree-of-freedom data, where the wrist degree-of-freedom data includes six degree-of-freedom data, and the wrist state in the first image is represented by these six degree-of-freedom data. Each finger in the second image corresponds to finger degree-of-freedom data, where the finger degree-of-freedom data includes twenty degree-of-freedom data, and the finger state in the second image is represented by these twenty degree-of-freedom data.
[0040] S103. Combine the wrist degree-of-freedom data and the finger degree-of-freedom data to obtain multiple hand degree-of-freedom data.
[0041] In the present disclosure, referring to Figure 2 , the first video includes first images a1 to an, where n is a positive integer. The second video includes second images c1 to cm, where m is a positive integer. The time stamp directions of the multiple first images and the multiple second images are in the direction of the arrow t.
[0042] Further, each hand degree-of-freedom data includes one wrist degree-of-freedom data and one finger degree-of-freedom data. That is, the hand degree-of-freedom data includes: twenty-six degree-of-freedom data. These twenty-six degree-of-freedom data represent a hand posture.
[0043] In an alternative embodiment, the wrist degree-of-freedom data and the finger degree-of-freedom data are combined to obtain a plurality of hand degree-of-freedom data, including: when aligning the first video and the second video, combining a plurality of wrist degree-of-freedom data and a plurality of finger degree-of-freedom data to obtain a plurality of hand degree-of-freedom data, wherein the wrist degree-of-freedom data of the first image of the i-th frame in the first video is combined with the finger degree-of-freedom data of the second image of the j-th frame in the second video, and the wrist degree-of-freedom data of the first image of the (i + x)-th frame in the first video is combined with the finger degree-of-freedom data of the second image of the (j + x)-th frame in the second video, where both i and j are positive integers, and x sequentially takes values from 1 to y, and y is a positive integer greater than 1.
[0044] It can be understood that the first video and the second video can be from different videos, or from different or partially different parts of the same video.
[0045] Referring to Figure 2 , there are n wrist degree-of-freedom data and m finger degree-of-freedom data. Among them, if h hand degree-of-freedom data are required in the hand degree-of-freedom data set, then h consecutive wrist degree-of-freedom data need to be combined with h consecutive finger degree-of-freedom data, where h is a positive integer greater than 1 and less than n and m. There are n - h + 1 selection methods for selecting h consecutive wrist degree-of-freedom data from the n wrist degree-of-freedom data. There are m - h + 1 selection methods for selecting h consecutive finger degree-of-freedom data from the m finger degree-of-freedom data. Thus, the obtained hand degree-of-freedom data set has (n - h + 1) × (m - h + 1) elements, and (n - h + 1) × (m - h + 1) hand videos can be rendered. It can be seen that the present disclosure expands the sample size of the training samples. Exemplarily, in Figure 2 , when n is 5, m is 4, and h is 3, referring to Table 1 for the obtained hand degree-of-freedom data set, 6 hand degree-of-freedom data sets can be obtained, and 6 hand videos can be obtained after corresponding rendering.
[0046] Table 1
[0047]
[0048]
[0049] That is, after the first image of the i-th frame and the second image of the j-th frame are aligned, the subsequent first images and second images are aligned in sequence, and the wrist degree-of-freedom data and the finger degree-of-freedom data are combined in sequence to obtain a plurality of hand degree-of-freedom data. Among them, a continuous hand video can be obtained after rendering the plurality of hand degree-of-freedom data. Exemplarily, referring to Figure 2, after combination, p + 1 hand degree - of - freedom data are obtained, where p is a positive integer. Among them, the hand degree - of - freedom data is obtained by combining the wrist degree - of - freedom data bi and the finger degree - of - freedom data dj. The hand degree - of - freedom data e2 is obtained by combining the wrist degree - of - freedom data b(i + 1) and the finger degree - of - freedom data d(j + 1), and so on until the hand degree - of - freedom data e(p + 1) is obtained by combining the wrist degree - of - freedom data b(i + p) and the finger degree - of - freedom data d(j + p).
[0050] In the present disclosure, when aligning the first video and the second video, combining multiple wrist degree - of - freedom data and multiple finger degree - of - freedom data can make the obtained multiple hand degree - of - freedom data maintain the continuity in the real situation when the first video and the second video are originally collected. These continuous hand degree - of - freedom data are used to generate training samples, which can improve the accuracy of the hand motion recognition model for recognizing continuous hand motions.
[0051] S104. Render the set of hand degree - of - freedom data to obtain a hand video.
[0052] Among them, the set of hand degree - of - freedom data includes some or all of the hand degree - of - freedom data among multiple hand degree - of - freedom data.
[0053] In the present disclosure, the hand video can be directly used as a training sample for training the hand motion recognition model.
[0054] In an alternative embodiment, the set of hand degree - of - freedom data is composed of randomly selected partial hand degree - of - freedom data from multiple hand degree - of - freedom data. Refer to Figure 2 , for example, the set of hand degree - of - freedom data includes: hand degree - of - freedom data e1, hand degree - of - freedom data e3, hand degree - of - freedom data e5, hand degree - of - freedom data e7, etc. In another alternative embodiment, the set of hand degree - of - freedom data is composed of randomly selected partial continuous hand degree - of - freedom data from multiple hand degree - of - freedom data. Refer to Figure 2 , for example, the set of hand degree - of - freedom data includes: hand degree - of - freedom data e1, hand degree - of - freedom data e2, hand degree - of - freedom data e3, hand degree - of - freedom data e4, hand degree - of - freedom data e5, etc.
[0055] In the embodiments of the present disclosure, if the hand degree - of - freedom data in the set of hand degree - of - freedom data is continuous, multiple hand degree - of - freedom data in the set of hand degree - of - freedom data can form a hand motion, such as clicking, swiping or other motions.
[0056] In the present disclosure, before rendering the set of hand degree - of - freedom data to obtain a hand video, it further includes: extracting a preset number of continuous hand degree - of - freedom data from multiple hand degree - of - freedom data; determining that the preset number of continuous hand degree - of - freedom data constitutes the set of hand degree - of - freedom data.
[0057] In an embodiment of the present disclosure, if the hand degree-of-freedom data in the hand degree-of-freedom data set is continuous, multiple hand degree-of-freedom data in the hand degree-of-freedom data set can form a hand movement, such as a click, a slide, or other movements.
[0058] Among them, since both the wrist degree-of-freedom data and the finger degree-of-freedom data are collected according to real hand movements, the obtained hand degree-of-freedom data set can represent real and continuous hand degree-of-freedom data. The hand video generated using this hand degree-of-freedom data is more in line with reality, so the robustness of the trained hand movement recognition model can be improved.
[0059] In an alternative embodiment, rendering the hand degree-of-freedom data set to obtain a hand video includes: rendering the hand degree-of-freedom data set and preset hand data to obtain a hand video. The preset hand data includes: hand shape data, or at least one of lighting data, hand texture data, and color data and hand shape data.
[0060] In an embodiment of the present disclosure, a hand shape data set, a lighting data set, a hand texture data set, and a color data set can be pre-stored in a database. Then, random sampling is performed on the hand shape data set, the lighting data set, the hand texture data set, and the color data set to obtain the preset hand data. In an alternative embodiment, the preset hand data includes: hand shape data. In another alternative embodiment, the preset hand data includes: hand shape data and at least one of (lighting data, hand texture data, and color data).
[0061] After rendering the hand degree-of-freedom data set and these preset hand data together through a rendering engine, a hand video is obtained.
[0062] In an embodiment of the present disclosure, the hand video includes multiple consecutive hand images, such as hand image f1, hand image f2, hand image f3,..., hand image fz, where z is a positive integer. The hand degree-of-freedom data of the hand image fs is the s-th hand degree-of-freedom data in the hand degree-of-freedom data set, and s takes values from 1 to z in sequence.
[0063] Further, after rendering the hand degree-of-freedom data set and the preset hand data to obtain a hand video, it further includes: obtaining a preset background image, and synthesizing the hand video and the preset background image to obtain a training sample, where the training sample is used to train a hand movement recognition model.
[0064] Among them, the obtained hand video only has the area of the hand, and the background is blank. By pasting the hand in the hand video onto the background image, the obtained training sample can improve the robustness of the hand movement recognition model to the background.
[0065] In an embodiment of the present disclosure, a background image set can be pre-stored in a database, and then a background image is randomly sampled as a preset background image. This preset background image serves as the background of the hand video, and after synthesis, a hand video with a background is obtained. This hand video with a background is used as a training sample to train a hand gesture recognition model.
[0066] In the present disclosure, the first video corresponding to the wrist and the second video corresponding to the fingers are collected separately, which can avoid the situation where a large amount of self-occluded data cannot be collected when collected together, and further avoid the hand gesture recognition model trained from being unable to predict such self-occluded hand gestures (such as the state of the occluded fingers when making a fist). Using the present disclosure, a large amount of such simulated self-occluded data can be obtained for training the hand gesture recognition model.
[0067] Furthermore, the hand degree-of-freedom data set adopts continuous wrist degree-of-freedom data and continuous finger degree-of-freedom data, which can make the hand gestures in the obtained training samples have continuity and improve the robustness of the hand gesture recognition model for recognizing continuous hand gestures.
[0068] In addition, using this embodiment can save the high cost of collecting real data, efficiently obtain a large number of training samples, and the accuracy and robustness of the hand gesture recognition model trained with these training samples are relatively high.
[0069] Reference Figure 3 , is a schematic flowchart of the method for training a hand gesture recognition model provided by an embodiment of the present disclosure. As Figure 3 shown, the method for training a hand gesture recognition model specifically includes the following steps:
[0070] S301, Obtain training samples.
[0071] Among them, the training samples are determined based on the hand videos obtained by the above sample generation method. Among them, the training samples can be hand videos obtained by rendering the hand degree-of-freedom data set, or training samples obtained by acquiring a preset background image and synthesizing the hand video and the preset background image.
[0072] S302, Obtain the label data of the training samples.
[0073] The label data is the hand degree-of-freedom data set for generating the training samples.
[0074] In an embodiment of the present disclosure, the label data of the training samples is the degree-of-freedom data set for generating the training samples.
[0075] S303, Train a hand gesture recognition model according to the training samples and the label data to obtain a trained hand gesture recognition model.
[0076] Specifically, the training samples can be input into the hand gesture recognition model for recognition to obtain a preset degree-of-freedom data set. Then, the loss value of the preset degree-of-freedom data set relative to the label data is calculated, and the model parameters of the hand gesture recognition model are adjusted according to the loss value. After training for a preset number of rounds and / or when the model parameters of the hand gesture recognition model meet the convergence requirements, a trained hand gesture recognition model can be obtained.
[0077] In the embodiments of the present disclosure, the hand gesture recognition model can adopt a GAN (Generative Adversarial Network), or other network algorithms, which are not limited herein.
[0078] In this embodiment, since the training samples for training the hand gesture recognition model are generated by the above sample generation method, the hand gesture recognition model has good robustness in recognizing various hand gestures, can accurately recognize hand gestures, and also has good robustness to the background of hand gestures.
[0079] Reference Figure 4 , is a schematic flowchart of the hand gesture recognition method provided by the embodiments of the present disclosure. As Figure 4 shown, the hand gesture recognition method specifically includes the following steps:
[0080] S401, obtain the hand video of the user.
[0081] Among them, when the user uses the VR device, in order to interact with the VR device, the user makes some hand gestures, and the camera of the VR device collects the hand video according to the hand gestures.
[0082] S402, input the hand video into the hand gesture recognition model for recognition to obtain a hand degree-of-freedom data set.
[0083] Among them, the hand degree-of-freedom data set is used to represent the hand gestures of the user, and the hand gesture recognition model is obtained by the above hand gesture recognition model training method.
[0084] Specifically, the hand degree-of-freedom data set includes multiple hand degree-of-freedom data, and each hand degree-of-freedom data corresponds to a frame image in the hand video. After obtaining the hand degree-of-freedom data set, the hand gestures of the user can be determined.
[0085] In the present disclosure, since the hand degree-of-freedom data set is recognized by the hand gesture recognition model, the hand video can be accurately recognized.
[0086] Corresponding to the sample generation method in the above embodiments, Figure 5 , is a structural block diagram of the sample generation device 50 provided by the embodiments of the present disclosure. For the sake of convenience of description, only the parts related to the embodiments of the present disclosure are shown. As Figure 5As shown, the sample generation device 50 specifically includes: an acquisition unit 501, a determination unit 502, a combination unit 503, and a rendering unit 504, where:
[0087] The acquisition unit 501 is configured to acquire a first video and a second video. The first video includes multiple consecutive first images of a first hand, and the second video includes multiple consecutive second images of a second hand.
[0088] The determination unit 502 is configured to determine the wrist degree - of - freedom data of the first images and the finger degree - of - freedom data of the second images.
[0089] The combination unit 503 is configured to combine the wrist degree - of - freedom data and the finger degree - of - freedom data to obtain multiple hand degree - of - freedom data.
[0090] The rendering unit 504 is configured to render a set of hand degree - of - freedom data to obtain a hand video. The set of hand degree - of - freedom data includes some or all of the multiple hand degree - of - freedom data, and the hand video is used as a training sample for training a hand gesture recognition model.
[0091] In some embodiments, the combination unit 503 is specifically configured to, when aligning the first video and the second video, combine multiple wrist degree - of - freedom data and multiple finger degree - of - freedom data to obtain multiple hand degree - of - freedom data. Among them, the wrist degree - of - freedom data of the i - th frame of the first image in the first video is combined with the finger degree - of - freedom data of the j - th frame of the second image in the second video, and the wrist degree - of - freedom data of the (i + x)-th frame of the first image in the first video is combined with the finger degree - of - freedom data of the (j + x)-th frame of the second image in the second video. Both i and j are positive integers, and x takes values from 1 to y in sequence, where y is a positive integer greater than 1.
[0092] In some embodiments, the determination unit 502: is further configured to, before rendering the set of hand degree - of - freedom data to obtain a hand video, extract a preset number of consecutive hand degree - of - freedom data from the multiple hand degree - of - freedom data; and determine that the preset number of consecutive hand degree - of - freedom data constitutes the set of hand degree - of - freedom data.
[0093] In some embodiments, rendering the set of hand degree - of - freedom data to obtain a hand video includes: rendering the set of hand degree - of - freedom data and preset hand data to obtain a hand video. The preset hand data includes: hand - shape data, or at least one of illumination data, hand texture data, color data, and hand - shape data.
[0094] In some embodiments, it further includes a synthesis module, which is configured to, after rendering the set of hand degree - of - freedom data and preset hand data to obtain a hand video, acquire a preset background image, and synthesize the hand video and the preset background image to obtain a training sample, and the training sample is used to train a hand gesture recognition model.
[0095] The sample generation device provided in this embodiment can be used to implement the technical solutions of the embodiments of the above sample generation method. The implementation principle and technical effects are similar, and will not be elaborated here in this embodiment.
[0096] Corresponding to the hand gesture recognition model training method in the above embodiment, a hand gesture recognition model training device (not shown) is provided, which specifically includes: an acquisition unit and a training unit, where:
[0097] The acquisition unit is used to acquire training samples, where the training samples are obtained according to the above sample generation method; and acquire the label data of the training samples, where the label data is the set of hand freedom degree data for generating the training samples;
[0098] The training unit is used to train the hand gesture recognition model according to the training samples and the label data to obtain a trained hand gesture recognition model.
[0099] The hand gesture recognition model training device provided in this embodiment can be used to implement the technical solutions of the embodiments of the above hand gesture recognition model training method. The implementation principle and technical effects are similar, and will not be elaborated here in this embodiment.
[0100] Corresponding to the hand gesture recognition method in the above embodiment, a hand gesture recognition device (not shown) is provided, which specifically includes: an acquisition unit and a recognition unit, where:
[0101] The acquisition unit is used to acquire the hand video of the user;
[0102] The recognition unit is used to input the hand video into the hand gesture recognition model for recognition to obtain a set of hand freedom degree data, where the set of hand freedom degree data is used to represent the hand gesture of the user, and the hand gesture recognition model is obtained according to the above hand gesture recognition model training method.
[0103] The hand gesture recognition device provided in this embodiment can be used to implement the technical solutions of the embodiments of the above hand gesture recognition method. The implementation principle and technical effects are similar, and will not be elaborated here in this embodiment.
[0104] Reference Figure 6, which shows a schematic structural diagram of an electronic device 60 suitable for implementing the embodiments of the present disclosure. The electronic device 60 can be a terminal device or a server. Among them, the terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0105] As Figure 6 shown, the electronic device 60 can include a processing device (such as a central processing unit, a graphics processing unit, etc.) 61, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 62 or the program loaded from the storage device 68 into the random access memory (RAM) 63. In the RAM 63, various programs and data required for the operation of the electronic device 60 are also stored. The processing device 61, the ROM 62, and the RAM 63 are connected to each other through a bus 64. The input / output (I / O) interface 65 is also connected to the bus 64.
[0106] Generally, the following devices can be connected to the I / O interface 65: an input device 66 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 67 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 68 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 69. The communication device 69 can allow the electronic device 60 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device 60 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.
[0107] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 69, or installed from a storage device 68, or installed from a ROM 62. When the computer program is executed by a processing device 61, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.
[0108] It should be noted that the above computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0109] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.
[0110] The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to perform the methods shown in the above embodiments.
[0111] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a unit, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0113] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".
[0114] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0115] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read-Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM) or flash memory, optical fibers, portable Compact Disc Read-Only Memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0116] In a first aspect, according to one or more embodiments of the present disclosure, there is provided a method for generating a sample, including: obtaining a first video and a second video, where the first video includes multiple consecutive first images of a first hand, and the second video includes multiple consecutive second images of a second hand; determining wrist degree-of-freedom data of the first images and finger degree-of-freedom data of the second images; combining the wrist degree-of-freedom data and the finger degree-of-freedom data to obtain multiple hand degree-of-freedom data; rendering a set of hand degree-of-freedom data to obtain a hand video, where the set of hand degree-of-freedom data includes some or all of the multiple hand degree-of-freedom data, and the hand video is used as a training sample for training a hand gesture recognition model.
[0117] According to one or more embodiments of the present disclosure, wrist degree-of-freedom data and finger degree-of-freedom data are combined to obtain a plurality of hand degree-of-freedom data, including: when aligning a first video and a second video, combining a plurality of wrist degree-of-freedom data and a plurality of finger degree-of-freedom data to obtain a plurality of hand degree-of-freedom data, wherein the wrist degree-of-freedom data of the first image of the i-th frame in the first video is combined with the finger degree-of-freedom data of the second image of the j-th frame in the second video, and the wrist degree-of-freedom data of the first image of the (i + x)-th frame in the first video is combined with the finger degree-of-freedom data of the second image of the (j + x)-th frame in the second video, where both i and j are positive integers, and x sequentially takes values from 1 to y, and y is a positive integer greater than 1.
[0118] According to one or more embodiments of the present disclosure, before rendering a set of hand degree-of-freedom data to obtain a hand video, it further includes: extracting a preset number of consecutive hand degree-of-freedom data from the plurality of hand degree-of-freedom data; determining that the preset number of consecutive hand degree-of-freedom data constitutes a set of hand degree-of-freedom data.
[0119] According to one or more embodiments of the present disclosure, rendering a set of hand degree-of-freedom data to obtain a hand video includes: rendering a set of hand degree-of-freedom data and preset hand data to obtain a hand video, where the preset hand data includes: hand shape data or at least one of illumination data, hand texture data, color data and hand shape data.
[0120] According to one or more embodiments of the present disclosure, after rendering a set of hand degree-of-freedom data and preset hand data to obtain a hand video, it further includes: obtaining a preset background image, and synthesizing the hand video and the preset background image to obtain a training sample, and the training sample is used to train a hand action recognition model.
[0121] In a second aspect, according to one or more embodiments of the present disclosure, a method for training a hand action recognition model is provided, including:
[0122] Obtaining a training sample, where the training sample is obtained according to the sample generation method described above;
[0123] Obtaining label data of the training sample, where the label data is the set of hand degree-of-freedom data for generating the training sample;
[0124] Training a hand action recognition model according to the training sample and the label data to obtain a trained hand action recognition model.
[0125] In a third aspect, according to one or more embodiments of the present disclosure, a method for recognizing a hand action is provided, including:
[0126] Obtaining a hand video of a user;
[0127] Input the hand video into the hand gesture recognition model for recognition to obtain a set of hand freedom data, which is used to represent the user's hand gestures. The hand gesture recognition model is obtained according to the aforementioned hand gesture recognition model training method.
[0128] Fourthly, according to one or more embodiments of the present disclosure, a sample generation device is provided, including:
[0129] An acquisition unit for acquiring a first video and a second video. The first video includes multiple consecutive first images of a first hand, and the second video includes multiple consecutive second images of a second hand;
[0130] A determination unit for determining the wrist freedom data of the first image and the finger freedom data of the second image;
[0131] A combination unit for combining the wrist freedom data and the finger freedom data to obtain a plurality of hand freedom data;
[0132] A rendering unit for rendering the set of hand freedom data to obtain a hand video. The set of hand freedom data includes some or all of the plurality of hand freedom data, and the hand video is used as a training sample for training the hand gesture recognition model.
[0133] Fifthly, according to one or more embodiments of the present disclosure, an electronic device is provided, including: at least one processor and a memory;
[0134] The memory stores computer execution instructions;
[0135] At least one processor executes the computer execution instructions stored in the memory, so that at least one processor executes the methods provided in the first aspect, the second aspect, and / or the third aspect as above.
[0136] Sixthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided. Computer execution instructions are stored in the computer-readable storage medium. When the processor executes the computer execution instructions, the methods provided in the first aspect, the second aspect, and / or the third aspect as above are implemented.
[0137] Seventhly, according to one or more embodiments of the present disclosure, a computer program product is provided. The computer program product includes computer execution instructions. When the processor executes the computer execution instructions, the methods provided in the first aspect, the second aspect, and / or the third aspect as above are implemented.
[0138] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0139] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0140] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms for implementing the claims.
Claims
1. A method for generating samples, comprising: Obtaining a first video and a second video, where the first video includes multiple consecutive first images of a first hand, and the second video includes multiple consecutive second images of a second hand; Determining the wrist degree - of - freedom data of the first images and the finger degree - of - freedom data of the second images; Combining the wrist degree - of - freedom data and the finger degree - of - freedom data to obtain multiple hand degree - of - freedom data; Rendering a set of hand degree - of - freedom data to obtain a hand video, where the set of hand degree - of - freedom data includes some or all of the multiple hand degree - of - freedom data.
2. The method for generating samples according to claim 1, where the combining the wrist degree - of - freedom data and the finger degree - of - freedom data to obtain multiple hand degree - of - freedom data includes: When aligning the first video and the second video, combining multiple wrist degree - of - freedom data and multiple finger degree - of - freedom data to obtain multiple hand degree - of - freedom data, where the wrist degree - of - freedom data of the i - th frame of the first image in the first video is combined with the finger degree - of - freedom data of the j - th frame of the second image in the second video, and the wrist degree - of - freedom data of the (i + x)-th frame of the first image in the first video is combined with the finger degree - of - freedom data of the (j + x)-th frame of the second image in the second video, where both i and j are positive integers, and x takes values from 1 to y in sequence, and y is a positive integer greater than 1.
3. The method for generating samples according to claim 1 or 2, before rendering the set of hand degree - of - freedom data to obtain a hand video, further comprising: Extracting a preset number of consecutive hand degree - of - freedom data from the multiple hand degree - of - freedom data; Determining that the preset number of consecutive hand degree - of - freedom data constitutes the set of hand degree - of - freedom data.
4. The sample generation method according to claim 1 or 2, wherein rendering the hand degree-of-freedom data set to obtain a hand video includes: Rendering the set of hand degree - of - freedom data and preset hand data to obtain the hand video, where the preset hand data includes: hand - shape data or at least one of lighting data, hand texture data, color data and the hand - shape data.
5. The method for generating samples according to claim 4, after rendering the set of hand degree - of - freedom data and preset hand data to obtain the hand video, further comprising: Obtaining a preset background image, and synthesizing the hand video and the preset background image to obtain a training sample, where the training sample is used to train the hand - action recognition model.
6. A method for training a hand - action recognition model, comprising: Obtaining a training sample, where the training sample is determined from a hand video obtained by the method for generating samples according to any one of claims 1 to 5; Obtaining label data of the training sample, where the label data is the set of hand degree - of - freedom data for generating the training sample; Training a hand - action recognition model according to the training sample and the label data to obtain a trained hand - action recognition model.
7. A method for recognizing hand actions, comprising: Obtaining a hand video of a user; Input the hand video into a hand motion recognition model for recognition to obtain a set of hand freedom data, where the set of hand freedom data is used to represent the hand motion of the user, and the hand motion recognition model is obtained according to the hand motion recognition model training method described in claim 6.
8. A sample generation device, comprising: An acquisition unit configured to acquire a first video and a second video, where the first video includes multiple consecutive first images of a first hand, and the second video includes multiple consecutive second images of a second hand; A determination unit configured to determine the wrist freedom data of the first image and the finger freedom data of the second image; A combination unit configured to combine the wrist freedom data and the finger freedom data to obtain a plurality of hand freedom data; A rendering unit configured to render a set of hand freedom data to obtain a hand video, where the set of hand freedom data includes some or all of the plurality of hand freedom data, and the hand video is used as a training sample for training a hand motion recognition model.
9. An electronic device, comprising: At least one processor and a memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the method described in any one of claims 1 to 7.
10. A computer-readable storage medium, in which computer-executable instructions are stored, and when a processor executes the computer-executable instructions, the method described in any one of claims 1 to 7 is implemented.