Model training, virtual human hand action control method and device, and storage medium

By training a neural network model and using a single camera to identify key points of the hand, the problems of precision and cost in virtual human hand control have been solved, achieving efficient virtual human hand motion control.

CN117036863BActive Publication Date: 2026-08-25杭州米络星科技(集团)有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311058462.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-21
Publication Date
2026-08-25
Estimated Expiration
2043-08-21

AI Technical Summary

Technical Problem

Existing virtual human hand control methods suffer from insufficient control precision, high cost, high environmental requirements, and large data processing volume.

Method used

By acquiring a dataset of human hand movements, a dataset of hand joint rotation degrees and a dataset of key points are generated. A neural network model is trained, and a single camera is used to acquire hand images and identify key points, thereby achieving precise control of the virtual human hand.

Benefits of technology

It achieves precise control over the virtual human's hand, reduces environmental requirements and data processing volume, reduces reliance on camera matrices, and lowers costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117036863B_ABST
    Figure CN117036863B_ABST
Patent Text Reader

Abstract

The application discloses a model training, virtual human hand action control method and device and a storage medium, wherein the training method comprises the following steps: acquiring a human hand action data set collected by a preset sensor system, and generating a hand joint node rotation degree data set based on the human hand action data set; generating a hand image set based on the hand joint node rotation degree data set, and identifying hand images in the hand image set to acquire a hand key point data set; generating a training data set based on the hand key point data set and the hand joint node rotation degree data set, training a preset neural network model based on the training data set, and acquiring a hand action recognition model. The application can accurately control the hand posture of a virtual human, and solves the problem of low recognition accuracy of an existing human hand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hand recognition technology, and more particularly to model training, virtual human hand motion control methods and devices, and storage media. Background Technology

[0002] In 3D virtual live streaming, the implementation of virtual human driving is often used, and hand driving is an important part of virtual human driving.

[0003] Existing methods for controlling virtual human hands mainly fall into two categories. One involves using wearable sensors to collect data on hand joint points, then controlling the virtual hand based on this sensor data. However, this method requires wearable sensors, and the limited number of sensor points results in imprecise control. The other method uses a camera matrix to capture hand images from various angles, performs image recognition, and controls the virtual hand based on the recognition results. However, this method requires deploying a camera matrix in a specific space, leading to high costs and a large amount of data processing. Furthermore, using cameras with ranging capabilities is expensive and requires a suitable operating environment, making it cumbersome. Using a single camera, current technology results in relatively coarse hand movement control, lacking fine precision and making the virtual hand movements appear artificial and providing a poor user experience. Summary of the Invention

[0004] The purpose of this application is to provide a model training, virtual human hand motion control method, device, and storage medium to solve the problems of insufficient control precision in current methods of virtual human hand control using sensor data, and the high environmental requirements, high cost, and large data processing volume of methods of virtual human hand control using camera matrices.

[0005] Firstly, this application provides a method for training a hand motion recognition model, including:

[0006] Acquire a human hand movement dataset collected by a preset sensor system, and generate a hand joint rotation dataset based on the human hand movement dataset;

[0007] A set of hand images is generated based on the hand joint rotation degree dataset, and the hand images in the set of hand images are identified to obtain a set of hand key points.

[0008] A training dataset is generated based on the hand key point dataset and the hand joint rotation dataset. A preset neural network model is trained based on the training dataset to obtain a hand motion recognition model.

[0009] In one embodiment of this application, the preset neural network model includes a sequentially connected sequence of up-dimensional feature extraction, a hand pose feature acquisition module, and a fully connected layer;

[0010] The hand pose feature acquisition module includes a hybrid feature layer, a first-level fully connected layer, and a second-level fully connected layer connected in sequence.

[0011] In one embodiment of this application, the preset sensor system is a Neuron motion capture system.

[0012] In one embodiment of this application, the hand image is identified using a mediapipe to obtain a dataset of key hand points.

[0013] Secondly, this application provides a training device for a hand action recognition model, including a rotation degree dataset acquisition module, a key point dataset acquisition module, and a training module;

[0014] The rotation degree dataset acquisition module is used to acquire the human hand movement dataset collected by the preset sensor system, and generate a hand joint rotation degree dataset based on the human hand movement dataset.

[0015] The key point dataset acquisition module is used to generate a hand image set based on the hand joint rotation degree dataset, and to identify the hand images in the hand image set to obtain a hand key point dataset.

[0016] The training module is used to generate a training dataset based on the hand key point dataset and the hand joint rotation dataset, and to train a preset neural network model based on the training dataset to obtain a hand action recognition model.

[0017] Thirdly, this application provides a method for controlling the movement of a virtual human hand, including:

[0018] Acquire an image of a human hand, and obtain key point data of the hand based on the image;

[0019] The key hand data is transmitted to the hand motion recognition model, and the hand motion recognition model outputs the rotational data of the hand joints.

[0020] The virtual human hand movements are controlled based on the rotational data of the hand joints;

[0021] The hand motion recognition model is the hand motion recognition model obtained by the hand motion recognition model training method.

[0022] In one embodiment of this application, an image of a human hand is acquired using a single camera.

[0023] In one embodiment of this application, key points of the hand in the human hand image are identified using a mediapipe to obtain key point data of the hand.

[0024] Fourthly, this application provides a virtual human hand motion control device, including a key point data acquisition module, a rotational data acquisition module, and a control module;

[0025] The key point data acquisition module is used to acquire human hand images and acquire key point data of the hand based on the human hand images;

[0026] The rotational data acquisition module is used to transmit the key hand points data to the hand motion recognition model, and the hand motion recognition model outputs the rotational data of the hand joints.

[0027] The control module is used to control the virtual human's hand movements based on the hand joint rotation data;

[0028] The hand motion recognition model is the hand motion recognition model obtained by the hand motion recognition model training method.

[0029] Fifthly, this application provides a storage medium on which a computer program is stored, which, when executed by a processor, implements the hand motion recognition model training method described above.

[0030] Alternatively, the virtual human hand motion control method may be implemented when the program is executed by the processor.

[0031] Compared with the prior art, one or more embodiments of the above solutions may have the following advantages or beneficial effects:

[0032] The hand movement recognition model training method provided in this invention collects human hand data through sensors to obtain a hand joint rotation degree dataset, which is used as the label in the original training data. Based on the hand joint rotation degree data, a hand image is obtained, and hand key points are identified through hand key point recognition, which is also used as the original training data. The data trained through this training dataset can accurately obtain the hand joint rotation degree based on the hand key point data, thereby enabling accurate control of the virtual human's hand posture and solving the problem of low accuracy in existing human hand recognition methods.

[0033] The virtual human hand motion control method provided in this embodiment of the invention employs the aforementioned hand motion recognition model. Due to its high accuracy in recognizing hand rotation, this control method acquires images of the human hand using only a single camera and identifies key points in the image. This allows the hand motion recognition model to determine the joint rotation of the human hand, thereby achieving precise control of the virtual human hand. Furthermore, this method has low environmental requirements, eliminates the need for camera matrix layout, and minimizes the need for extensive data processing.

[0034] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the description, claims, and drawings. Attached Figure Description

[0035] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0036] Figure 1 The diagram shows a flowchart of the hand motion recognition model training method of the present invention.

[0037] Figure 2 The diagram shown is a structural schematic of the hand motion recognition model training device of the present invention.

[0038] Figure 3 The diagram shown is a flowchart illustrating the virtual human hand motion control method of the present invention.

[0039] Figure 4 The diagram shown is a structural schematic of the virtual human hand motion control device of the present invention. Detailed Implementation

[0040] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.

[0041] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the drawings only show the components related to this application and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0042] The following embodiments of this application provide a video super-resolution reconstruction method and apparatus, storage medium and terminal, which solve the problems of inconsistent timing of the generated results and insufficient repair of high-frequency detail information in current video super-resolution methods.

[0043] The following will describe in detail, with reference to the accompanying drawings, the principles and implementation methods of a hand motion recognition model training method, apparatus, storage medium and terminal of this embodiment, so that those skilled in the art can understand the hand motion recognition model training method, apparatus, storage medium and terminal of this embodiment without creative effort.

[0044] like Figure 1 As shown in the figure, this embodiment provides a method for training a hand motion recognition model, which includes the following steps.

[0045] Step S101: Obtain the human hand movement dataset collected by the preset sensor system, and generate a hand joint rotation dataset based on the human hand movement dataset.

[0046] Specifically, real human hand data is acquired through a pre-set sensor system. During data acquisition, the human body needs to make various hand gestures, and a certain amount of data needs to be acquired for each gesture to maximize data coverage and result accuracy. Then, the pre-set sensor system software can be used to generate corresponding joint rotation data based on each set of acquired human hand movement data. All virtual human joint rotation data form a hand joint rotation dataset. Alternatively, other software can be selected to convert human hand movement data into joint rotation data. Preferably, the Neuron motion capture system can be selected as the pre-set sensor system.

[0047] Step S102: Generate a hand image set based on the hand joint rotation degree dataset, and identify the hand images in the hand image set to obtain the hand key point dataset.

[0048] Specifically, a corresponding hand image is generated based on the hand joint rotation data in the hand joint rotation dataset. Then, the hand image is recognized to obtain a hand keypoint dataset. This method acquires the hand keypoint data corresponding to all hand joint rotation data in the hand joint rotation dataset. All hand keypoint data form a hand keypoint dataset.

[0049] The hand keypoint data includes 21 keypoints for the left hand and 21 keypoints for the left hand. Furthermore, hand keypoint data can be obtained by recognizing hand images using MediaPipe.

[0050] Step S103: Generate a training dataset based on the hand key point dataset and the hand joint rotation dataset, and train a preset neural network model based on the training dataset to obtain a hand motion recognition model.

[0051] Specifically, the hand keypoint data in the hand keypoint dataset is correlated with the hand joint rotation data in the hand joint rotation dataset. The hand keypoint data is used as the training data in the training dataset, and the hand joint rotation data is used as the label in the training dataset to form the training dataset. Then, a preset neural network model is trained based on the training dataset, and the value of the preset neural network model's loss function is calculated. When the value of the loss function is stable, the relevant model is saved, and a neural network model for virtual human hand movement control, i.e., a hand movement recognition model, can be obtained.

[0052] The preset neural network model includes a sequentially connected sequence of up-dimensional feature extraction sequences, a hand pose feature acquisition module, and a fully connected layer. The hand pose feature acquisition module includes a sequentially connected hybrid feature layer, a first-level fully connected layer, and a second-level fully connected layer.

[0053] The hand pose feature acquisition module is further implemented using MPalmNet, which consists of three parts: a hybrid feature layer, a first-level fully connected layer, and a second-level fully connected layer. The hybrid feature layer primarily processes the input features through convolutional operations and activation functions, and uses max pooling layers to reduce spatial dimensionality, making the model more efficient and robust. The first-level and second-level fully connected layers map the input features to different dimensional spaces and use activation functions to ensure the model's non-linear performance. For example, a tensor x can be set as input with a shape of (batch_size, 85*6, 84), where 6 represents 5 fingers and 1 palm, and 84 represents the dimensions of the hand's keypoints. The hand pose feature acquisition module first reshapes x to (batch_size, 85, 84*6), then extracts features through the hybrid feature layer, reshapes the output to (batch_size, 1660), and passes it through the first-level and second-level fully connected layers to finally compute a tensor of (batch_size, 256) representing the hand pose features.

[0054] The default neural network model is MPHandNet, primarily used for feature extraction and prediction. Specifically, it predicts hand pose (including the pose of the palm and each finger) based on input hand keypoint information. The default neural network model includes two upscaling feature extraction sequences (e.g., expand1Seq and expand2Seq), followed by a hand pose feature acquisition module. This module uses two fully connected layers (level1FC and level2WithHandFC), and then a final fully connected layer (lastFC). For example, a tensor x can be used as input, representing the input keypoint data. First, local feature extraction is performed on the keypoint information of each finger and palm (10 points per finger, 10 points per palm). Then, mpPalmNet is used to calculate the overall hand pose. Finally, the finger pose information and palm pose information are combined, merging these two features, and the rotation data is calculated using a fully connected layer.

[0055] The hand movement recognition model training method provided in this invention collects human hand data through sensors to obtain a hand joint rotation degree dataset, which is used as the label in the original training data; based on the hand joint rotation degree data, a hand image is obtained, and hand key points are identified through hand key point recognition, which is also used as the original training data; the data trained through this training dataset can accurately obtain the hand joint rotation degree based on the hand key point data, thereby accurately controlling the hand posture of the virtual human, solving the problem of low accuracy in existing human hand recognition.

[0056] like Figure 2 As shown, this embodiment provides a hand motion recognition model training device, including a rotation degree dataset acquisition module, a key point dataset acquisition module, and a training module.

[0057] The rotation degree dataset acquisition module is used to acquire the human hand movement dataset collected by the preset sensor system, and generate a hand joint rotation degree dataset based on the human hand movement dataset.

[0058] The key point dataset acquisition module is used to generate a hand image set based on the hand joint rotation degree dataset, and to identify the hand images in the hand image set to obtain the hand key point dataset;

[0059] The training module is used to generate a training dataset based on the hand key point dataset and the hand joint rotation dataset, and to train a preset neural network model based on the training dataset to obtain a hand motion recognition model.

[0060] The hand movement recognition model training device provided in this invention collects human hand data through sensors to obtain a hand joint rotation degree dataset, which is used as a label in the original training data; it acquires hand images based on the hand joint rotation degree data, and identifies hand key points through hand key point recognition, which is also used as the original training data; the data trained through this training dataset can accurately obtain the hand joint rotation degree based on the hand key point data, thereby accurately controlling the hand posture of the virtual human, solving the problem of low accuracy in existing human hand recognition.

[0061] like Figure 3 As shown in the figure, this embodiment provides a virtual human hand motion control method, which includes the following steps.

[0062] Step S301: Obtain an image of a human hand and obtain key point data of the hand based on the image.

[0063] Specifically, a real human hand image is acquired using a camera or video camera. Since the hand movement recognition model acquired in this embodiment is obtained using the method described in the previous embodiment, this embodiment can acquire a human hand image using a single camera. Then, the key points of the hand in the human hand image are identified using MediaPipe to obtain hand key point data.

[0064] Step S302: Transmit the key hand data to the hand motion recognition model, and the hand motion recognition model outputs the rotational data of the hand joints.

[0065] Specifically, the hand gesture recognition model is obtained by the method described in the above embodiments.

[0066] Step S303: Control the virtual human's hand movements based on the rotational data of the hand joints.

[0067] Specifically, a data transmission protocol is built between the application and the UE program (virtual human rendering program). Then, the rotational data of the hand joints is transmitted to the UE program, generating a DLL program for the Unreal Engine to call, thereby enabling control of the virtual human's hand movements based on the rotational data of the hand joints.

[0068] The virtual human hand motion control method provided in this embodiment of the invention employs the aforementioned hand motion recognition model. Due to its high accuracy in recognizing hand rotation, this control method acquires images of the human hand using only a single camera and identifies key points in the image. This allows the hand motion recognition model to determine the joint rotation of the human hand, thereby achieving precise control of the virtual human hand. Furthermore, this method has low environmental requirements, eliminates the need for camera matrix layout, and minimizes the need for extensive data processing.

[0069] like Figure 4 As shown, this embodiment provides a virtual human hand motion control device, including a key point data acquisition module, a rotational data acquisition module, and a control module.

[0070] The key point data acquisition module is used to acquire human hand images and acquire key point data of the hand based on the human hand images.

[0071] The rotational data acquisition module is used to transmit the key hand points to the hand motion recognition model, and the hand motion recognition model outputs the rotational data of the hand joints.

[0072] The control module is used to control the virtual human's hand movements based on the hand joint rotation data.

[0073] The hand gesture recognition model is obtained by the method described in the above embodiments.

[0074] The virtual human hand motion control device provided in this embodiment of the invention uses the aforementioned hand motion recognition model. Because of its high accuracy in recognizing hand rotation, this control device acquires images of the human hand using only a single camera and identifies key points in the image. This allows it to obtain the joint rotation of the human hand through the hand motion recognition model, thereby achieving precise control of the virtual human hand. Furthermore, this device has low environmental requirements and does not require a camera matrix layout or extensive data processing.

[0075] This application also provides a computer-readable storage medium. Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing a processor. The program can be stored in a computer-readable storage medium, which is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof. The storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state drive (SSD)).

Claims

1. A method for training a hand motion recognition model, comprising: Acquire a human hand movement dataset collected by a preset sensor system, and generate a hand joint rotation dataset based on the human hand movement dataset; A hand image set is generated based on the hand joint rotation degree dataset, and the hand images in the hand image set are identified to obtain a hand key point dataset. The hand images are acquired through a single camera of a camera. A training dataset is generated based on the hand key point dataset and the hand joint rotation dataset. A preset neural network model is trained based on the training dataset to obtain a hand motion recognition model. The preset neural network model includes a sequentially connected sequence of up-dimensional feature extraction, a hand pose feature acquisition module, and a fully connected layer. The hand pose feature acquisition module includes a sequentially connected hybrid feature layer, a first-level fully connected layer, and a second-level fully connected layer. The hybrid feature layer processes the input features through convolution operations and activation functions, and uses a max pooling layer to reduce the spatial dimension. The first-level and second-level fully connected layers map the input features to different dimensional spaces and use activation functions to ensure the nonlinear performance of the model.

2. The training method according to claim 1, characterized in that, The preset sensor system is the Neuron motion capture system.

3. The training method according to claim 1, characterized in that, The hand image is identified using MediaPipe to obtain key hand point data.

4. A hand motion recognition model training device, characterized in that, It includes a rotation degree dataset acquisition module, a key point dataset acquisition module, and a training module; The rotation degree dataset acquisition module is used to acquire the human hand movement dataset collected by the preset sensor system, and generate a hand joint rotation degree dataset based on the human hand movement dataset. The key point dataset acquisition module is used to generate a hand image set based on the hand joint rotation degree dataset, and to identify the hand images in the hand image set to obtain a hand key point dataset, wherein the hand images are acquired through a single camera of a camera. The training module is used to generate a training dataset based on the hand key point dataset and the hand joint rotation degree dataset, and to train a preset neural network model based on the training dataset to obtain a hand action recognition model. The preset neural network model includes a sequentially connected sequence of up-dimensional feature extraction, a hand pose feature acquisition module, and a fully connected layer. The hand pose feature acquisition module includes a sequentially connected hybrid feature layer, a first-level fully connected layer, and a second-level fully connected layer. The hybrid feature layer processes the input features through convolution operations and activation functions, and uses a max pooling layer to reduce the spatial dimension. The first-level and second-level fully connected layers map the input features to different dimensional spaces and use activation functions to ensure the nonlinear performance of the model.

5. A method for controlling the movement of a virtual human hand, comprising: Acquire an image of a human hand, and obtain key point data of the hand based on the image; The key hand data is transmitted to the hand motion recognition model, and the hand motion recognition model outputs the rotational data of the hand joints. The virtual human hand movements are controlled based on the rotational data of the hand joints; The hand motion recognition model is the hand motion recognition model obtained by the hand motion recognition model training method according to any one of claims 1-3.

6. The control method according to claim 5, characterized in that, Images of a human hand are captured using a single camera lens.

7. The control method according to claim 5, characterized in that, The key points of the hand in the human hand image are identified using MediaPipe to obtain key hand point data.

8. A virtual human hand motion control device, characterized in that, It includes a key point data acquisition module, a rotational degree data acquisition module, and a control module; The key point data acquisition module is used to acquire human hand images and acquire key point data of the hand based on the human hand images; The rotational data acquisition module is used to transmit the key hand points data to the hand motion recognition model, and the hand motion recognition model outputs the rotational data of the hand joints. The control module is used to control the virtual human's hand movements based on the hand joint rotation data; The hand motion recognition model is the hand motion recognition model obtained by the hand motion recognition model training method according to any one of claims 1-3.

9. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the hand motion recognition model training method as described in any one of claims 1 to 3; Or, when the program is executed by the processor, it implements the virtual human hand motion control method as described in any one of claims 5 to 7.

Citation Information

Patent Citations

  • Method and system for generating neural network model capable of being used for hand action recognition

    CN112215112A