Gesture posture estimation method based on multi-view and semi-supervised learning

By using multi-view and semi-supervised learning methods in 3D gesture pose estimation, unsupervised and supervised learning models are combined, and the problems of labeling difficulties and real-time requirements are solved, and a low-cost and high-precision gesture 3D gesture estimation is achieved.

CN120236327APending Publication Date: 2025-07-01JIANGSU KAIBO SOFTWARE DEV CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510300197.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Vision-based 3D gesture estimation faces challenges such as occlusion problems, perspective changes, large-scale labeling data requirements and real-time requirements, especially the difficulty in labeling 3D gesture pose data.

Method used

A gesture estimation method based on multi-view and semi-supervised learning is adopted. By integrating the unsupervised learning model and supervised learning model, a semi-supervised learning model is constructed, and unsupervised training is used for unsupervised training with multiple perspectives, and a small number of labeled samples are appropriately used to form a low-cost and high-precision gesture 3D gesture estimation system.

Benefits of technology

In the case of only a small number of label samples, efficiently utilizing a large number of label-free samples is achieved to achieve high-precision gesture estimation, reducing the demand for label samples, and on the public gesture estimation dataset Panoptic, the deviation of 3D estimation is less than 10mm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236327A_ABST
    Figure CN120236327A_ABST
Patent Text Reader

Abstract

The invention discloses a gesture posture estimation method based on multiple visual angles and semi-supervised learning, and belongs to the technical field of man-machine interaction, the gesture posture estimation method comprises the following steps: collecting gesture images through multiple visual angles, constructing a label-free data set, labeling a small amount of gesture posture data to form a label data set with hand key points; then randomly selecting one from a plurality of visual angles as input, forming a hidden code through an encoder constructed by a neural network, mapping the hidden code into a gesture image of another visual angle through a decoder to realize consistency constraint of different visual angles, mapping the hidden code into 2D key point coordinates through a projection converter, and finally obtaining a gesture image of the other visual angle; the method comprises the following steps of: randomly inputting a multi-view label-free sample or a small number of label samples for training under the combined action of two constraint mechanisms, and generating 3D key point coordinate estimation of a gesture posture only by inputting a gesture image into an encoder in practical application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer interaction, and particularly to a gesture pose estimation method based on multi-view and semi-supervised learning. Background Technique

[0002] With the rapid development of fields such as human-computer interaction, virtual reality, and augmented reality, the demand for gesture recognition and understanding is increasing day by day. Vision-based 3D gesture pose estimation aims to recover the three-dimensional spatial positions of hand joint points from images or videos, and is one of the key technologies for realizing natural and intuitive human-computer interaction;

[0003] In recent years, the rise of deep learning technology has brought new opportunities to vision-based 3D gesture pose estimation. Methods based on convolutional neural networks (CNNs) can learn hand feature representations from a large amount of data and have made significant progress;

[0004] However, due to factors such as complex hand structures, severe occlusions, and large perspective changes, vision-based 3D gesture pose estimation still faces many challenges. Occlusion problem: The hand often has self-occlusions or occlusions with other objects during movement, resulting in the loss of information in some hand regions and making it difficult to accurately estimate the positions of occluded joint points; Perspective change: The hand poses vary greatly under different perspectives, and the model needs to have strong generalization ability to adapt to various perspective changes; Data shortage: Obtaining large-scale and high-quality 3D gesture pose annotation data is costly, which limits the performance improvement of deep learning models; Real-time requirement: Many application scenarios have high real-time requirements for 3D gesture pose estimation, and it is necessary to improve the calculation efficiency while ensuring accuracy; Complex background interference: Complex backgrounds will increase image noise, interfere with hand region detection and feature extraction, and affect the pose estimation accuracy;

[0005] In summary, the problem of difficult annotation of 3D gesture pose data is particularly prominent. Although the current deep learning has strong capabilities, a large number of labeled training samples need to be provided, which is very difficult for the application of 3D gesture pose estimation. Therefore, a gesture pose estimation method based on multi-view and semi-supervised learning is proposed. Summary of the Invention

[0006] The purpose of the present invention is to provide a gesture pose estimation method based on multi-view and semi-supervised learning to solve the problems raised in the above background technique.

[0007] To achieve the above purpose, the present invention provides the following technical solution: A gesture pose estimation method based on multi-view and semi-supervised learning, including the following steps:

[0008] S1. Collect continuous multiple-view gesture image data, define this data set as D1, and D1 is an unlabeled gesture data set;

[0009] S2. Collect a small data set of gesture images and their hand key point coordinates, define this data set as D2, and D2 is a labeled gesture data set;

[0010] S3. Based on the data set D1, construct an unsupervised learning model;

[0011] S4. Based on the data set D2, construct a supervised learning model;

[0012] S5. Integrate the unsupervised learning model and the supervised learning model to form a semi-supervised learning model;

[0013] S6. When conducting actual applications after training is completed, use the trained autoencoder.

[0014] As a further optimization of this technical solution: In S3, the unsupervised learning model includes an encoder and a decoder;

[0015] Among them, the input of the decoder is the latent code output by the encoder, and the output target is to approximate the gesture image of another view.

[0016] As a further optimization of this technical solution: In S4, the supervised learning model includes an encoder and a projection transformer;

[0017] Among them, the input of the projection transformer is the output of the encoder, and the output of the projection transformer is to approximate the hand key point coordinates as a reference.

[0018] As a further optimization of this technical solution: The input of the encoder in both the unsupervised learning model and the supervised learning model is the gesture image of a certain view, and their outputs are both latent codes, that is, the 3D coordinates of the gesture posture.

[0019] As a further optimization of this technical solution: In S5, the semi-supervised learning model is to share the autoencoders of the unsupervised learning model and the supervised learning model, and the structures of other parts remain unchanged;

[0020] Among them, after adding the unsupervised loss function and the supervised loss function, it is used as the total loss function for training the semi-supervised learning model.

[0021] As a further optimization of this technical solution: In S6, input the gesture image to be tested into the autoencoder, and the autoencoder will output the 3D coordinate information of the test gesture.

[0022] Compared with the prior art, the beneficial effects of the present invention are:

[0023] 1. In the present invention, an unsupervised learning model and a supervised learning model are fused to form a semi-supervised learning model. By using the semi-supervised learning model, the problem of effectively utilizing a large number of unlabeled samples in the case of a small number of labeled samples is effectively solved. That is, an unsupervised learning model is constructed using a large number of unlabeled samples, and this process does not require labeled samples. At the same time, in order to further improve the accuracy of gesture pose estimation, a supervised learning process is constructed using a small number of labeled samples and fused with the unsupervised learning process to form a semi-supervised learning mechanism, so that both accurate estimation of gesture poses can be obtained and the demand for labeled samples can be greatly reduced.

[0024] 2. In the present invention, unlabeled samples provided from multiple perspectives are used for unsupervised training. A loss function is constructed by using the internal consistency constraints of different perspectives, and a small number of labeled samples are appropriately utilized, thereby constructing a semi-supervised learning mechanism that can achieve low-cost and high-precision 3D gesture pose estimation based on multi-perspective visual information.

[0025] 3. In the present invention, a semi-supervised learning system can be constructed by using a large number of unsupervised multi-perspective samples on the premise of only having a small number of labeled samples, which can achieve accurate estimation of high-precision gesture poses and has good economy. On the publicly available gesture pose estimation dataset Panoptic, the deviation of 3D estimation is less than 10 mm. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a schematic flowchart of a gesture pose estimation method based on multi-perspective and semi-supervised learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention.

[0028] Embodiment

[0029] Please refer to Figure 1 , the present invention provides a technical solution: a gesture pose estimation method based on multi-perspective and semi-supervised learning, including the following steps:

[0030] S1. Collect continuous gesture image data from multiple perspectives, and define this dataset as D1, and D1 is an unlabeled gesture dataset;

[0031] S2. Collect a dataset of a small number of gesture images and their hand key point coordinates, and define this dataset as D2, and D2 is a labeled gesture dataset;

[0032] S3. Based on the dataset D1, construct an unsupervised learning model;

[0033] S4. Based on the dataset D2, construct a supervised learning model;

[0034] S5. Integrate the unsupervised learning model and the supervised learning model to form a semi-supervised learning model;

[0035] S6. When performing actual applications after training, use the trained autoencoder.

[0036] In this embodiment, specifically: in S1, multiple cameras or multiple industrial CCDs can be used for gesture image acquisition, and they are evenly distributed in a circular shape in the scene;

[0037] Among them, during the acquisition process, the tester needs to change various gestures so as to collect gesture data at various angles and in various forms.

[0038] In this embodiment, specifically: in S3, the unsupervised learning model includes an encoder and a decoder;

[0039] Among them, the input of the decoder is the latent code output by the encoder, and the output target is to approximate the gesture image from another perspective.

[0040] In this embodiment, specifically: in S4, the supervised learning model includes an encoder and a projection transformer;

[0041] Among them, the input of the projection transformer is the output of the encoder, and the output of the projection transformer is to approximate the hand key point coordinates as a reference.

[0042] In this embodiment, specifically: the input of the encoder in both the unsupervised learning model and the supervised learning model is the gesture image from a certain perspective, and their outputs are both latent codes, that is, the 3D coordinates of the gesture posture.

[0043] In this embodiment, specifically: in S5, the semi-supervised learning model is to share the autoencoders of the unsupervised learning model and the supervised learning model, and the structures of other parts remain unchanged;

[0044] Among them, after adding the unsupervised loss function and the supervised loss function, it is used as the total loss function for training the semi-supervised learning model.

[0045] In this embodiment, specifically: in S6, input the gesture image to be tested into the autoencoder, and the autoencoder will output the 3D coordinate information of the test gesture.

[0046] Experimental example

[0047] The preprocessing process of the original signal is completed using the integrated tool MATLAB or the open-source Python algorithm, and the neural network model is constructed based on Pytorch of the deep learning framework;

[0048] A1. Based on multiple cameras that are evenly distributed in a circular pattern in the scene, collect multiple segments of continuous video signals, thereby obtaining gesture images from multiple consecutive viewpoints. During the collection process, the tester needs to change various gestures in order to collect gesture data from various angles and in various forms, and define this dataset as D1, which is an unlabeled gesture dataset;

[0049] A2. Collect a small dataset of gesture images and their hand key-point coordinates, and define this dataset as D2, which is a labeled gesture dataset. This data can utilize publicly available gesture pose datasets or be extracted based on open-source hand key-point detection algorithms. The hand key-points adopt the standard 21-point pattern, that is, 4 key-points for each finger and 1 key-point for the palm part;

[0050] A3. Based on dataset D1, construct an unsupervised learning model:

[0051] Select multi-viewpoint images at a certain moment t, defined as where i represents a certain viewpoint. If there are a total of 4 viewpoints, the value range of i is [1, 2, 3, 4]. Arbitrarily select an image from one of the viewpoints as the input of the encoder. The encoder is constructed using a multi-layer convolutional neural network, whose input dimension is the same as the size of the gesture image, and its output is a latent code, with the dimension defined as N×3, where N is the number of key-points representing the three-dimensional gesture pose. Set N = 21. The input of the decoder is this latent code, and the decoder is also constructed using a convolutional neural network or a fully connected neural network, and its output dimension is the same as the input dimension of the encoder. Select an image from another viewpoint different from the above as the target image for the decoder output. Define the loss function at the pixel level and the feature level respectively, and the specific operations are as follows:

[0052]

[0053] In formulas (1) and (2), E() represents the encoder, D() represents the decoder, and T() represents the feature transformer. The feature transformer adopts the pre-trained residual network ResNet18, that is, an 18-layer residual network;

[0054] A4. Based on dataset D2, construct a supervised learning model:

[0055] The encoder in the unsupervised learning model is the same as that in the supervised learning model, thus realizing the mechanism of sharing the encoder in unsupervised learning and supervised learning. The input of the key-point projection transformer is the output of the encoder, and its function is to project the 3D coordinates of 21 standard key points of the hand into a two-dimensional space. The goal of supervised training is to make the extracted hand key-point information consistent with the reference information. The loss function operation is defined as follows:

[0056]

[0057] In formula (3) thereof, P represents the key-point projection transformer, and G() represents the reference information of the labeled 2D coordinates of the key points;

[0058] A5. After adding the unsupervised loss function and the supervised loss function, the total loss function for training the semi-supervised model is defined as:

[0059] L(θ e ,θ d ,θ p ) = L1(θ e ,θ d ) + L2(θ e ,θ d ) + L3(θ p ) (4);

[0060] A6. When actually applied after training, only the trained autoencoder needs to be used. Input the gesture image to be tested into the autoencoder, and the autoencoder will output the 3D key-point coordinate information of the gesture.

[0061] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A hand gesture estimation method based on multi-view and semi-supervised learning, characterized in that: The following steps are involved: S1, collect continuous multi-view gesture image data, define the data set as D1, and D1 is an unlabeled gesture data set; S2, collect a small number of gesture images and their hand key point coordinates data set, define this data set as D2, and D2 is a labeled gesture data set; S3, build an unsupervised learning model based on data set D1; S4, build a supervised learning model based on data set D2; S5. Fusion of the unsupervised learning model with the supervised learning model to form a semi-supervised learning model; S6. After the training is completed, the trained autoencoder can be used for actual application.

2. The method for hand gesture estimation based on multi-view and semi-supervised learning according to claim 1, characterized in that: In S3, the unsupervised learning model consists of an encoder and a decoder; The input of the decoder is the hidden code output by the encoder, and the output goal is to approximate the gesture image from another perspective.

3. The method for hand gesture estimation based on multi-view and semi-supervised learning according to claim 2, characterized in that: In S4, the supervised learning model consists of an encoder and a projection transformer; The input of the projective transformer is the output of the encoder, and the output of the projective transformer is the coordinates of the hand key points approximated as a reference.

4. The method for hand gesture estimation based on multi-view and semi-supervised learning according to claim 3, characterized in that: The input of the encoder in both the unsupervised learning model and the supervised learning model is a gesture image from a certain perspective, and its output is a hidden code, that is, the 3D coordinates of the gesture posture.

5. The method for hand gesture estimation based on multi-view and semi-supervised learning according to claim 1, characterized in that: In S5, the semi-supervised learning model is to share the autoencoder of the unsupervised learning model with the supervised learning model, and the other parts of the structure remain unchanged; Among them, the unsupervised loss function and the supervised loss function are accumulated as the total loss function for training the semi-supervised learning model.

6. The method for hand gesture estimation based on multi-view and semi-supervised learning according to claim 1, characterized in that: In S6, the gesture image to be tested is input into the autoencoder, and the autoencoder outputs the 3D coordinate information of the test gesture.

Citation Information

Cited By

  • An automatic labeling-oriented self-supervised hand pose estimation method

    CN122799457A