Human Skeleton Data Generation Method Based on GAN Network

Through the human skeleton data generation method based on GAN network, the problems of high production cost of human behavior recognition data sets and lack of action data are solved, and virtual bone sequence data is generated, which improves the diversity of the data set and the generalization ability of the model, and reduces the overfitting problem.

CN114596635BActive Publication Date: 2025-05-30XIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210228558.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-08
Publication Date
2025-05-30
Estimated Expiration
2042-03-08

AI Technical Summary

Technical Problem

In the prior art, the production cost of human behavior recognition data sets is high, and some categories of action data are scarce, resulting in overfitting problems during training.

Method used

Using the human skeleton data generation method based on the GAN network, a space-time GAN network model is built, including a spatial human body generation network and a human body action sequence generation network, and using W distance and gradient penalty terms as loss functions to generate virtual skeleton sequence data to expand the data set.

Benefits of technology

It effectively reduces the cost of producing human behavior recognition data sets, increases the diversity of action data, reduces the overfitting problem, and improves the generalization ability of human behavior recognition network models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114596635B_ABST
    Figure CN114596635B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating human skeleton data based on a GAN (Generative Adversarial Network) network. A spatio-temporal GAN network model is built, which is divided into a spatial human body generation network and a human body motion sequence generation network. Overall, using the idea of GAN adversarial generation, new internal structures are proposed for the generator and discriminator respectively on different models in time and space, so as to generate human skeleton data using the spatio-temporal GAN network model. The present invention solves the problems that the production cost of the current mainstream behavior recognition data set is high, and the action data of some categories is scarce. The samples generated by the generation method can be used to expand the human behavior recognition data set, effectively reducing the overfitting problem in the training process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of human body bone pose generation, and particularly relates to a method for generating human body bone data based on a GAN network. Background Art

[0002] After development, visual human behavior recognition has become a forefront hot issue in related academic fields and has been widely valued by the academic community. The algorithm research of visual human behavior recognition can be roughly divided into traditional algorithms based on bone models and deep learning recognition algorithms.

[0003] In traditional algorithms, a method of combining manually extracting features from 3D human body bone data (such as the covariance of 3D joint positions on an action sequence, the histogram of 3D skeleton joint positions (HOJ3D), joint vectors, and angle information, etc.) with conventional classifiers (such as support vector machines (SVM), hidden Markov models (HMM), and random forest algorithms, etc.) is adopted. Currently, most mainstream human behavior recognition research uses deep learning methods. Existing deep learning-based human behavior recognition methods can be roughly divided into three types: recurrent neural network models (RNN), convolutional neural network models (CNN), and graph convolutional network models (GCN). Among them, the data structures of the connection points are respectively represented as vector sequences, pseudo-images, and undirected graphs.

[0004] Currently, the method with the best recognition effect is the human behavior recognition method based on graph convolution (GCN). The dynamic skeleton can be naturally represented in the form of 2D or 3D coordinates through the time series of human joint positions. Then, by analyzing the coordinate changes of human joint points, the recognition of the overall human behavior can be achieved. Early methods for action recognition based on skeletons only formed feature vectors using joint coordinates at each time step and performed temporal analysis on them. However, taking the 3D coordinates of human body bone joint points in the video time series requires huge human and material resources, and the collected human actions also have certain limitations. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for generating human body bone data based on a GAN (Generative Adversarial Network) network, which solves the problems that the production cost of the current mainstream behavior recognition data set is high, and the action data of some categories is scarce. The samples generated by the generation method can be used to expand the human behavior recognition data set and effectively reduce the overfitting problem in the training process.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A method for generating human body bone data based on a GAN network includes the following steps:

[0008] Step 1: Obtain a dataset of human skeletal actions;

[0009] Step 2: Preprocess the dataset in Step 1 as the training set, and preprocess the action sequence to be recognized as the test set;

[0010] Step 3: Build a spatio-temporal GAN network model. The spatio-temporal GAN network model includes a spatially human body generation network and a human action sequence generation network connected in series. The spatially human body generation network includes a single-frame generator and a single-frame discriminator, and the human action sequence generation network includes a sequence generator and a sequence discriminator;

[0011] Step 4: Use the training set to train the spatially human body generation network and the human action sequence generation network in turn; When training, use the Wasserstein distance and a gradient penalty term as the loss function;

[0012] Step 5: Use the trained spatio-temporal GAN network model to generate skeletal sequence data for the test set.

[0013] The features of the present invention also lie in:

[0014] Specifically, Step 2 is as follows: First, exclude samples that deviate from the actual due to sensor errors; Secondly, perform normalization processing on the data, that is, for each set of data, divide all coordinate points by the maximum value so that the sizes of all coordinate points are distributed between [0, 1]; Finally, convert the relative coordinates into absolute coordinates to obtain the skeletal joint point coordinate data.

[0015] The single-frame generator in Step 3 includes an input layer and 5 fully connected layers connected in sequence, and a normalization layer is connected after each fully connected layer. The single-frame discriminator is composed of 5 fully connected layers connected in sequence.

[0016] The sequence generator in Step 3 includes an input layer, 3 fully connected layers and 5 LSTM networks connected in sequence. The sequence discriminator includes 5 LSTM networks connected in sequence.

[0017] The operation formula of the LSTM network is as follows:

[0018] h t =σ(W xh x t +W hh h t-1 +b) (1)

[0019] Among them, W xh 、W hh represent the network parameters to be trained, x t represents the eigenvalue of the current node, h t-1 represents the eigenvalue accumulated at the previous moment, b represents the bias eigenvalue, and σ() represents a non-linear activation function.

[0020] Step 4 specifically includes:

[0021] Step 4.1, initialize network parameters;

[0022] Step 4.2, use the mapping of the input to the single-frame generator through the fully connected layer. The generated fake samples and the real labels are sent to the discriminator together, so that the single-frame discriminator can better distinguish the two. Then, the gradients obtained by the single-frame discriminator are fed back to the generator to train the generator, and the training is repeated until the network converges.

[0023] Step 4.3, perform equidistant sampling on the training set, so that 16 frames represent a human action sequence. Take 16 generators of the spatial human body generation network that have been trained in the previous step, and through the fully connected layer mapping, then send this network to the LSTM network layer with 16 nodes. The vectors output by each node are used as the input of each spatial generator, so that the vectors input to the 16 generators are correlated in time. In this way, the generated virtual 16-frame human behavior and the samples obtained by equidistant sampling of the real data are sent to the discriminator together, so that the discriminator can distinguish the generated human action sequence and the real human action sequence. Then, the gradients are fed back to the generator until the network converges.

[0024] The initialized parameters in Step 4.1 are: the number of times to traverse all data during training is set to an integer between 100 and 200, the number of samples for each batch training is set to one of {8, 16, 32, 64}, and the initial learning rate is 0.000001.

[0025] The generators and discriminators of the spatial human body generation network and the human action sequence generation network both adopt an alternating training method. For each batch of samples, the discriminator is trained 6 times first, and then the generator is trained 1 time.

[0026] The formula for the loss function is:

[0027]

[0028] L D = D(G(z)) - D(x) + λ(||gradD(x)|| 2 - 1) 2 (3)

[0029] where, L G represents the loss function of the generator, L D represents the loss function of the discriminator; x represents the real samples in the dataset, G(z) represents the samples generated by the generator, D(x) represents the score of the real samples through the discriminator, D(G(z)) represents the score of the generated samples through the discriminator, ||gradD(x)|| 2Denotes the second norm of the gradient of the generated sample in the discriminator, which constitutes the main part of the gradient penalty term, and λ represents the weight of the gradient penalty term.

[0030] The beneficial effects of the present invention are:

[0031] The overall generation method is divided into two modules: a spatial generation model for human skeletal postures and a generation model for human skeletal posture sequences. Overall, the network utilizes the idea of GAN adversarial generation and proposes new internal structures for the generator and discriminator respectively under different constraints in time and space.

[0032] In the generation model of spatial human skeletal postures, it aims to utilize the existing data to obtain a generator for generating a single-frame human skeletal posture; during the training process of the generation model of human skeletal time series, the spatial human skeletal posture generator that has been generated in the previous step is used for fine-tuning operations, so that on the time axis, the frame data generated by multiple generators has a certain correlation in the time dimension. Then, a complete virtual sequence sample is obtained through the principle of inter-frame interpolation.

[0033] The present invention can generate a large number of virtual samples using the current popular generation model, which can reduce the cost of manual labeling. By expanding the data with virtual samples, the generalization ability of the current mainstream human behavior recognition network model can be effectively improved, and overfitting can be reduced to a certain extent. And through virtual domain training of the network and then through the method of transfer learning, the test recognition rate of real human behavior samples can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is the method flow chart of the human skeletal data generation method based on the GAN network of the present invention;

[0035] Figure 2 is the framework diagram of the spatial human generation network in the present invention.

[0036] Figure 3 is the framework diagram of the human action sequence generation network in the present invention;

[0037] Figure 4 is the training process diagram in Embodiment 1 of the present invention;

[0038] Figure 5 is the generation result diagram of the spatial human generation network in Embodiment 1 of the present invention;

[0039] Figure 6 is the final generation result of the skeletal sequence data in Embodiment 1 of the present invention.

[0040] Figure 7 is the loss change of the generator and discriminator during the training of the spatial human generation model in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0041] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0042] The method for generating human skeleton data based on the GAN network of the present invention, such as Figure 1 , includes the steps:

[0043] Step 1, obtaining a dataset of human skeleton actions;

[0044] Step 2, preprocessing the dataset in Step 1 as a training set, and preprocessing the action sequence to be recognized as a test set;

[0045] Step 3, building a spatio-temporal GAN network model, where the spatio-temporal GAN network model includes a spatially human body generation network and a human action sequence generation network connected in series in sequence. The spatially human body generation network includes a single-frame generator and a single-frame discriminator, and the human action sequence generation network includes a sequence generator and a sequence discriminator;

[0046] Step 4, using the training set to train the spatially human body generation network and the human action sequence generation network in sequence; when training, the Wasserstein distance (Wasserstein, earth mover's distance) and a gradient penalty term are used as loss functions;

[0047] Step 5, using the trained spatio-temporal GAN network model to generate skeleton sequence data for the test set.

[0048] Among them, for the dataset in Step 1, on the one hand, existing datasets for human behavior recognition tasks can be used, such as the UCF101, Kinetics, or NTU RGBD datasets, etc. These datasets have been widely used in human behavior recognition tasks. On the other hand, a self-made dataset can also be used. Sensors are used to collect the joint point coordinates during human actions, and the collected data is classified and integrated, and labeled with the action type to make a self-made dataset.

[0049] Specifically, Step 2 is as follows: First, artificially exclude samples that deviate from the actual due to sensor errors. Secondly, normalize the data, that is, for each set of data, divide all coordinate points by the maximum value so that the sizes of all coordinate points are distributed between [0,1]. Finally, convert the relative coordinates to absolute coordinates, that is, take the joint point of the human hip as the origin and perform a translation transformation. Specifically, it can be implemented by subtracting the joint point coordinates of the hip from all joint point coordinates.

[0050] In Step 3, such as Figure 2 Figure 3, in the spatial human body generation network, the input of the single-frame generator is a random noise vector of length 100 that satisfies Z~N(0, 1). The network consists of 5 fully connected layers, and a batch normalization layer (BN layer) is connected after each fully connected layer to accelerate the network convergence, which can realize the mapping from the noise to the 25*3 coordinates, where 25 represents the number of human joint points and 3 represents the (x, y, z) coordinates of the joint points. The single-frame discriminator also consists of 5 fully connected layers, which respectively input the generated fake samples and the real sample data, and finally obtain the scores of true and false;

[0051] The human body motion sequence generation model, where part of the input of the sequence generator is a noise vector of length 10 that satisfies Z~N(0, 1). First, it is mapped to a feature vector of length 1200 through 3 fully connected layers, and then split into 16*25*3, where 16 represents the number of frames, 25 represents the total number of joint points, and 3 represents 3 dimensions. Then, 5 LSTM layers are connected. The operation formula of the LSTM layer is as follows:

[0052] h t =σ(W xh x t +W hh h t-1 +b) (1)

[0053] Among them, W xh 、W hh represent the network parameters to be trained, x t represents the eigenvalue of the current node, h t-1 represents the eigenvalue accumulated at the previous moment, b represents the bias eigenvalue, and σ() represents a non-linear activation function. That is, the eigenvalue of the next moment depends on the weight of its own eigenvalue plus the weight of the previous eigenvalue plus the offset and then passes through a non-linear activation function. In this way, the obtained eigenvalue has local semantic information, and the eigenvalue of the previous frame can be passed to the next frame. After that, the random vector with sequence semantics is input into the trained human body motion space generator to generate a motion sequence. The structure of the sequence discriminator is similar. It inputs the generated sequence samples and the real sequence samples into a 5-layer LSTM network, and finally gives the score situation of true and false.

[0054] The spatio-temporal GAN network model uses the Wasserstein distance plus a gradient penalty term as the loss function, which can help the training to be more stable and reliable. The loss function of the model is as follows:

[0055]

[0056] L D =D(G(z))-D(x)+λ(||gradD(x)|| 2 -1) 2 (3)

[0057] Among them, L G represents the loss function of the generator, and L D represents the loss function of the discriminator; x represents the real samples in the dataset, G(z) represents the samples generated by the generator, D(x) represents the score of the real samples passing through the discriminator, D(G(z)) represents the score of the generated samples passing through the discriminator, and ||gradD(x)|| 2 represents the two-norm of the gradient of the generated samples in the discriminator, which constitutes the main part of the gradient penalty term, and λ represents the weight of the gradient penalty term.

[0058] Among them, step 4 is specifically as follows:

[0059] Step 4.1, initialize the parameters. The training hyperparameters are set as follows: epoch is the number of times to train through all the data, which is set as an integer between 100 and 200, batch_size is the number of samples in each batch of training, which is set as one of {8, 16, 32, 64}, and learning_rate is the learning rate, with the initial learning rate being 0.000001.

[0060] Step 4.2, train the spatial human body generation network. At the beginning, input a random noise vector that follows a Gaussian distribution to the single-frame generator. The single-frame generator network outputs a 25*3 feature map representing the 3D coordinates of 25 human body nodes. Send the generated samples and the samples of the real data to the single-frame discriminator for training together. The single-frame discriminator will score and judge the input data. Through the loss function, it is hoped that the score of the generated samples is lower and the score of the real data samples is higher. When training the generator, it is hoped that the score of the discriminator is higher and closer to the real samples until the network converges.

[0061] When training the human body action sequence generation network, first, it is necessary to perform equally spaced sampling on the existing human body action sequences so that 16 frames represent a human body action sequence. Then, read 16 generators of the human body space models that have been trained in the previous step. The input is a noise with a length of 100 that follows a Gaussian distribution with a mean of 0 and a variance of 1. Map it to a vector with a length of 1200 through a fully connected layer, and then send this network to an LSTM network layer with 16 nodes. Each node outputs a vector with a length of 100 as the input of each sequence generator. In this way, it can be ensured that the vectors input to the 16 sequence generators are correlated in time. Obtain the generated virtual 16-frame human body behavior and the samples obtained by equally spaced sampling of the real data according to this method and send them to the sequence discriminator together, so that the sequence discriminator can distinguish the generated human body action sequences from the real human body action sequences, and then backpropagate the gradient to the generator to make the generated human body action sequences more and more realistic.

[0062] Example 1

[0063] In this embodiment, a human body drinking action sequence is used to generate human body bone data, which is carried out through the following steps:

[0064] Execute step 1. In this embodiment, the NTU-RGBD dataset is adopted. This dataset contains human body actions of 60 categories, with more than about 50,000 sample sequences. Among these action categories, 40 categories are daily behavior actions, 9 categories are life-related behaviors, and 11 categories are two-person interaction actions.

[0065] This dataset is collected by the Kinect sensor developed by Microsoft and uses three cameras at different angles. The collected data forms include 3D bone information, RGB frames, infrared sequences, and depth information. In this embodiment, only the posture generation of a single person is concerned.

[0066] Execute step 2, preprocessing. The human body actions in the NTU-RGBD dataset are composed of 25 joint points, denoted by X ij , representing the j-th joint point in the i-th frame (i = 1, 2,... n, j = 1, 2,..., 25). Each joint point has three-dimensional coordinates, that is, X ij =(x ij , y ij , z ij ). The normalization formula is as follows:

[0067]

[0068] That is, each coordinate is divided by the maximum point coordinate value in this action sequence, so that X ij ∈[0, 1], that is, the normalization operation is completed. Since the data collected by different sensors may have different origins relative to the sensors, such data is not convenient for the network to learn its distribution. Therefore, a coordinate absolutization operation is required. In the operation, first, a joint point of a certain part of the human body should be selected as the origin. Here, we select the joint point of the human hip as the origin, denoted by X io , then the coordinate absolutization formula is as follows:

[0069] X ij =X ij -X io (5)

[0070] Execute steps 3 to 5

[0071] As Figure 4 is the training process diagram in Embodiment 1 of this example; Figure 5 is the generation result diagram of the spatial human body generation network in Embodiment 1 of this example; Figure 6 is the final generation result of the bone sequence data in Embodiment 1 of this example.

[0072] As Figure 7 , during the process of training the spatial human body generation network, the losses of the generator and the discriminator oscillate back and forth without divergence and tend to stabilize at the 150th training epoch, and better human body spatial samples can be generated.

Claims

1. A method for generating human bone data based on a spatio-temporal GAN network, characterized in that, it includes the following steps: Step 1, obtain a dataset of human bone actions; Step 2, preprocess the dataset in Step 1 as a training set, and preprocess the action sequence to be recognized as a test set; Step 3, build a spatio-temporal GAN network model, the spatio-temporal GAN network model includes a spatially human generation network and a human action sequence generation network cascaded in sequence, the spatially human generation network includes a single-frame generator and a single-frame discriminator, and the human action sequence generation network includes a sequence generator and a sequence discriminator; The single-frame generator includes an input layer and 5 fully connected layers connected in sequence, and a normalization layer is connected after each fully connected layer, and the single-frame discriminator is composed of 5 fully connected layers connected in sequence; The sequence generator includes an input layer, 3 fully connected layers and 5 LSTM networks connected in sequence, and the sequence discriminator includes 5 LSTM networks connected in sequence; Step 4, use the training set to train the spatially human generation network and the human action sequence generation network in sequence; When training, use the Wasserstein distance and a gradient penalty term as the loss function; Specifically include: Step 4.1, initialize network parameters; Step 4.2, use the mapping through the fully connected layer input to the single-frame generator, and send the generated fake samples and real labels to the discriminator together, so that the single-frame discriminator can better distinguish the two, and then send the gradient obtained by the single-frame discriminator back to the generator to train the generator, and train repeatedly until the network converges; Step 4.3, sample the training set at equal intervals, so that 16 frames represent a human action sequence, take 16 generators of the spatially human generation network that have been trained in the previous step, map through the fully connected layer, and then send this network to the LSTM network layer with 16 nodes, and the vectors output by each node are used as the input of each spatial generator, so that the vectors input to the 16 generators are correlated in time, so as to obtain the generated virtual 16-frame human behavior and the samples obtained by equal-interval sampling of real data and send them to the discriminator together, so that the discriminator can distinguish the generated human action sequence from the real human action sequence, and then send the gradient back to the generator until the network converges; Step 5, use the trained spatio-temporal GAN network model to generate bone sequence data for the test set.

2. The method for generating human bone data based on a spatio-temporal GAN network according to claim 1, characterized in that, the specific content of Step 2 is: First, exclude the samples that deviate from the actual due to sensor errors; Secondly, perform normalization processing on the data, that is, for each group of data, divide all coordinate points by the maximum value, so that the sizes of all coordinate points are distributed between [0,1]; Finally, convert the relative coordinates into absolute coordinates to obtain the bone joint point coordinate data.

3. The method for generating human bone data based on a spatio-temporal GAN network according to claim 1, characterized in that, the operation formula of the LSTM network is as follows: (1) Among them, represents the network parameters to be trained, represents the eigenvalue of the current node, represents the eigenvalues accumulated at the previous moment, represents the bias eigenvalue, represents a non-linear activation function.

4. The method for generating human bone data based on a spatio-temporal GAN network according to claim 1, characterized in that, The initialization parameters in step 4.1 are as follows: the number of times to train through all data is set as an integer between 100 and 200, the number of samples for each batch training is set as one of {8, 16, 32, 64}, and the initial learning rate is 0.000001.

5. The method for generating human skeleton data based on the spatio-temporal GAN network according to claim 1, characterized in that, for the generators and discriminators of the spatial human body generation network and the human action sequence generation network, an alternating training method is adopted. For each batch of samples, the discriminator is first trained 6 times, and then the generator is trained 1 time.

6. The method for generating human skeleton data based on the spatio-temporal GAN network according to claim 1, characterized in that, the formula of the loss function is: (2) (3) Among them, L G represents the loss function of the generator, L D represents the loss function of the discriminator; represents the real samples in the dataset, represents the samples generated by the generator, represents the score of the real samples passing through the discriminator, represents the score of the generated samples passing through the discriminator, represents the second norm of the gradient of the generated samples in the discriminator, which constitutes the main part of the gradient penalty term, represents the weight of the gradient penalty term.

Citation Information

Patent Citations

  • Character image generation method guided by text based on generative adversarial network

    CN110021051A

  • Behavior recognition method based on local scene perception graph convolutional network

    CN113255514A