A Behavior Recognition Method Based on Multi-Stream Fusion Graph Convolutional Network

Through the multi-stream fusion graph convolution network, the multi-class behavior characteristics are extracted and fused by skeleton normalization and spatiotemporal graph convolution network, the problem of under-explored information complementarity in the existing methods is solved, and the accuracy and robustness of behavior recognition are improved.

CN114187653BActive Publication Date: 2025-07-08FUDAN UNIVERSITY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111356801.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-16
Publication Date
2025-07-08
Estimated Expiration
2041-11-16

AI Technical Summary

Technical Problem

The existing behavior recognition methods fail to fully explore the complementarity between multiple types of information, resulting in insufficient recognition accuracy, large amount of calculations of image-based methods, and insufficient robustness of methods based on human skeletons.

Method used

Multi-stream fusion graph convolution network is adopted to extract and fuse multiple behavioral features, including joint nodes, bones and motion information, and build local and global connection graphs to reduce training difficulty and improve recognition accuracy through skeleton normalization, space-time graph convolution network and multi-stream feature fusion network.

Benefits of technology

It improves the accuracy of behavior recognition, reduces the amount of calculation, and enhances the robustness and recognition effect of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114187653B_ABST
    Figure CN114187653B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of behavior recognition, and specifically relates to a behavior recognition method based on a multi-stream fusion graph convolutional network. The present invention mainly conducts behavior recognition by extracting and fusing multiple types of behavior information, which is divided into three stages: data preprocessing, feature extraction, and feature fusion. In the data preprocessing stage, three skeleton normalization measures are proposed to reduce the influence of factors such as human position, camera perspective, and the distance between the human body and the camera on the representation of human skeleton data; in the feature extraction stage, a global connection graph of the skeleton is constructed to directly learn the mutual relationship between distant joints; in the feature fusion stage, the features of three types of information are fused in two stages. The method proposed by the present invention makes more effective use of the complementary information of multiple types of behaviors. The proposed skeleton normalization measures make the representation of the human skeleton have affine invariance, reduce the training difficulty of the network, and achieve good results on the public dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of behavior recognition, and particularly relates to a behavior recognition method based on a multi-stream fusion graph convolutional network. Background Art

[0002] The goal of behavior recognition is to identify the behaviors of people in videos. This technology plays an important role in fields such as intelligent security, video retrieval, intelligent care, and advanced human-computer interaction, and thus has received extensive attention from the academic and industrial communities. Behavior recognition is divided into two major research directions: behavior recognition based on static images and behavior recognition based on videos. The former only recognizes the behaviors of people in an image based on a single image, ignoring motion information; while the latter recognizes based on the image sequence obtained from the video. The behavior recognition method based on videos can be divided into two types according to different input data: the behavior recognition method based on images and the behavior recognition method based on human skeletons. The input of the former is an image sequence, while the input of the latter is a human skeleton sequence. The behavior recognition method based on images is vulnerable to factors such as the background environment, illumination, and perspective of the image data, and this type of method requires a large amount of computation and high computing power in practical applications. Compared with the behavior recognition method based on images, the behavior recognition method based on human skeletons is more robust, not affected by the background environment, and has less computation, becoming a research hotspot in recent years. The human skeleton contains joint point information, bone information, and motion information, and these three types of information are closely related and complementary to each other. However, the mainstream methods have a relatively simple way of fusing multiple types of information and do not fully explore the complementarity between multiple types of information. Summary of the Invention

[0003] To solve the problems existing in the prior art, the present invention proposes a behavior recognition method based on a multi-stream fusion graph convolutional network. This method is an improvement aiming at the defect that the existing models do not well explore the complementarity between multiple types of information. The skeleton normalization measure proposed by the present invention makes the representation of the human skeleton have affine invariance and reduces the training difficulty of the network; aiming at the problem that the existing methods have a relatively simple way of fusing multiple types of information and do not fully explore the complementarity between multiple types of information, the method proposed by the present invention can better extract and fuse multiple types of behavior features, make more effective use of the complementary information of multiple types of behaviors, and improve the accuracy of behavior recognition.

[0004] The present invention mainly conducts behavior recognition by extracting and fusing multiple types of behavior information, which is divided into three stages: data preprocessing, feature extraction, and feature fusion. In the data preprocessing stage, three skeleton normalization measures are proposed to reduce the influence of factors such as human position, camera perspective, and the distance between the human body and the camera on the representation of human skeleton data; in the feature extraction stage, a global connection graph of the skeleton is constructed to directly learn the mutual relationship between distant joint points; in the feature fusion stage, the features of three types of information are fused in two stages. The technical solution of the present invention is specifically introduced as follows.

[0005] The present invention proposes a behavior recognition method based on a multi-stream fusion graph convolutional network, which is divided into 3 stages: data preprocessing, feature extraction, and feature fusion; among them:

[0006] In the data preprocessing stage, the skeleton normalization module is used to process the input human skeleton sequence data, that is, joint point data, to obtain normalized human skeleton data, and then the bone data and motion data are further obtained. The bone data is obtained by obtaining the vector formed between adjacent joint points, and the motion data is obtained by obtaining the displacement of the same joint point between adjacent frames. Among them, the human skeleton sequence data can be expressed as T represents the length of the skeleton sequence. In the present invention, T takes 300, and x t ∈R V×C represents the joint point coordinates of the t-th skeleton, V represents the number of joint points in the human skeleton. In the present invention, V = 14, and C represents the dimension of the joint point coordinates. In the present invention, C = 3, indicating that each joint point has three coordinates: x, y, and z.

[0007] Among them, the human joint numbers and their meanings are as follows:

[0008] 0: Neck; 1: Head; 2: Right shoulder; 3: Right elbow; 4: Right wrist; 5: Left shoulder; 6: Left elbow; 7: Left wrist;

[0009] 8: Right hip; 9: Right knee; 10: Right ankle; 11: Left hip; 12: Left knee; 13: Left ankle.

[0010] In the feature extraction stage, the spatio-temporal graph convolutional network is used to extract the spatio-temporal features of the joint point data, bone data, and motion data respectively, to obtain joint point features, bone features, and motion features;

[0011] In the feature fusion stage, the multi-stream feature fusion network is used to further fuse the joint point features, bone features, and motion features, and then the prediction result of the behavior is obtained through the classifier; among them, the method of using the multi-stream feature fusion network for fusion is as follows:

[0012] In the first stage, first, the three features are spliced pairwise, and the spliced features are input into two consecutive graph convolutional units to fuse the features of two types of information; then, the fused features are input into the pooling layer.

[0013] In the second stage, two fully connected layers are connected after the pooling layer, and there is a ReLU layer in the two fully connected layers to obtain three classification features f0, f1, and f2. Then, the three-way features are fused to obtain the overall classification feature f3, and f3 = f0 + f1 + f2.

[0014] In the present invention, the skeleton normalization module in the data preprocessing stage proposes a skeleton normalization method, which includes three processing steps: position normalization, perspective normalization, and scale normalization, as follows:

[0015] (1) Position normalization

[0016] First, perform position normalization processing on the input skeleton sequence, that is, given the human skeleton sequence where, x t represents the t-th skeleton in the sequence, and T represents the length of the sequence. Update the coordinates of all joint points according to the following formula:

[0017]

[0018] where, x t,i represents the coordinate of the i-th joint point of the skeleton x t , and i = 0, 1,..., 13. Denote the skeleton sequence after position normalization processing as X 1 , and in the above formula is the coordinate of the i-th joint point of the t-th skeleton 1 of X.

[0019] (2) Perspective normalization

[0020] Then, perform a rotation transformation on the skeleton sequence X 1 after position normalization. Specifically, first determine the rotation matrix R according to the first skeleton x1 of the sequence X, and the formula is as follows:

[0021]

[0022] where, the vectors v x , v y , v z are determined by x1, and are calculated respectively as follows:

[0023] (a) Determine the horizontal direction vector v x according to the 2nd joint and the 5th joint of x1, as follows:

[0024] v x = x1,5 -x 1,2

[0025] (b) Determine v according to the following formula y :

[0026]

[0027] where v 1,0 represents the vector from joint point 1 to joint point 0 in the skeleton x1, that is:

[0028] v 1,0 = x 1,1 - x 1,0

[0029] represents the projection of v 1,0 on v x ;

[0030] (c) After obtaining v x and v y , then find the vector v z perpendicular to these two vectors according to the following formula:

[0031] v z = v x × v y

[0032] Then rotate the coordinates of all joint points in X1 according to the following formula:

[0033]

[0034] where is the coordinate of the j-th joint point, j = 0, 1,..., 13. Denote the skeleton sequence after view normalization as X 2 , and in the above formula is the coordinate of the j-th joint point of the t-th skeleton 2 of X .

[0035] (3) Scale normalization

[0036] Finally, perform scale normalization. For the skeleton sequence X 2 , first scale the distance between joint point 0 and joint point 1 to 1, that is, calculate the scaling factor r according to the following formula:

[0037]

[0038] Then update the coordinates of all joint points in X 2 according to the following formula:

[0039]

[0040] Denote the skeleton sequence after scale normalization as X 3 , where in the above formula is X 3 the t-th skeleton of the k-th joint coordinate of

[0041] In the present invention, in the feature extraction stage, spatio-temporal features of joint data, bone data, and motion data are extracted through a spatio-temporal graph convolutional network. The implementation steps of the spatio-temporal graph convolutional network are as follows:

[0042] (1) Construct a spatio-temporal graph of the human skeleton

[0043] The construction of the spatio-temporal graph of the human skeleton is divided into three steps:

[0044] (a) For the skeleton sequence X 3 and the set H of joints that are physiologically adjacent to each other in the human body. The definition of H is as follows. For each 3 in X connect its physiologically adjacent joints to obtain partial spatial edges, thereby constructing a local connection graph

[0045] H ={(0,1),(0,2),(0,5),(2,3),(3,4),(5,6),(6,7),(8,9),(9,10),(11,12),(12,13)}

[0046] (b) Given the set M, where M is a set of joints that are not physiologically adjacent but are closely related. Its definition is as follows. For the given skeleton sequence X 3 for each in it, establish edges according to M to obtain a global connection graph. Combine it with the local connection graph obtained in step (a) to form a skeleton spatial graph G S ={V, E S}), where V represents the set of joints, V ={v t,i |t = 1…T, i = 0…N - 1}, T is the length of the skeleton sequence, N is the number of joints in the skeleton, and E S is the set of spatial edges, E S ={(v t, i v t,j )|(i, j) ∈ U}, and U is the union of H and M

[0047] M ={(1,4),(1,7),(4,7),(4,13),(4,10),(7,10),(7,13),(10,13)}

[0048] (c) For the skeleton spatio-temporal graph G obtained in step (b) S , establish temporal edges between the same joint points in the skeleton spatio-temporal graph between adjacent frames to obtain the set E of temporal edges T , E T = {(v t,i v t+1,i ) | t = 1…T - 1, i = 0…N - 1}, thus obtaining the skeleton temporal graph G T = {V, E T}, and finally obtaining the skeleton spatio-temporal graph G = {V, E}, where E = {E S , E T}, G = {G S , G T}.

[0049] (2) Spatio-temporal graph convolution

[0050] Perform spatio-temporal graph convolution on the human skeleton spatio-temporal graph obtained in step (1). The graph convolution in the space is implemented by ST-GCN, and two adaptive graphs proposed in 2S-AGCN are introduced. The graph convolution in the time is implemented by a 9×1 one-dimensional convolution.

[0051] The convolution operation adopted in the space is as follows:

[0052]

[0053] Among them, f in and f out are the input and output skeleton sequence matrices respectively; K v = 3 represents the convolution kernel size; k is the serial number of the set; w k is the weight parameter used for the k-th set; A k ∈ R N×N is the adjacency matrix; B k and C k are the weight parameters learned through the network. Among them, the calculation method of C k can be expressed as:

[0054]

[0055] Among them, W θk and represent the parameters of two 1×1 convolutions respectively. represents two embedded features obtained through convolution.

[0056] In the present invention, in the feature extraction stage, the spatio-temporal graph convolutional network is stacked by a batch normalization (BN) layer and six consecutive spatio-temporal graph convolutional units; each spatio-temporal graph convolutional unit has the same structure, including a spatial graph convolution (GCN-S), a BN layer, a ReLU layer, a Dropout layer, a temporal graph convolution (GCN-T), a BN layer, a ReLU layer, and a residual connection.

[0057] In the present invention, in the feature fusion stage, the method for designing the loss function in the multi-stream feature fusion network is as follows:

[0058] First, use the softmax classifier to process the four features f0, f1, f2, and f3 to obtain their predicted probability values, which are p0, p1, p2, and p3 respectively, and then construct the loss function as:

[0059] L = αL0 + βL1 + γL2 + δL3

[0060] where L0, L1, L2, and L3 are the losses corresponding to each type of feature respectively,

[0061]

[0062] where c represents the number of behaviors; y represents the true label of the sample, and α, β, γ, and δ are the weights of each type of loss respectively.

[0063] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0064] The representation of the human skeleton is made affine invariant through the skeleton normalization method, thereby reducing the training difficulty of the network; a local and global connection graph is constructed in the spatio-temporal graph convolutional network, enabling the network to simultaneously focus on the local and global parts of the human body; the proposed multi-stream feature fusion network effectively fuses various motion information, fully explores the complementarity between information, and improves the accuracy of action recognition. Description of the Drawings

[0065] Figure 1 is the flow chart of the action recognition method based on the multi-stream fusion graph convolutional network proposed by the present invention.

[0066] Figure 2 is the spatial graph of the human skeleton, (a) is the local connection graph, (b) is the global connection graph, and (c) is the skeleton spatial graph.

[0067] Figure 3 is the spatio-temporal graph of the human skeleton.

[0068] Figure 4 is the network structure of the spatio-temporal graph convolutional network.

[0069] Figure 5 is the network structure of the multi-stream fusion graph convolutional network. Detailed implementation manner

[0070] The present invention proposes a behavior recognition method based on a multi-stream fusion graph convolutional network, which is mainly divided into three stages: data preprocessing, feature extraction, and feature fusion; the structure of the entire network is as Figure 1 shown. In the data preprocessing stage, the skeleton normalization module is used to process the input human skeleton sequence data, that is, joint point data, to obtain normalized human skeleton data. Then, the human bone data and human motion data are respectively obtained from it. In the feature extraction stage, three spatio-temporal graph convolutional networks are used to extract the spatio-temporal features of the joint point data, bone data, and motion data respectively. In the feature fusion stage, the multi-stream feature fusion network is used to further fuse the features of the three kinds of information in two stages, and finally the prediction result of the behavior is obtained through the classifier.

[0071] In the data preprocessing stage, the skeleton normalization module is used to process the input human skeleton sequence data, that is, joint point data, to obtain normalized human skeleton data, and then the bone data and motion data are further obtained. The bone data is obtained by obtaining the vector formed between adjacent joint points, and the motion data is obtained by obtaining the displacement of the same joint point between adjacent frames. Among them, the human skeleton sequence data can be expressed as T represents the length of the skeleton sequence. In the present invention, T takes 300, and x t ∈R V×C represents the joint point coordinates of the t-th skeleton, V represents the number of joint points in the human skeleton. In the present invention, V = 14, and C represents the dimension of the joint point coordinates. In the present invention, C = 3, indicating that each joint point has three coordinates of x, y, and z.

[0072] Among them, the human joint numbers and their meanings are as follows:

[0073] 0: Neck; 1: Head; 2: Right shoulder; 3: Right elbow; 4: Right wrist; 5: Left shoulder; 6: Left elbow; 7: Left wrist;

[0074] 8: Right hip; 9: Right knee; 10: Right ankle; 11: Left hip; 12: Left knee; 13: Left ankle.

[0075] In the feature extraction stage, the spatio-temporal graph convolutional network is used to extract the spatio-temporal features of the joint point data, bone data, and motion data respectively, to obtain joint point features, bone features, and motion features;

[0076] In the feature fusion stage, the multi-stream feature fusion network is used to further fuse the joint point features, bone features, and motion features, and then the prediction result of the behavior is obtained through the classifier; among them, the method of fusing using the multi-stream feature fusion network is as follows:

[0077] In the first stage, first splice the three features pairwise, input the spliced features into two consecutive graph convolution units to fuse the features of two types of information; then, input the fused features into the pooling layer.

[0078] In the second stage, two fully connected layers are connected after the pooling layer, and there is a ReLU layer in the two fully connected layers to obtain three classification features f0, f1, and f2. Then, fuse the three-way features to obtain the overall classification feature f3, where f3 = f0 + f1 + f2.

[0079] The following are the specific steps:

[0080] 1. Data preprocessing

[0081] In the present invention, the skeleton normalization module in the data preprocessing stage proposes a skeleton normalization method, which includes three processing steps: position normalization, perspective normalization, and scale normalization, as follows:

[0082] (1) Position normalization

[0083] First, perform position normalization processing on the input skeleton sequence, that is, given the human skeleton sequence where, x t represents the t-th skeleton in the sequence, T represents the length of the sequence, and update the coordinates of all joint points according to the following formula:

[0084]

[0085] where, x t,i represents the i-th joint point coordinate of the skeleton x t , and i = 0, 1,..., 13. Denote the skeleton sequence after position normalization processing as X 1 , and the in the above formula is the i-th joint point coordinate of the t-th skeleton 1 of X .

[0086] (2) Perspective normalization

[0087] Then, perform a rotation transformation on the skeleton sequence X 1 after position normalization. Specifically, first determine the rotation matrix R according to the first skeleton x1 of the sequence X, and the formula is as follows:

[0088]

[0089] where, the vectors v x , v y , v z are determined by x1, and are calculated respectively as follows:

[0090] (a) Determine the horizontal direction vector v based on the 2nd joint and 5th joint of x1 x ,:

[0091] v x = x 1,5 - x 1,2

[0092] (b) Determine v according to the following formula y :

[0093]

[0094] where v 1,0 represents the vector from the 1st joint point to the 0th joint point in the skeleton x1, that is:

[0095] v 1,0 = x 1,1 - x 1,0

[0096] represents the projection of v 1,0 onto v x ;

[0097] (c) After obtaining v x and v y , then find the vector v perpendicular to these two vectors according to the following formula z :

[0098] v z = v x × v y

[0099] Then rotate the coordinates of all joint points in X1 according to the following formula:

[0100]

[0101] where is the coordinate of the jth joint point, j = 0, 1,..., 13. Denote the skeleton sequence after perspective normalization as X 2 , and in the above formula is the coordinate of the jth joint point of the tth skeleton 2 of X .

[0102] 2. Feature Extraction

[0103] Feature extraction is performed through a spatio-temporal graph convolutional network to extract spatio-temporal features of joint data, bone data, and motion data. The implementation steps of the spatio-temporal graph convolutional network are as follows:

[0104] (1) Construct a spatio-temporal graph of the human skeleton

[0105] The construction of the spatio-temporal graph of the human skeleton is divided into three steps:

[0106] (a) For the skeleton sequence X 3 and the set H of adjacent joint points in the human body physiology, the definition of H is as follows. For each 3 in X connect its adjacent joint points in physiology to obtain partial spatial edges, thereby constructing a local connection graph (as shown in Figure 2 (a)).

[0107] H = {(0,1),(0,2),(0,5),(2,3),(3,4),(5,6),(6,7),(8,9),(9,10),(11,12),(12,13)}

[0108] (b) Given the set M, M is a set of joint points that are not adjacent but closely related in physiology. Its definition is as follows. For the given skeleton sequence X 3 in each establish edges according to M to obtain a global connection graph (as shown in Figure 2 (b)). It forms the skeleton spatial graph G S = {V,E S} with the local connection graph obtained in step (a). The skeleton spatial graph is as shown in Figure 2 (c), where V represents the set of joint points, V = {v t,i |t = 1…T,i = 0…N - 1}, T is the length of the skeleton sequence, N is the number of joint points in the skeleton, and E S is the set of spatial edges, E S = {(v t,i v t,j )|(i,j) ∈ U}, and U is the union of H and M.

[0109] M = {(1,4),(1,7),(4,7),(4,13),(4,10),(7,10),(7,13),(10,13)}

[0110] (c) For the skeleton spatial graph G S obtained in step (b), establish temporal edges between the same joint points in the skeleton spatial graph between adjacent frames to obtain the set E T of temporal edges, E T = {(v t,i v t+1,i )|t = 1…T - 1,i = 0…N - 1}, thereby obtaining the skeleton temporal graph G T = {V,E T}, and finally obtaining the skeleton spatio-temporal graph G = {V,E}, as shown in Figure 3 , where E = {ES , E T},G = {G S , G T}.

[0111] (2) Spatiotemporal Graph Convolution

[0112] Perform spatiotemporal graph convolution on the human skeleton spatiotemporal graph obtained in step (1). The graph convolution in space is implemented using ST-GCN, and two adaptive graphs proposed in 2S-AGCN are introduced. The graph convolution in time is implemented using a 9×1 one-dimensional convolution.

[0113] The convolution operation adopted in space is as follows:

[0114]

[0115] Among them, f in and f out are the input and output skeleton sequence matrices respectively; K v = 3 represents the convolution kernel size; k is the serial number of the set; w k is the weight parameter used for the k-th set; A k ∈R N×N is the adjacency matrix; B k and C k are the weight parameters learned through the network. Among them, the calculation method of C k can be expressed as:

[0116]

[0117] Among them, W θk and represent the parameters of two 1×1 convolutions respectively. represents two embedded features obtained through convolution.

[0118] The spatiotemporal graph convolution network is stacked by a batch normalization (BN) layer and six consecutive spatiotemporal graph convolution units (G1 to G6). Each spatiotemporal graph convolution unit has the same structure: spatial graph convolution (GCN-S), BN layer, ReLU layer, Dropout layer, temporal graph convolution (GCN-T), BN layer, ReLU layer, and a residual connection. Its structure is as Figure 4 shown.

[0119] Among them, the input and output dimensions of the spatiotemporal graph convolution network are listed as follows:

[0120] The input dimension of G1 is 3×T×N, and the output dimension is 64×T×N.

[0121] The input dimension of G2 is 64×T×N, and the output dimension is 64×T×N.

[0122] The input dimension of G3 is 64×T×N, and the output dimension is 64×T×N.

[0123] The input dimension of G4 is 64×T×N, and the output dimension is

[0124] The input dimension of G5 is The output dimension is

[0125] The input dimension of G6 is The output dimension is

[0126] T is the length of the skeleton sequence, and N = 14 is the number of human joint points.

[0127] 3. Feature Fusion

[0128] The multi-stream fusion module is carried out in two stages; in the first stage, first, the three features output by the feature extraction stage are concatenated pairwise, and the dimension of the feature changes from to The concatenated features are input into two consecutive graph convolution units to fuse the features of two types of information. Then, the fused features are input into the pooling layer, and average pooling is performed on the two dimensions of N and T in the pooling layer. In the second stage, the pooling layer is followed by two fully connected layers, and there is a ReLU layer in the two fully connected layers. Then, three classification features f0, f1, and f2 are obtained. Then, the three-way features are fused to obtain the overall classification feature f3, and f3 = f0 + f1 + f2. The network structure of the multi-stream fusion module is as Figure 5 shown.

[0129] In the multi-stream fusion module, a loss function applicable to the present invention is designed. Specifically: first, the softmax classifier is used to process the four features f0, f1, f2, and f3 to obtain their predicted probability values, which are p0, p1, p2, and p3 respectively. Accordingly, the constructed loss function is:

[0130] L = αL0 + βL1 + γL2 + δL3

[0131] where L0, L1, L2, and L3 are the losses corresponding to each type of feature,

[0132]

[0133] where c represents the number of behaviors; y represents the true label of the sample. α, β, γ, and δ are the weights of each loss respectively. During the training process, the SGD optimizer is used, and the hyperparameters α, β, γ, and δ are set to 1, 1, 1, and 3 respectively.

[0134] Example 1

[0135] A behavior recognition method based on a multi-stream fusion graph convolutional network proposed by the present invention was experimented on the publicly available dataset NTU-RGB+D 60, and the results were compared with those of current mainstream methods. According to the mainstream practice, the experiments were carried out on two benchmarks, X-Sub and X-View, and Top1 was used as the evaluation metric.

[0136] The experimental parameters of the present invention are set as follows:

[0137] During training, continuous 300-frame human skeleton data is used as input. When the number of samples is less than 300 frames, the sample is repeatedly used for padding until 300 frames are reached.

[0138] During the training process, the SGD optimizer is adopted, and the hyperparameters α, β, γ, and δ in the loss function are set to 1, 1, 1, and 3 respectively. The learning rate is set to 0.01, and the learning rate is reduced by 10 times at the 10th and 20th epochs respectively. The batch size is set to 64, and a total of 30 epochs are trained.

[0139] The experimental environment of the present invention is as follows: the processor is Intel(R) Xeon(R) CPU E5-2603 v4 @ 1.70GHz, the graphics card is NVIDIA Titan XP 12GB, the memory is 64GB, the operating system is Ubuntu 16.04 (64-bit), the programming language is Python3.7.4, and the deep learning framework is PyTorch1.2.0.

[0140] The experimental results are shown in Table 1. It can be seen that the method proposed by the present invention outperforms the existing methods in terms of the metrics on both benchmarks, which verifies the effectiveness of the proposed method.

[0141] Table 1 Comparison results on the NTU-RGB+D dataset

[0142] Method Name X-Sub X-View 2S-AGCN[1] 88.5 95.1 PR-GCN[2] 85.2 91.7 PL-GCN[3] 89.2 95.0 The method proposed by the present invention 89.3 96.0

[0143] References:

[0144] [1] Shi L, Zhang Y, Cheng J, et al. Two-stream adaptive graph convolutional networks for skeleton-based action recognition[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019: 12026-12035.

[0145] [2] Li S, Yi J, Farha Y A, et al. Pose Refinement Graph Convolutional Network for Skeleton-Based Action Recognition[J]. IEEE Robotics and Automation Letters, 2021, 6(2): 1028-1035.

[0146] [3] Huang L, Huang Y, Ouyang W, et al. Part-Level Graph Convolutional Network for Skeleton-Based Action Recognition[C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2020, 34(07): 11045-11052.

Claims

1. A behavior recognition method based on a multi-stream fusion graph convolutional network, characterized in that It is divided into three stages: data preprocessing, feature extraction, and feature fusion; among which: In the data preprocessing stage, the input human skeleton sequence data, i.e., joint point data, is processed using the skeleton normalization module to obtain normalized human skeleton sequence data. Then, bone data is obtained by calculating the vectors formed between adjacent joint points, and motion data is obtained by calculating the displacements of the same joint point between adjacent frames. Among them: The human skeleton sequence data is represented as T represents the length of the skeleton sequence, T is taken as 300, x t ∈R V×C represents the joint point coordinates of the t-th skeleton, V represents the number of joint points in the human skeleton, V = 14, C represents the dimension of the joint point coordinates, C = 3, indicating that each joint point has three coordinates x, y, and z; Among them, the human joint numbers and their meanings are as follows: 0: Neck; 1: Head; 2: Right shoulder; 3: Right elbow; 4: Right wrist; 5: Left shoulder; 6: Left elbow; 7: Left wrist; 8: Right hip; 9: Right knee; 10: Right ankle; 11: Left hip; 12: Left knee; 13: Left ankle; In the feature extraction stage, the spatio-temporal graph convolutional network is used to extract the spatio-temporal features of joint point data, bone data, and motion data respectively, obtaining joint point features, bone features, and motion features; In the feature fusion stage, the multi-stream feature fusion network is used to further fuse the joint point features, bone features, and motion features, and then the prediction result of the behavior is obtained through the classifier; the method of using the multi-stream feature fusion network for fusion is as follows: In the first stage, first splice the three features in pairs, input the spliced features into two consecutive graph convolutional units to fuse the features of two types of information; then, input the fused features into the pooling layer; In the second stage, two fully connected layers are connected after the pooling layer, and there is a ReLU layer in the two fully connected layers to obtain three classification features f0, f1, and f2, and then fuse the three-way features to obtain the overall classification feature f3, f3 = f0 + f1 + f2; among which: In the feature extraction stage, the spatio-temporal graph convolutional network is used to extract the spatio-temporal features of joint data, bone data, and motion data. The implementation steps of the spatio-temporal graph convolutional network are as follows: (1) Construct the spatio-temporal graph of the human skeleton The construction of the spatio-temporal graph of the human skeleton is divided into three steps: (a) For the skeleton sequence X 3 and the set H of adjacent joint points in human physiology, the definition of H is as follows. For each x 3 in X t 3 , connect its adjacent joint points in physiology to obtain partial spatial edges, thereby constructing a local connection graph; H={(0,1),(0,2),(0,5),(2,3),(3,4),(5,6),(6,7), (8,9),(9,10),(11,12),(12,13)} (b) Given a set M, which is a set of physiologically non - adjacent but closely related joint points, and is defined as follows. For a given skeleton sequence X 3 For each in it, establish edges according to M to obtain a global connection graph; combine it with the local connection graph obtained in step (a) to form a skeleton space graph G S ={V, E S}}, where V represents the set of joint points, V = {v t,i |t = 1…T, i = 0…N - 1}, T is the length of the skeleton sequence, N is the number of joint points in the skeleton, and E S is the set of spatial edges, E S ={(v t,i v t,j )|(i, j) ∈ U}, and U is the union of H and M; M={(1,4),(1,7),(4,7),(4,13),(4,10),(7,10),(7,13),(10,13)} (c) For the skeleton spatial graph G obtained in step (b) S , establish time edges between the same joint points in the skeleton spatial graphs between adjacent frames to obtain the set E of time edges T , E T = {(v t,i v t+1,i ) | t = 1…T - 1, i = 0…N - 1}, thereby obtaining the skeleton time graph G T = {V, E T}, and finally obtaining the skeleton spatio - temporal graph G = {V, E}, where E = {E S , E T}, G = {G S , G T}; (2) Spatio-temporal graph convolution Perform spatio-temporal graph convolution on the spatio-temporal graph of the human skeleton obtained in step (1). The graph convolution in space is implemented by ST-GCN, and two adaptive graphs proposed in 2S-AGCN are introduced. The graph convolution in time is implemented by a 9×1 one-dimensional convolution; In the feature extraction stage, the spatio-temporal graph convolutional network is stacked by a batch normalization BN layer and six consecutive spatio-temporal graph convolutional units; each spatio-temporal graph convolutional unit has the same structure, including a spatial graph convolution GCN-S, a BN layer, a ReLU layer, a Dropout layer, a temporal graph convolution GCN-T, a BN layer, a ReLU layer, and a residual connection.

2. The behavior recognition method based on the multi-stream fusion graph convolutional network according to claim 1, wherein The skeleton normalization module in the data preprocessing stage proposes a skeleton normalization method, which includes three processing steps: position normalization, view normalization, and scale normalization, specifically as follows: (1) Position normalization First, perform position normalization on the input skeleton sequence, that is, given a human skeleton sequence where, x t represents the t-th skeleton in the sequence, T represents the length of the sequence, and update the coordinates of all joint points according to the following formula: Among them, x t,i represents the coordinate of the i-th joint point of the skeleton x t , where i = 0, 1, …, 13. Denote the skeleton sequence after position normalization as X 1 . In the above formula, is the coordinate of the i-th joint point of the t-th skeleton 1 of X ; (2) View normalization Then, perform a rotation transformation on the skeleton sequence X after position normalization 1 Specifically, first determine the rotation matrix R based on the first skeleton x1 of the sequence X. The formula is as follows: Among them, the vector v x , v y , v z are determined by x1 and are calculated as follows: (a) Determine the horizontal direction vector v based on the 2nd joint and the 5th joint of x1 x : v x = x 1,5 - x 1,2 (b) Determine v according to the following formula y : where v 1,0 represents the vector from joint point No. 1 to joint point No. 0 in the skeleton x1, that is: v 1,0 = x 1,1 -x 1,0 denote v 1,0 on v x projection; (c) Obtain v x and v y After that, calculate the vector v perpendicular to these two vectors according to the following formula z : v z = v x × v y Then rotate the coordinates of all joint points in X1 according to the following formula: Among them, The coordinates of the j-th joint point, where j = 0, 1, …, 13. Denote the skeleton sequence after perspective normalization as X 2 , and in the above formula is the j-th joint point coordinate of the t-th skeleton of X 2 ; ​ (3) Scale normalization Finally, perform scale normalization on the skeleton sequence X 2 , first scale the distance between joint points 0 and 1 to 1, that is, calculate the scaling factor r according to the following formula: Then update X according to the following formula 2 the coordinates of all the joint points in Denote the skeleton sequence after scale normalization as X 3 , where the in the above formula is X 3 's t-th skeleton 's k-th joint coordinate.

3. The behavior recognition method based on a multi-stream fusion graph convolutional network according to claim 1, wherein In the feature fusion stage, the method of designing the loss function in the multi-stream feature fusion network is as follows: First, use the softmax classifier to process the four features f0, f1, f2, and f3 to obtain their predicted probability values, which are p0, p1, p2, and p3 respectively, and then construct the loss function as: L = αL0 + βL1 + γL2 + δL3 Among them, L0, L1, L2, and L3 are the losses corresponding to each type of feature respectively, Among them, c represents the number of behaviors; y represents the true label of the sample, and α, β, γ, and δ are the weights of each type of loss respectively.

Citation Information

Patent Citations

  • Method for constructing human body behavior recognition model based on graph convolution network

    CN111652124A