A Multi-View Skeleton Sequence Fusion Method Based on Graph Convolution

By introducing a multi-view angle skeleton sequence fusion method based on graph convolution in multi-view angle action recognition, combining the natural topological relationship of the human body and the correspondence between the nodes between the perspectives, the problems of low accuracy of multi-view angle action recognition and unbalanced input in the prior art are solved, and higher recognition accuracy and model flexibility are achieved.

CN115497164BActive Publication Date: 2025-06-27TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211157830.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-22
Publication Date
2025-06-27
Estimated Expiration
2042-09-22

AI Technical Summary

Technical Problem

When processing multi-viewing skeleton sequences, the existing multi-viewing action recognition method ignores the spatial information of the skeleton sequence, resulting in low recognition accuracy and imbalance in the input of the multi-branch model, resulting in low overall accuracy.

Method used

A multi-view skeleton sequence fusion method based on graph convolution is proposed. Through data enhancement and multi-branch network structure, the multi-view fusion graph is constructed by combining the natural topological relationship of the human body and the correspondence between the nodes between the perspectives, and the multi-view fusion graph is fusionized through graph convolution, and finally the loss function based on deviation weighting is used for optimization.

Benefits of technology

The recognition accuracy of multi-view skeleton sequence fusion is improved, the problem of input imbalance of multi-branch models is solved, and the flexibility and cross-domain performance of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115497164B_ABST
    Figure CN115497164B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-view skeleton sequence fusion method based on graph convolution, comprising the following steps: performing data augmentation, adjusting it into a multi-view skeleton sequence as the input of a multi-branch network; using a spatio-temporal graph convolution network in each branch to extract the time-domain graph integration features of each view, extracting some end points and connection points from the human skeleton as joint points, and performing segmentation in the spatial dimension to obtain the time-domain graph integration representation of each joint point; constructing a multi-view fusion graph by combining the natural topological relationship of the human body, the corresponding relationship of joint points between views, and the graph integration features of joint points in the integration space; performing graph convolution according to the multi-view fusion graph to fuse features and obtain the graph integration features of multi-view fusion; and jointly predicting by multiple branches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence and computer vision, relates to feature fusion technology, and specifically provides a multi-view skeleton sequence fusion method based on graph convolution. Background Art

[0002] Multiple cameras can simultaneously capture the same action executor from different perspectives, thus providing supplementary information for many important visual tasks (such as human-computer interaction). In this case, an important issue is multi-view complementarity, which aims to use a multi-camera system to make up for occlusions and missing parts that may occur in a single view.

[0003] The background art related to the present invention includes:

[0004] (1) Adaptive Graph Convolutional Network (Reference [1]): Most existing works usually use predefined graphs for graph convolution. However, in action recognition tasks, the natural connection relationships of the human body may not be the most suitable edges. In addition, neural networks are hierarchical, and the optimal graph corresponding to each layer of graph convolution may not be the same. Therefore, the present invention performs adaptive graph convolution by combining the natural topological relationship of the human body, the corresponding relationship of joint points between perspectives, and the similarity of joint points in the integrated spatial features.

[0005] (2) Feature Fusion (Reference [4]): Most existing feature fusion methods follow the following process. First, the features to be fused are preprocessed to unify the feature space. Next, common fusion strategies such as maximum fusion, average fusion, and bilinear pooling are used for fusion. Although this process is general, it ignores the spatial information of the skeleton sequence. Therefore, on the basis of following the traditional process at the decision layer, the present invention introduces a graph convolutional network at the intermediate layer and combines spatial information, i.e., joint point-level information, for fusion.

[0006] (3) Multi-objective Optimization (Reference [2]): Most multi-view action recognition methods adopt a multi-branch structure. Each branch takes the skeleton sequence under a specified perspective as input and jointly optimizes the loss functions of different branches. Grid search aims to find a fixed optimal weight for each loss, but its transferability is very poor. Another widely used strategy is automatic weighted loss, where the weights of each loss are jointly optimized by learning. However, this is task-independent and more suitable for losses of different scales. In addition, data loss in certain perspectives is common, which can lead to imbalance among different branches. Therefore, the present invention proposes a multi-view data augmentation method and a deviation-weighted multi-loss function to solve the above problems.

[0007] (4) Similarity calculation method: In machine learning, the similarity between two targets is often evaluated by measuring the distance between samples. Common similarity measurement methods include Euclidean distance, cosine similarity, Hamming distance, Manhattan distance, etc. The present invention uses Euclidean distance as the measurement method for the similarity between integrated space features, thereby constructing a similarity adjacency matrix.

[0008] References

[0009] [1] Shi L, Zhang Y, Cheng J, et al. Skeleton-based action recognition with multi-stream adaptive graph convolutional networks[J]. IEEE Transactions on Image Processing, 2020, 29: 9532 - 9545.

[0010] [2] Alex Kendall, Yarin Gal, and Roberto Cipolla, “Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7482–7491.

[0011] [3] Le Zhang, Zenglin Shi, Ming-Ming Cheng, Yun Liu, Jia-Wang Bian, Joey Tianyi Zhou, Guoyan Zheng, and Zeng Zeng, “Nonlinear regression via deep negative correlation learning,” IEEE transactions on pattern analysis and machine intelligence, 2019.

[0012] [4] Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman, “Convolutional two-stream network fusion for video action recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 1933–1941.

[0013] [5] Azis N A, Jeong Y S, Choi H J, et al. Weighted averaging fusion for multi-view skeletal data and its application in action recognition[J]. IET Computer Vision, 2016, 10(2): 134-142. Summary of the Invention

[0014] The objective of the present invention is to provide a multi-view skeleton sequence fusion method that can improve recognition accuracy. The technical solution is as follows:

[0015] A multi-view skeleton sequence fusion method based on graph convolution includes the following steps:

[0016] Step 1: Perform data augmentation and adjust it into a multi-view skeleton sequence as the input of a multi-branch network;

[0017] (1) Combine the skeleton sequences of a single action sequence at all viewpoints into a multi-view skeleton sequence;

[0018] (2) Rearrange the obtained multi-view skeleton sequence at the viewpoint level;

[0019] (3) Use the multi-view skeleton sequence after viewpoint-level rearrangement as the input of the multi-branch network, and each branch receives all viewpoints of the samples at different times.

[0020] Step 2: Use a spatio-temporal graph convolutional network in each branch to extract the temporal graph integration features X1, X2,..., X n ;

[0021] (1) Select a spatio-temporal graph convolutional network with the same structure for each branch;

[0022] (2) Use the selected spatio-temporal graph convolutional network in each branch to extract the temporal graph integrated feature X of the input perspective. The method is as follows: Use the selected spatio-temporal graph convolutional network in each branch to input the multi-perspective skeleton sequence after perspective-level rearrangement. Where c1 is the number of channels of the skeleton sequence, t is the number of frames of the skeleton sequence, i.e., the time dimension, and v is the number of joint points of the skeleton sequence, i.e., the space dimension; Take the average of the time dimension while retaining the space dimension to extract the temporal graph integrated features X1, X2,..., X of each perspective. n , and form the temporal graph integrated feature of the entire sequence. Where c2 is the number of channels of the temporal graph integrated feature X.

[0023] (3) Extract some endpoint and connection points from the human skeleton as joint points, and segment them in the space dimension to obtain the temporal graph integrated representation X of each joint point. :,j ;

[0024] Step 3: Construct a multi-perspective fusion graph M by combining the natural topological relationship of the human body, the corresponding relationship of joint points between perspectives, and the graph integrated features of joint points in the integrated space.

[0025] (1) The connection between joint points is the natural topological relationship of the human body; Combine the natural topological relationship of the human body and the corresponding relationship of joint points between perspectives to construct an adjacency matrix A.

[0026] (2) Perform Laplace transform on the adjacency matrix A to obtain the natural connection graph N.

[0027] (3) Construct the corresponding similarity graph R according to the graph integrated representation of each extracted joint point.

[0028] (4) Both the natural connection graph N and the similarity graph R are stored in the form of an adjacency matrix, and the matrix element value represents the strength of the connection between nodes. Normalize the natural connection graph N and the similarity graph R at the matrix level respectively, and then sum the elements at the corresponding positions with weights. Weightedly fuse the natural connection graph N and the similarity graph R to obtain the multi-perspective fusion graph M.

[0029] Step 4: Perform graph convolution according to the multi-perspective fusion graph M to fuse the features X1, X2,…, X n Obtain the graph integrated feature X of multi-perspective fusion n+1 , and the method is as follows:

[0030] (1) Concatenate the temporal graph integrated features X1, X2,..., X of each extracted perspective in the space dimension to construct a feature vector X containing joint points under all perspectives. n ; n+1

[0031] (2) The feature vector X containing joint points under all perspectives.n+1 The input graph convolution network performs spatial-perspective graph convolution on the graph and the graph M, and obtains the graph integration feature Y of multi-perspective fusion n+1 ;

[0032] Step Five, multi-branch joint prediction, the method is as follows:

[0033] The obtained graph integration feature Y of multi-perspective fusion n+1 and the time-domain graph integration features X1, X2,..., X of each perspective extracted n are respectively subjected to global average pooling in the spatial dimension, and then input into n+1 independent fully connected layers to obtain the preliminary prediction of classification where class is the number of sample categories;

[0034] (2) Use the sigmoid function to map the data elements of the obtained preliminary prediction of classification to between (0,1);

[0035] (3) Perform decision-level fusion on the mapped preliminary prediction of classification by taking the maximum value as the final prediction result;

[0036] Step Six, if the loss function corresponding to the end-to-end network framework has not converged, repeat steps two to five iteratively.

[0037] Furthermore, the loss function corresponding to the end-to-end network framework is specifically:

[0038] (1) Use the cross-entropy loss function to constrain the difference between the prediction result and the true value of each branch, and use it as the loss function of each branch;

[0039] (2) The training samples of each branch are the same, so the commonly used automatic weighted loss function is not applicable; take the reciprocal of the prediction variance of each branch as the weight of the branch loss function;

[0040] (5) The loss function of the overall network is the weighted sum of the branch loss functions.

[0041] The beneficial effects of the technical solution provided by the present invention are as follows:

[0042] 1. When the present invention performs skeleton-based behavior recognition by fusing multi-perspective inputs, it uses the corresponding relationship of joint points under different perspectives in the graph integration space, and proposes spatial-perspective graph convolution starting from time-space graph convolution. Compared with traditional fusion methods, it is more intuitive, and the multi-perspective fusion features obtained are better for prediction. In addition, the spatio-temporal graph convolution network for extracting graph integration can be freely selected, greatly enhancing the flexibility of the model.

[0043] 2. The present invention proposes a multi - perspective data augmentation strategy and a supporting loss function to improve the overall accuracy of the model while solving the problem of unbalanced inputs of each branch in traditional multi - branch models.

[0044] 3. The present invention proposes an end - to - end trained multi - branch graph neural network model, which uses a graph model to fuse features from multiple perspectives and solves the problem of multi - perspective behavior recognition based on skeletons. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 : is the natural topological relationship of the human body (left figure) and the corresponding relationship of joint points between perspectives (right figure);

[0046] Figure 2 : is a case diagram of the problem of multi - perspective behavior recognition based on skeletons;

[0047] Figure 3 : is a diagram of a multi - temporal behavior recognition method based on skeletons;

[0048] Figure 4 : is a flowchart of a multi - temporal behavior recognition method based on skeletons;

[0049] Figure 5 : is a detailed diagram of multi - perspective graph convolution. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] Behavior recognition based on skeletons plays an important role in many applications of computer vision. The present invention studies the problem of fusing skeleton sequences from multiple perspectives simultaneously captured by different cameras for behavior recognition, that is, the fusion method of multi - perspective skeleton sequences. The present invention proposes a new multi - perspective skeleton sequence fusion method based on graph convolutional network to solve this problem. First, data augmentation is performed, and the view - level full permutation of the multi - perspective bone sequences is used as the input of the model. A spatio - temporal graph convolutional network is used to extract the time - domain graph integrated features of each joint point in each perspective at the single - perspective level, while compressing the time dimension and retaining the spatial dimension, that is, the joint points. Next, a natural connection graph is established using the natural topological relationship of the human body, and a corresponding relationship graph is calculated based on the corresponding relationship of joint points between perspectives based on the above integration. Combining the two, multi - perspective graph convolution is performed to obtain the graph integrated features of multi - perspective fusion. And it is jointly optimized with the graph integrated features at the single - perspective level for prediction. Considering the characteristics of this task, the present invention also proposes a data - driven deviation - based loss function to improve the training and prediction performance of the model. This fusion method improves the prediction accuracy of the existing action recognition methods based on skeletons and achieves good cross - domain performance.

[0051] The multi - perspective skeleton sequence fusion method based on graph convolution provided by the present invention improves the recognition accuracy by fusing the representations of the same action sequence from multiple perspectives at different stages, such as Figure 2As shown, the skeleton sequences from different perspectives will be fused and jointly used for action recognition.

[0052] The present invention models the problem of multi-view action recognition based on skeletons with any number of perspectives as a multi-objective optimization problem. Combining a multi-branch network, the present invention proposes a new graph convolution-based fusion method to solve this problem, and the process design is as Figure 4 shown. The overall structure of the model is as Figure 3 shown. For the sake of simplicity, only the case of two perspectives is shown, and the case of more perspectives can be extended on this basis. The first half of the network extracts the time-domain graph integrated features of each perspective through a spatio-temporal graph convolution network; in the second half of the network, an adaptive graph convolution is constructed by combining the natural topological relationship of the human body, the corresponding relationship of joint points between perspectives, and the similarity of joint points in the integrated space to obtain the graph integrated features of multi-view fusion and jointly optimize them with the single-view graph integrated features, so as to better solve the problem of action recognition based on multi-view skeleton sequences. The present invention can be used to utilize the action sequences captured by multiple cameras in the same scene, and fuse the skeleton sequences from different perspectives at multiple stages without restricting the number of perspectives, that is, multi-view skeleton sequence fusion.

[0053] The technical solution of the present invention is given below. The flowchart is as Figure 4 shown.

[0054] Step S1, perform data augmentation on the single-view skeleton sequence as the input of the multi-branch network.

[0055] Step S2, use a spatio-temporal graph convolution network in each branch to extract the time-domain graph integrated features X1, X2,..., X n .

[0056] Step S3, construct a multi-view fusion graph M by combining the topological relationship of the human body, the corresponding relationship of joint points between perspectives, and the similarity of joint points in the integrated space.

[0057] Step S4, perform multi-view graph convolution according to the multi-view fusion graph B to obtain the graph integrated features X n+1 of multi-view fusion.

[0058] Step S5, combine the single-view features X1, X2,..., X n and the multi-view fusion feature X n+1 for joint prediction.

[0059] Step S6, if the loss function corresponding to the end-to-end network framework has not converged, repeat steps two to five iteratively.

[0060] The specific implementation steps of the multi-view skeleton sequence fusion method based on graph convolution are as follows:

[0061] (1) Data augmentation:

[0062] Before the training starts, first adjust the single-view skeleton sequence to a multi-view skeleton sequence. The specific steps are as follows:

[0063] (1) Combine the skeleton sequences of a single action sequence at all viewpoints into a multi-view skeleton sequence.

[0064] (2) Rearrange the obtained multi-view skeleton sequence at the viewpoint level.

[0065] Explanation 1: Perspective-level rearrangement

[0066] The data used in the present invention includes the skeleton sequences obtained by photographing the same action sequence from multiple viewpoints at the same moment. On the premise of keeping the original temporal and spatial order of the skeleton sequence unchanged, the input branch of the single-view skeleton sequence is transformed.

[0067] (3) Use the full permutation of the multi-view skeleton sequence in the viewpoint dimension as the input of the network, that is, each branch receives all viewpoints of the samples at different times as input, so as to balance the input of each branch.

[0068] (2) Extraction of integrated features of the temporal domain graph of a single viewpoint: Use a spatio-temporal graph convolutional network in each branch to extract the integrated features X1, X2,..., X of the temporal domain graph of each viewpoint n

[0069] The specific method for extracting the integrated features of the temporal domain graph is as follows:

[0070] (1) Select a spatio-temporal graph convolutional network with the same structure for each branch.

[0071] Explanation 2: Selection of spatio-temporal graph convolutional network

[0072] The spatio-temporal graph convolutional network can be regarded as a feature extraction network. The current mainstream method in the field of skeleton-based action recognition is the spatio-temporal graph convolutional neural network. The skeleton sequence is regarded as a spatio-temporal graph containing temporal connections and spatial connections, and temporal dimension convolution and spatial dimension convolution are alternately performed. The graph convolutional neural network can be used as the spatio-temporal graph convolutional network to independently select the retained dimensions and specifications.

[0073] (2) Input the single-view skeleton sequence into the selected spatio-temporal graph convolutional network in each branch where c1 is the number of channels of the skeleton sequence (the actual meaning is the two-dimensional or three-dimensional coordinates of the joint points), t is the number of frames of the skeleton sequence, that is, the temporal dimension, and v is the number of joint points of the skeleton sequence, that is, the spatial dimension. On the basis of retaining the spatial dimension, take the average of the temporal dimension to extract the integrated features of the temporal domain graph of the entire sequence where c2 is the number of channels of the integrated feature X of the temporal domain graph.

[0074] (3) Perform spatial dimension segmentation to obtain the integrated representation X of the time-domain graph for each joint point :,j 。

[0075] Explanation 3: Graph integrated representation of joint points

[0076] The spatial dimension of the features extracted by the spatio-temporal graph convolutional network is consistent with the original skeleton sequence specification, while the time dimension is compressed. Therefore, the feature vector at the specified position j can be regarded as the integrated time graph of the joint points.

[0077] (III) Obtain the multi-view fusion graph M: Construct the multi-view fusion graph M by combining the natural topological relationship of the human body, the corresponding relationship of joint points between views, and the graph integration features of joint points in the integrated space

[0078] The specific method for calculating the multi-view fusion graph M through the natural connection relationship and joint point similarity is as Figure 5 :

[0079] (1) Extract some end points and connection points from the human skeleton as joint points. The connections between joint points are the natural topological relationship of the human body, and there is also a corresponding relationship between the same joint points under different views. Construct the adjacency matrix A by combining the natural topological relationship of the human body and the corresponding relationship of joint points between views.

[0080] (2) Perform Laplace transform on the adjacency matrix A to obtain the natural connection graph N.

[0081] (3) Construct the similarity graph R between joint points according to the integrated representation of each joint point extracted.

[0082] (4) Normalize and then weighted fuse the natural connection graph N and the similarity graph R to obtain the multi-view fusion graph M.

[0083] Explanation 4: Weighted fusion of graphs

[0084] Both the natural connection graph N and the similarity graph R are stored in the form of an adjacency matrix in reference [1]. The matrix element value represents the strength of the connection between nodes. Normalize them at the matrix level respectively, and then sum the elements at the corresponding positions with weights. The weights are trainable parameters shared by the entire matrix. The natural connection graph N and the similarity graph R share one weight respectively on the whole graph, and the sum of their corresponding weights is 1, thus giving the graph greater flexibility.

[0085] (IV) Fusion of single-view features: Perform graph convolution according to the multi-view fusion graph M to fuse the features X1, X2, …, X n Obtain the graph integration features X of multi-view fusion n+1

[0086] The method for fusing single-view graph integration features is:

[0087] (1) Concatenate the extracted integrated features X1, X2,..., X n in the spatial dimension to construct a feature vector X containing joint points under all perspectives n+1 .

[0088] (2) For the purpose of suppressing the over-smoothing trend of the graph, input the output in (1), i.e., the feature vector X containing joint points under all perspectives n+1 and the multi-perspective fusion graph M into the graph convolutional block. The graph convolutional block can adopt the spatio-temporal graph convolutional network structure for feature extraction to perform spatial-perspective graph convolution and obtain the graph integrated feature Y of multi-perspective fusion n+1 .

[0089] Explanation 5: Spatial-perspective graph convolution

[0090] After being compressed in the time dimension, its position is replaced by the perspective dimension. The original skeleton sequence can be regarded as a complete time-space graph, and at this time it can be regarded as a spatial-perspective graph. The difference is that the spatial coordinates corresponding to the original graph joints are replaced by the graph integrated representation. Consider performing spatial-perspective graph convolution on it in the same way as spatio-temporal graph convolution. Since the number of nodes is small at this time, it can be directly convolved as a whole

[0091] (V) Multi-branch joint prediction and optimization

[0092] The method for joint prediction and optimization of the single-perspective branch and the fusion multi-perspective branch is as follows

[0093] (1) The obtained graph integrated feature Y of multi-perspective fusion n+1 and the extracted integrated features X1, X2,..., X n respectively perform global average pooling in the spatial dimension, and then input them into n + 1 independent fully connected layers to obtain the preliminary prediction of classification where class is the number of sample categories

[0094] (2) Use the sigmoid function to map the data elements of the obtained preliminary prediction of classification to between (0, 1)

[0095] (3) Perform decision-level fusion on the mapped preliminary prediction of classification by taking the maximum value as the final prediction result Say Note 6: Decision-level fusion

[0096] The fusion of the multi-branch model can be divided into data fusion, feature fusion, model fusion, and decision-level fusion according to the processing stage (new reference [5]). The present invention adopts decision-level fusion

[0097] The prediction results of each branch belong to the same feature space. Therefore, various traditional fusion methods can be adopted. Reference 4 compares various traditional fusion methods at different levels of the model.

[0098] The present invention adopts maximum value fusion, and takes the maximum value of the prediction values of the same category of each branch as the output.

[0099] Explanation 7: Construction of loss function

[0100] Step 1: Single-branch loss function. Considering that in step (i), through the rearrangement of the input data, each branch is trained on the same data in different orders. Therefore, the traditional cross-entropy loss function is adopted:

[0101]

[0102] where represents the output result of the network, and the value range of this value is (0, 1); y represents the true predicted value, which can only take 0 or 1.

[0103] Step 2: Deviation-based loss function. Reference 3 proposes that the generalization error of the ensemble model depends on the covariance between sub-models on the premise that the deviation is less than that of the sub-models. Since each branch can make independent predictions, the multi-branch model can be regarded as an ensemble model that integrates each branch. Although the same training data reduces the covariance between branches, it may also make the ensemble model in a local optimum. Therefore, the present invention constructs a deviation-based loss function.

[0104]

[0105] Suppose there are inputs from m perspectives, where w i is the weight of branch i, and the sum of the weights of each branch is 1. When i ∈ [1, m], is the prediction result of the single-perspective branch; when i = m + 1, is the prediction result of the multi-time fusion branch. y represents the true predicted value. The weight w i is negatively correlated with the deviation of each branch. The variance loss function is used to measure the deviation of the branch. Let In this way, the branch with a higher fitting degree has a larger learning rate, and at the same time effectively suppresses the changes in other branches. On the premise of ensuring the basic accuracy, the deviation-based loss function can make each branch easier to converge to the global optimum, thereby preventing the model from overfitting and ultimately improving the overall accuracy of the model.

[0106] (VI) Determine whether the model training is completed

[0107] The specific method for determining whether the model training is completed is as follows:

[0108] During the training process of a neural network, it is possible to determine whether to stop training based on the value of the loss function. Training can be stopped when the loss function drops to a certain level and remains basically unchanged.

Claims

1. A multi-view skeleton sequence fusion method based on graph convolution, comprising the following steps: Step 1, perform data augmentation and adjust it into a multi-view skeleton sequence as the input of a multi-branch network; (1) Combine the skeleton sequences of a single action sequence under all views into a multi-view skeleton sequence; (2) Rearrange the obtained multi-view skeleton sequence at the view level; (3) Use the multi-view skeleton sequence after view-level rearrangement as the input of the multi-branch network, and each branch receives all views of the samples at different times. Step 2: Use a spatio-temporal graph convolutional network in each branch to extract the time-domain graph integrated features X1, X2, ..., X of each perspective n ; (1) Select a spatio-temporal graph convolution network with the same structure for each branch; (2) Extract the time-domain graph integrated feature X of the input perspective using the selected spatio-temporal graph convolutional network in each branch. The method is as follows: Use the selected spatio-temporal graph convolutional network in each branch to input the multi-perspective skeleton sequence after perspective-level rearrangement. Where c1 is the number of channels of the skeleton sequence, t is the number of frames of the skeleton sequence, i.e., the time dimension, and v is the number of joint points of the skeleton sequence, i.e., the space dimension; Take the average of the time dimension while retaining the space dimension to extract the time-domain graph integrated features X1, X2,..., X of each perspective. n to form the time-domain graph integrated feature of the entire sequence. Where c2 is the number of channels of the time-domain graph integrated feature X. (3) Extract some end points and connection points from the human skeleton as joint points, and segment them in the spatial dimension to obtain the time-domain graph integration representation of each joint point; Step 3, construct a multi-view fusion graph M by combining the natural topological relationship of the human body, the corresponding relationship of joint points between views, and the graph integration features of joint points in the integrated space; (1) The connection between joint points is the natural topological relationship of the human body; construct an adjacency matrix A by combining the natural topological relationship of the human body and the corresponding relationship of joint points between views; (2) Perform Laplace transform on the adjacency matrix A to obtain the natural connection graph N; (3) Construct the corresponding similarity graph R according to the graph integration representation of each extracted joint point; (4) Both the natural connection graph N and the similarity graph R are stored in the form of an adjacency matrix, and the matrix element value represents the connection strength between nodes. Normalize the natural connection graph N and the similarity graph R at the matrix level respectively, and then perform weighted summation of the elements at the corresponding positions. Weightedly fuse the natural connection graph N and the similarity graph R to obtain the multi-view fusion graph M; Step 4: Perform graph convolution according to the multi-view fusion graph M to fuse the features X1, X2, …, X n Obtain the graph integrated feature X of multi-view fusion n+1 , the method is as follows: (1) Integrate the time-domain graph integrated features X1, X2,..., X of each perspective extracted n Concatenate in the spatial dimension to construct a feature vector X containing joint points under all perspectives n+1 ; (2) Input the feature vector X containing the joint points from all perspectives n+1 and graph M into the graph convolutional network for spatial-perspective graph convolution to obtain the graph integrated feature Y with multi-perspective fusion n+1 ; Step 5, multi-branch joint prediction, the method is as follows: Integrate the obtained multi-view integrated graph feature Y n+1 and the time-domain graph integrated features X1, X2,..., X of each extracted view n Perform global average pooling in the spatial dimension respectively, and then input them into n+1 independent fully connected layers to obtain the preliminary prediction of classification where class is the number of sample classes; (2) Use the sigmoid function to map the data elements of the obtained preliminary prediction of classification to between (0, 1); (3) Perform decision-level fusion on the mapped preliminary prediction of classification by taking the maximum value as the final prediction result; Step 6, if the loss function corresponding to the end-to-end network framework does not converge, repeat steps 2 to 5 iteratively.

2. The multi-view skeleton sequence fusion method according to claim 1, wherein The loss function corresponding to the end-to-end network framework is specifically: (1) Use the cross-entropy loss function to constrain the difference between the prediction result and the true value of each branch, and use it as the loss function of each branch; (2) The training samples of each branch are the same, so the commonly used automatic weighted loss function is not applicable; take the reciprocal of the prediction variance of each branch as the weight of the branch loss function; (3) The overall loss function of the network is the weighted sum of the branch loss functions.

Citation Information

Patent Citations

  • Human skeleton reconstruction method based on double-view-angle Kinect joint point fusion

    CN110458944A

  • Multi-modal feature fusion method based on LSTM network

    CN111461166A