Two-person interaction behavior recognition method based on time decoupling hierarchical graph convolution
By using the time-decoupled hierarchical graph convolution method in the recognition of two-person interactive behaviors, the TB-GCN graph convolution network is constructed, which solves the problem of insufficient modeling of behavior time continuity in the prior art, and improves the accuracy and efficiency of recognition.
Patent Information
- Application Number
- CN202510155142.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The existing two-person interactive behavior recognition algorithm based on node data does not model the time continuity of behavior, cannot effectively capture subtle changes in time, and has poor real-time and low accuracy.
The method based on time decoupling hierarchical graph convolution is adopted to construct a hierarchical graph topology through optimized joint node data, and THGC graph convolution blocks and TBGC graph convolution blocks are used, and TB-GCN graph convolution network is constructed, deep feature extraction and classification are used to identify the two-person interactive behavior.
Effectively dealing with insufficient time continuity modeling for interactive behavior recognition improves the accuracy of complex interactive situations, reduces recognition time, and improves the efficiency of two-person interactive behavior recognition.
Smart Images

Figure CN120071441A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a method for recognizing dual-person interaction behaviors based on time-decoupled hierarchical graph convolution. Background Art
[0002] With the rapid development of intelligent devices and sensing technologies, the recognition of dual-person interaction behaviors based on skeleton data has become increasingly important in many application scenarios, especially in the fields of monitoring, human-computer interaction, virtual reality, and sports analysis. The skeleton data of the human body is captured by a depth camera or an inertial measurement unit, which can effectively reflect the key points and movement trajectories of the human body, providing rich information for understanding and analyzing human interaction behaviors.
[0003] Skeleton data can effectively represent the relative positions and angular relationships between individual joints, which is crucial for understanding an individual's pose and the interaction between people. Lee et al. proposed a hierarchical decomposition graph convolutional network (HD-GCN) architecture, which introduced the concept of a hierarchical decomposition graph (HD-Graph). HD-GCN effectively divides each joint node into multiple sets to extract the main structural adjacency edges and long-range edges, and uses these edges to construct an HD-Graph that contains them. These edges are in the same semantic space of the human skeleton. A hierarchical aggregation attention module is introduced to highlight the integration of key hierarchical edges in the hierarchical decomposition graph. Although such HD-GCN effectively extracts the local features of points, HD-GCN needs to process complex graph structures, especially in the hierarchical decomposition of human skeleton models, where the computational cost will increase significantly, resulting in high time costs for training and inference. Liu et al. proposed a temporal decoupling graph convolutional network (TD-GCN), which applies different adjacency matrices to skeletons from different frames. In each convolutional layer of TD-GCN, high-level spatio-temporal features are extracted from the skeletal data to obtain deep spatio-temporal information. Channel-dependent and time-dependent adjacency matrices corresponding to different channels and frames are calculated to capture the spatio-temporal dependencies between skeleton joints. The spatio-temporal features of skeleton nodes and the topological information of adjacent skeleton nodes are fused using the channel-dependent and time-dependent adjacency matrices. TD-GCN captures spatio-temporal features through channel and time dependencies, but key information is omitted when designing and selecting the adjacency matrix, affecting the recognition effect. Zhou et al. proposed to use the advantages of graph distance to describe physical topology, thereby encoding bone connections. By combining persistent homology analysis with topology representations of specific actions to describe system dynamics, important topological nuances that are often lost in traditional graph convolutional networks are effectively retained. The redundancy problem in multi-relationship modeling of existing traditional graph convolutional networks is revealed, and an effectively improved graph convolutional method (BlockGC) is proposed to solve the redundancy problem. Although this BlockGC reduces the number of parameters, under the condition that the training data is relatively dense, BlockGCN will exhibit overfitting. Li et al. input the two-person joint point data as a whole into the network, constructed a two-person graph that can retain the basic interaction relationship, and proposed four manual labeling strategies to further enhance the interaction relationship between two people. Due to the significant differences in the connection relationships between different actions, the model is difficult to adapt to complex interaction scenarios.
[0004] Generally speaking, the existing two-person interactive behavior recognition algorithms based on joint point data have poor adaptability and insufficient modeling of the temporal continuity of behaviors, and cannot effectively capture the subtle changes that occur over time. Manually designed features often cannot comprehensively capture the interaction relationship and miss key dynamic information. The existing methods run slowly, resulting in the two-person interactive behavior recognition algorithm based on joint point data still being a challenging research topic.
[0005] In summary, it is very meaningful to provide a two-person interaction behavior recognition method applicable to recognizing complex interaction behaviors. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a two-person interaction behavior recognition method based on time-decoupled hierarchical graph convolution to solve the problems of insufficient modeling of the temporal continuity of behaviors in existing recognition methods, inability to effectively capture subtle action changes occurring in time, and poor real-time performance and low accuracy in interaction behavior recognition.
[0007] The technical solution of the present invention is: A two-person interaction behavior recognition method based on time-decoupled hierarchical graph convolution, including:
[0008] S1: Using optimized joint point data as the input of the network, constructing a hierarchical graph topology to obtain a hierarchical relationship matrix;
[0009] S2: Constructing the CTR-GC channel topology optimized graph convolution into a time-decoupled hierarchical graph convolution, and using the hierarchical relationship matrix obtained in S1 to construct a THGC graph convolution block that can simultaneously capture time channel information and space channel information;
[0010] S3: Constructing the graph convolution in THGC into BLOCKGC, and further constructing it into a TBGC graph convolution block;
[0011] S4: Stacking the TBGC graph convolution block and the multi-scale time convolution block to obtain a TB-GCN graph convolution network;
[0012] S5: Using the TB-GCN graph convolution network to perform deep feature extraction and classification on the input skeleton data of the human body to obtain the two-person interaction behavior recognition result.
[0013] Preferably, the optimized joint point data in S1 is the joint point data after removing noise and lost frames, and the optimization processing method is: obtaining two-person joint point data through the depth camera Kinect v2, performing denoising and filtering lost frames on the obtained joint point data, using three strategies of calculating the frame length, dispersion degree and motion value of the node data to perform denoising and lost frame filtering on the two-person joint point data, taking the abdominal node as the center, constructing a rooted tree, dividing the human body joints into five layers, and then dividing the graph so that the nodes with concentrated edges in the same hierarchical structure are stored in the same semantic space.
[0014] Preferably, the THGC graph convolutional block constructed in S2 consists of a residual structure composed of two branches and a 1×1 convolution. One branch further extracts more detailed spatial feature information through a 1×1 convolution and a temporal pooling module, processes the obtained spatial features using self-contrast subtraction and the tanh activation function, and adjusts the number of channels through a 1×1 convolution to obtain spatial features;
[0015] The other branch consists of a bottleneck block composed of two flow branches. One flow extracts more detailed temporal feature information through a spatial pooling module, processes the obtained temporal features using self-contrast subtraction and the tanh activation function, and the other flow adjusts the number of channels through a 1×1 convolution to retain its original feature information. The features of the two flows are fused to obtain temporal features. The obtained temporal features are element-wise multiplied and added to the spatial features obtained in the previous branch and the weighted hierarchical relationship matrix to obtain spatio-temporal features. The features that fuse time and space are fused with their own features after passing through a 1×1 convolution to form a time-decoupled hierarchical graph convolutional block.
[0016] Preferably, in S3, the graph convolution in THGC is constructed as BLOCKGC, and then constructed into a TBGC graph convolutional block, including: processing the input data by normalizing the weights, introducing relative position encoding into the graph convolution, generating k-hop graph distance encoding for the adjacency matrix and calculating the weights, performing batch normalization and using a 1×1 convolution as the residual connection of the downsampling layer, and outputting the result through the ReLU activation function.
[0017] The present invention provides a method for recognizing dual-person interaction behaviors based on time-decoupled hierarchical graph convolution. This method can effectively address the problem of insufficient modeling of the time continuity in interaction behavior recognition, solve the problem of poor real-time performance in recognizing interaction behaviors with a large amount of data, and can effectively improve the accuracy in complex interaction scenarios, greatly improving the efficiency of dual-person interaction behavior recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is the overall framework of a method for recognizing dual-person interaction behaviors based on time-decoupled hierarchical graph convolution provided by the disclosed embodiments of the present invention;
[0021] Figure 2 Schematic diagram of the multi-scale temporal convolution block structure provided by the disclosed embodiments of the present invention;
[0022] Figure 3 Schematic diagram of the TBGC graph convolution block structure provided by the disclosed embodiments of the present invention;
[0023] Figure 4 Schematic diagram of the hierarchical topology structure provided by the disclosed embodiments of the present invention;
[0024] Figure 5 Overall structure diagram of the TB-GCN graph convolution network provided by the disclosed embodiments of the present invention;
[0025] Figure 6 Confusion matrix of the results of the NTU RGB+D interaction data provided by the disclosed embodiments of the present invention. Detailed implementation manners
[0026] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of the system consistent with some aspects of the present invention as detailed in the appended claims.
[0027] In view of the problems existing in the method for recognizing dual-person interaction behaviors based on joint point data in the prior art, this implementation provides a method for recognizing dual-person interaction behaviors based on temporally decoupled hierarchical graph convolution to effectively improve the accuracy of recognizing dual-person interaction behaviors. The overall framework flowchart is as Figure 1 shown. Specifically, it includes the following steps:
[0028] S1: Use the joint point data as the network input, process the noise and frame loss in the input joint point data to obtain optimized joint point data as the network input, and construct a hierarchical graph topology through the optimized joint data to obtain a hierarchical relationship matrix that can simultaneously capture the natural and non-natural connections of the human body;
[0029] Among them, processing the noise and frame loss in the input joint point data to obtain optimized joint point data includes: obtaining dual-person joint point data by using the depth camera Kinect v2, calculating the frame length of the joint point data and filtering out the shorter skeleton data, identifying the dispersion degree of the joint points on the plane and removing the frames with excessive noise, calculating the motion value and screening out the skeleton data with the motion amplitude not within the predetermined range, so as to obtain optimized joint point data:
[0030] Use the optimized joint point data as the input of the time-decoupled hierarchical graph convolutional network model, and construct a hierarchical graph topology using the optimized joint data. Among them, the specific steps for constructing the hierarchical graph topology are as follows:
[0031] 1) First, decompose the graph through the physical connection edges of the joint points to construct a rooted tree, and decompose the given joint points into the rooted tree;
[0032] 2) Taking the abdominal node as the centroid, assign the nodes concentrated on the same hierarchical structure edges to the same semantic space, and define the directed adjacency matrix of the H-level edge set with L levels as As shown in formula (1), where H k represents the k-th level node set, and ε(H k →H k+1 ) represents a set of edges from H k to H k+1 ;
[0033]
[0034] 3) Apply the fully connected edges to the rooted tree to construct an adjacency matrix with a larger receptive field and far connectivity significance The adjacency matrix As shown in formula (2), the constructed hierarchical edge set is as shown in formula (3), where H k ∪H k+1 represents the union of the k-th level node set and the k-th level node set;
[0035]
[0036] ε k =ε(H k ∪H k+1 ,H k →H k+1 ,H k+1 →H k ) (3)
[0037] The input joint point features are in the form of three-dimensional coordinates, with a size of C*T*2N, where V is the number of human joint points and N is the number of joint points. Treat the two action performers as a whole and send them into the network. The number of video frames is set to 120 frames, and the size of the joint point input data is 3*300*50. Normalize the obtained hierarchical set adjacency matrix with the degree matrix to ensure the stability of training, and use all elements of the matrix as learnable parameters to ensure the adaptability of training.
[0038] S2: Construct the CTR-GC channel topology optimized graph convolution into a time-decoupled graph convolution, as Figure 2As shown in the figure, the relationship matrix obtained in S1 is used to construct a THGC graph convolution block that can capture both temporal channel information and spatial channel information;
[0039] 1) To obtain more spatio-temporal feature information, the joint point data is used as input, which is transformed into a relationship matrix in S1 and refined into three parts. The first branch extracts spatial feature information through a spatial channel module; the second branch extracts temporal feature information through a temporal decoupling module; the third branch is a branch using 1×1 convolution to obtain a residual connection of its own features;
[0040] 2) The input feature of the first branch is C / r*T*2N. After passing through a 1×1 convolution and a temporal average pooling module to extract spatial feature information, the feature shape becomes C / r*2N. The obtained spatial features are adjusted to the shape of C / r*2N*2N using self-contrast subtraction, and the output values are adjusted to the range of -1 to 1 using the tanh activation function. Then, a 1×1 convolution is used to adjust the feature shape to C'*2N*2N;
[0041] 3) The input feature of the second branch is C'*T*2N. After passing through a 1×1 convolution and a temporal decoupling module to extract spatial feature information, the feature shape becomes T*2N. The obtained spatial features are adjusted to the shape of T*2N*2N using self-contrast subtraction, and the output value range is adjusted using the tanh activation function. The obtained features are fused with the self-features after 1×1 convolution to obtain temporal feature information;
[0042] 4) The features obtained from the first branch and the second branch are subjected to an element-wise addition operation with the hierarchical relationship matrix constructed in step 1 to obtain spatio-temporal features. The obtained spatio-temporal features are fused with the self-features obtained from the branch using 1×1 convolution to obtain the output spatio-temporal features;
[0043] Compared with the original CTR-GC module, the improved new module uses a temporal decoupling method to obtain rich temporal features, uses different methods in two branches to extract temporal and spatial joint point features, and adopts a residual network structure in the second branch. Secondly, the temporal decoupling convolution block adds the features passing through the temporal channel and the spatial channel in each layer element-wise with the hierarchical relationship matrix and sends them to the next layer for fusion, so that the final output contains a feature map with temporal and spatial information.
[0044] S3: The graph convolution in THGC is constructed as BLOCKGC, and then further constructed as a TBGC graph convolution block;
[0045] Specifically,
[0046] 1) The input data is processed by normalizing the weights, relative position encoding is introduced into the graph convolution, and the shortest distance method is used for the skeleton graph G sThe relative distance between the upper two joints is encoded as shown in Equation (4), where P 1 and P P represent the first vertex and the last vertex on path P, and the weight parameter B i,j can be retrieved from the training parameter table. Then, according to the shortest path distance d i,j through the bone connection, it is assigned to each pair of joints.
[0047]
[0048] 2) Batch normalize the obtained graph encoding and use a 1×1 convolution as the residual connection of the downsampling layer. Process the result through the ReLU activation function and output it.
[0049] Follow the CTR-GCN architecture. As Figure 4 shown, stack the TB-GC graph convolution block and the multi-scale temporal convolution block to obtain the time-decoupled hierarchical graph convolution network, ensuring that the network has a high recognition speed while improving the recognition accuracy. The multi-scale temporal convolution block is as Figure 2 shown;
[0050] Step 5: Feed the fused input data into the improved time-decoupled hierarchical graph convolution network for training and testing, extract the deep spatio-temporal features with natural connections, non-natural connections, and interaction connection relationships, and obtain the final classification result through the pooling layer, fully connected layer, and activation function. The final recognition accuracy is 97.9%, proving the effectiveness of the method of the present invention. The confusion matrix obtained from the test is as Figure 6 shown.
[0051] Experiments prove that the method for recognizing dual-person interaction behaviors based on time-decoupled hierarchical graph convolution proposed by the present invention can effectively handle the problem of insufficient modeling of the time continuity of interaction behavior recognition, solve the problem of poor real-time performance of recognizing interaction behaviors with a large amount of data, and can effectively improve the accuracy of complex interaction scenarios, greatly improving the efficiency of dual-person interaction behavior recognition.
[0052] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these changes and deformations should also be regarded as the protection scope of the present invention.
Claims
1. A two-person interaction behavior recognition method based on time-decoupled hierarchical graph convolution, characterized in that: include: S1: The optimized joint point data is used as the input of the network to construct the hierarchical graph topology to obtain the hierarchical relationship matrix; S2: The CTR-GC channel topology optimization graph convolution is constructed as a time-decoupled hierarchical graph convolution. The hierarchical relationship matrix obtained in S1 is used to construct a THGC graph convolution block that can capture both temporal channel information and spatial channel information. S3: The graph convolution in THGC is constructed as BLOCKGC, which is then constructed as TBGC graph convolution block; S4: Stack the TBGC graph convolution block and the multi-scale temporal convolution block to obtain the TB-GCN graph convolution network; S5: Use the TB-GCN graph convolutional network to perform deep feature extraction and classification on the input human skeleton data to obtain the results of two-person interaction behavior recognition.
2. The method for identifying two-person interactive behavior based on time-decoupled hierarchical graph convolution according to claim 1 is characterized in that: The optimized joint point data described in S1 is the joint point data after removing noise and frame loss, and the optimization processing method is: obtaining the joint point data of two people through the depth camera Kinect v2, denoising and filtering the obtained joint point data for frame loss processing, and using three strategies of calculating the frame length, dispersion degree and motion value of the node data to denoise and filter the double-person joint point data for frame loss filtering. With the abdomen node as the center, a rooted tree is constructed to divide the human body joints into five layers, and then the graph is divided so that the nodes with concentrated edges of the same hierarchical structure are stored in the same semantic space.
3. The method for identifying two-person interactive behavior based on time-decoupled hierarchical graph convolution according to claim 1 is characterized in that: The THGC graph convolution block constructed in S2 consists of a residual structure consisting of two branches and 1×1 convolution. One of the branches is further extracted through a 1×1 convolution and a temporal pooling module to extract more detailed spatial feature information. The acquired spatial features are processed using self-contrastive subtraction and tanh activation function, and the number of channels is adjusted through 1×1 convolution to obtain spatial features. The other branch is composed of a bottleneck block consisting of two stream branches. One stream extracts more detailed temporal feature information through the spatial pooling module, and uses self-contrastive subtraction and tanh activation function to process the acquired temporal features. The other stream adjusts the number of channels through 1×1 convolution to retain its original feature information. The two stream features are fused to obtain the temporal features, and the obtained temporal features are multiplied and added element by element with the spatial features obtained in the previous branch and the weighted hierarchical relationship matrix to obtain the spatiotemporal features. The features that integrate time and space are fused with their own features after 1×1 convolution to form a time-decoupled hierarchical graph convolution block.
4. The method for identifying two-person interactive behavior based on time-decoupled hierarchical graph convolution according to claim 1 is characterized in that: In S3, the graph convolution in THGC is constructed as BLOCKGC, and then constructed as TBGC graph convolution block, including: processing input data by standardizing weights, introducing relative position coding into graph convolution, generating k-hop graph distance coding from the adjacency matrix and calculating weights, performing batch normalization and using 1×1 convolution as residual connection for the downsampling layer, and outputting the results through ReLU activation function processing.
Citation Information
Patent Citations
Human body behavior recognition method and system based on adaptive space-time convolutional network
CN114463837A
Action recognition method based on dynamic local-global graph convolutional neural network
CN114998525A
Human body behavior recognition method and system based on multi-channel directed graph convolution
CN116895097A
Skeleton behavior recognition method based on space-time re-aggregation graph convolutional network
CN118506452A
Skeleton behavior recognition method based on self-adaptive multi-dimensional dynamic graph convolutional network
CN118587479A