Human body behavior prediction method based on double-flow space-time diagram convolutional network
Through the dual-stream spatiotemporal graph convolution network, multi-space category modeling and frequency weighting mechanisms are used to solve the problems of long-range dependence and frequency information adjustment in human behavior prediction, and improve the prediction accuracy and robustness.
Patent Information
- Application Number
- CN202510616186.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-22
AI Technical Summary
Traditional graph convolutional networks (GCNs) are difficult to effectively model long-range relationships across the limbs in human behavior prediction, and it is difficult to dynamically adjust the weight of information of different frequencies, resulting in insufficient feature expression capabilities.
Using a dual-stream space-time graph convolution network, a multi-space category modeling graph convolution module and frequency weighting and channel feature importance attention mechanism is used to capture the short-range and long-range joint dependencies, and dynamically adjust the high-low-frequency feature weights.
It significantly improves the accuracy of human behavior prediction and enhances the robustness and feature extraction capabilities of the model in complex scenarios and diverse actions.
Smart Images

Figure CN120526477A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and deep learning technology, and specifically relates to a human behavior prediction method based on a dual-stream spatiotemporal graph convolutional network. Background Art
[0002] Human action prediction is a cutting-edge technology that aims to predict complete action categories from partial action sequences. Its core goal is to accurately predict the final action category using limited observation fragments, even before the action is fully unfolded. Unlike traditional action recognition, which relies on complete video sequences, action prediction utilizes only partial video fragments for analysis and prediction. This feature gives it significant advantages in computational efficiency and real-time response. With the widespread adoption of 5G technology and the rapid development of artificial intelligence, action prediction technology has shown broad application prospects in a variety of fields, such as intelligent surveillance, autonomous driving, medical rehabilitation, and home service robots.
[0003] In related technologies, graph convolutional networks (GCNs) have been widely used for skeletal behavior prediction due to their advantage in processing non-Euclidean data. However, traditional GCN methods have the following shortcomings: limited by a single neighborhood connectivity pattern, they can only capture short-range dependencies between physically adjacent nodes and have difficulty modeling long-range relationships across limbs, resulting in spatial domain limitations; they cannot effectively distinguish between high-frequency details and low-frequency trends in actions, which limits the ability to extract temporal features and leads to insufficient temporal domain performance; and channel attention mechanisms typically ignore frequency domain information, making it difficult to dynamically adjust the weights of high- and low-frequency features, resulting in insufficient distribution of feature channel importance. Although many GCN-based studies have attempted to address the above issues, these methods have not fully overcome the challenges of modeling long-range dependencies in the spatial domain and extracting multi-frequency features in the temporal domain. Furthermore, in the temporal domain, actions typically contain information of different frequencies (such as low-frequency overall trends and high-frequency local details), but traditional GCNs have difficulty dynamically adjusting the weights of this information, resulting in insufficient feature expression capabilities. Summary of the Invention
[0004] To address the challenges of traditional GCN methods in modeling long-range dependencies in the spatial domain and extracting multi-frequency features in the temporal domain, as well as the difficulty of dynamically adjusting the weights of information containing different frequencies, resulting in insufficient feature expression capabilities, this paper provides a human behavior prediction method based on a two-stream spatiotemporal graph convolutional network, aiming to improve the accuracy of human behavior prediction.
[0005] The technical solution adopted by the human behavior prediction method based on a dual-stream spatiotemporal graph convolutional network of the present invention is:
[0006] A human behavior prediction method based on a dual-stream spatiotemporal graph convolutional network includes the following steps:
[0007] S1. Obtain a human skeleton dataset, input a sequence of skeleton point data representing a single behavior, and construct a spatiotemporal skeleton graph based on the natural connections in the human skeleton space and the connections of the same nodes between adjacent frames to obtain input data;
[0008] S2. Preprocess the skeleton point data, construct a spatiotemporal graph and fill in the number of frames for the skeleton point data to generate a standardized spatiotemporal skeleton graph;
[0009] S3. Processing the preprocessed skeleton point data through a two-stream spatiotemporal graph convolutional network; each stream of the two-stream network contains an improved spatiotemporal graph convolutional layer to generate initial spatiotemporal features;
[0010] S4, through the multi-space category modeling graph convolution module to process the initial spatiotemporal features, construct multi-scale joint point spatial dependencies, and capture the short-range and long-range spatiotemporal features of the skeleton nodes;
[0011] S5. Process the output features of the first stream through the frequency weighting and channel feature importance attention mechanism, dynamically adjust the feature weights of high-frequency and low-frequency channels, and generate enhanced spatiotemporal features;
[0012] S6. Fuse the output features of the dual-stream network to generate the final behavior prediction results and realize human behavior prediction.
[0013] A further improvement of the technical solution of the present invention is that the spatiotemporal skeleton graph is formed in step S1, specifically,
[0014] Original 3D joint coordinate X S , directly represents the absolute spatial position of each joint point; the relative coordinate X R , obtained by calculating the vector difference between each joint and the central node of the human body; motion information X T , obtained by calculating the temporal difference of joint coordinates between adjacent frames; the bone vector X K , generating bone connection features with the joints close to the center of the skeleton as source nodes;
[0015] Then, X is transformed along the feature dimension S 、X R 、X T and X K Multimodal splicing is performed to form a fused feature tensor to achieve cross-modal information complementarity; finally, a spatiotemporal skeleton graph is formed based on the natural connection of the human skeleton and the connection between adjacent frames.
[0016] A further improvement of the technical solution of the present invention is that the skeleton point data is pre-processed in step S2, specifically,
[0017] Linear interpolation is used to fill in samples with insufficient frame numbers, and random interception is performed on samples with excessive frame numbers to ensure the consistency of the frame number of the input data. The specific calculation formula of the interpolation process is:
[0018]
[0019] Among them, P(t) represents the joint coordinates of the t-th frame, t1 and t2 are the time points of the adjacent frames before and after the missing frame, and are the joint coordinate values of the corresponding frames respectively.
[0020] A further improvement of the technical solution of the present invention is that the first stream of the dual-stream spatiotemporal graph convolutional network in step S3 includes the following processing flow:
[0021] First, the input data is a human skeleton sequence tensor with the dimension Where N is the batch size, C in is the number of input channels, T is the time step, and V is the number of joint points; then, the input data is standardized through the batch normalization layer to eliminate data distribution differences and accelerate model convergence; then, a 10-layer cascaded spatiotemporal graph convolution module is adopted, each layer contains: improved adjacency matrix modeling and adaptive edge weight learning mechanism, after layer-by-layer processing, high-order spatiotemporal feature maps are output; on this basis, the feature map is compressed through the spatiotemporal joint global average pooling layer to generate an aggregated feature vector, which is mapped to the unnormalized prediction score for category prediction through the fully connected layer; finally, the enhanced feature map is generated through the class activation mapping mechanism to locate the discriminative spatiotemporal regions and provide fine-grained feature support for the auxiliary branches.
[0022] A further improvement of the technical solution of the present invention is that the basic branch predicts the output category, and the specific calculation formula is:
[0023]
[0024] Where T' (or t) and V' (or v) represent the time dimension and spatial dimension after temporal convolution downsampling, N is the batch size, and K is the number of action categories.
[0025] A further improvement of the technical solution of the present invention is that the second stream of the dual-stream spatiotemporal graph convolutional network in step S3 includes the following processing flow:
[0026] First, the input data is the enhanced feature map generated by the base branch, which is adjusted to Among them C aux=256 is a fixed number of channels, and M is the multi-scale feature dimension. Next, a multi-layer spatiotemporal graph convolution module similar to the basic branch is adopted, but each layer additionally integrates the following mechanisms: frequency weighting and channel feature importance attention mechanism to extract local enhanced features. After layer-by-layer optimization, an enhanced feature map integrating multi-granularity spatiotemporal information is output. On this basis, the feature dimension is compressed through a spatiotemporal joint global average pooling layer to generate an aggregated feature vector, which is mapped to the unnormalized prediction score of the auxiliary branch through a fully connected layer.
[0027] A further improvement of the technical solution of the present invention is that the auxiliary branch predicts the output category, and the specific calculation formula is:
[0028]
[0029] in, C aux =256 is fixed as the number of channels of the auxiliary branch, N is the batch size, T is the number of time frames, V is the number of nodes, M is the number of different modalities, T' (or t) and V' (or v) represent the time dimension and spatial dimension after temporal convolution downsampling, respectively; finally, the prediction scores of the basic branch and the auxiliary branch are fused through the learnable weight coefficient to generate a more robust behavior prediction result.
[0030] A further improvement of the technical solution of the present invention is that in step S4, the specific calculation process of the multi-space category modeling graph convolution module is:
[0031] Input features The initial feature representation is generated by interacting with the weight matrix and bias vector through a two-dimensional convolution operation, where N is the batch size, T is the number of time frames, and V is the number of nodes. The specific calculation formula is:
[0032]
[0033] Where W is the weight matrix and b is the bias vector; then, the expanded features are split into multiple subsets along the channel dimension. The specific calculation formula is:
[0034]
[0035] Where s is the number of feature subsets, C out is the number of output feature channels; through channel splitting, each spatial category obtains an independent feature subset x" :,k,:,:,: ](k=1,2,…,s), thereby achieving the decoupling of multi-scale spatial relationships; then, the model uses the adjacency matrix A∈R s×V×V The first s subsets of , perform graph convolution operation on the split feature x', the specific calculation formula is:
[0036]
[0037] Among them, s_kernel_size is the size of the spatial convolution kernel.
[0038] A further improvement of the technical solution of the present invention is that the frequency weighting and channel feature importance attention mechanism in step S5 is implemented according to the following processing steps:
[0039] The input feature map is subjected to average pooling and maximum pooling operations in the spatial dimension to extract low-frequency response features that characterize the law of motion and high-frequency detail features that contain detail changes. The low-frequency response features and high-frequency detail features are concatenated in the channel dimension and sequentially passed through the nonlinear transformation module composed of the fully connected layer and the Sigmoid activation function to generate the frequency attention weight matrix. The specific calculation formula is:
[0040] A freq =σ(FC(α·A avg +β·A max ))
[0041] Among them, α and β are learnable parameters, A freq is the frequency attention weight matrix, A avg and A max are low-frequency features and high-frequency features respectively; the frequency attention weight matrix, channel importance weight and the original input feature map are multiplied channel by channel to output an optimized feature map with frequency selection enhancement characteristics. The specific calculation formula is:
[0042] X final =X·A freq ·S channel
[0043] Among them, X freq is the frequency weighted feature, S channel is the channel importance weight.
[0044] A further improvement of the technical solution of the present invention is that the fusion of the dual-stream network output features in step S6 includes:
[0045] First, the unnormalized prediction scores for the first stream output Unnormalized prediction scores from the second stream output Perform weighted fusion of learnable weight coefficients; then, normalize the fusion results through the softmax function to generate the probability distribution of the final behavior prediction; the fusion weights are adaptively optimized through an end-to-end training process, and the contribution ratio of the dual-stream features is dynamically balanced based on the gradient backpropagation mechanism, thereby maximizing the model's generalization ability for complex action patterns.
[0046] Due to the adoption of the above technical solution, the technical advancements achieved by the present invention include:
[0047] This paper further proposes a dual-stream spatiotemporal graph convolutional network for multi-space category modeling through a multi-space category modeling graph convolution module and a frequency weighted and channel feature importance attention mechanism, which significantly improves the accuracy of behavior prediction.
[0048] The multi-spatial category modeling graph convolution module effectively captures short-range and long-range joint dependencies by decomposing node neighborhood relationships into multiple spatial categories (such as physical connections, dynamic semantic connections, and global statistical connections), overcoming the limitations of traditional GCNs in the spatial domain. The frequency weighting mechanism dynamically adjusts the weights of high- and low-frequency features, enhancing the model's ability to model frequency information and assign feature channel importance in the temporal domain. This multi-scale, multi-frequency feature extraction strategy enables the model to remain robust even in the presence of sparse spatiotemporal cues and ambiguous initial patterns, allowing it to adapt to complex scenarios and diverse actions. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a flow chart of a human behavior prediction method based on a dual-stream spatiotemporal graph convolutional network of the present invention;
[0050] Figure 2 This is a model framework diagram of a human behavior prediction method based on a dual-stream spatiotemporal graph convolutional network of the present invention;
[0051] Figure 3 This is a structural diagram of a multi-space category modeling graph convolution module of a human behavior prediction method based on a dual-stream spatiotemporal graph convolutional network in the present invention;
[0052] Figure 4 It is a human joint distribution map of the NTU-RGB+D dataset of a human behavior prediction method based on a dual-stream spatiotemporal graph convolutional network in the present invention. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. In the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.
[0054] like Figure 1 and Figure 2 As shown, the present invention discloses a human behavior prediction method based on a dual-stream spatiotemporal graph convolutional network, comprising the following steps:
[0055] S1. Obtain a human skeleton dataset, input a sequence of skeleton point data representing a single behavior, and construct a spatiotemporal skeleton graph according to the natural connection of the human skeleton space and the connection of the same nodes between adjacent frames to obtain the input data.
[0056] In practical applications, a depth camera (such as Kinect v2) can be used to capture the movements of participants and generate a human skeleton dataset containing multiple action categories and samples. Each skeleton frame is represented by the three-dimensional coordinates of multiple joints, and a frame may contain one or more skeletons. During the data acquisition process, data diversity can be enriched by recording the movements of different participants or capturing them from different perspectives. Different criteria can be used for evaluation, such as grouping participants (dividing participants into training and test groups) or dividing them by perspective (using sequences from certain perspectives as training data and sequences from other perspectives as test data). When constructing and using the data, it is necessary to identify and remove damaged samples to ensure data integrity and reliability.
[0057] S2. Preprocess the skeleton point data, construct a spatiotemporal graph and fill in the number of frames for the skeleton point data to generate a standardized spatiotemporal skeleton graph. Linear interpolation is used to fill in samples with insufficient frame numbers, and random interception is performed on samples with excessive frame numbers to ensure the consistency of the frame number of the input data. The specific calculation formula for the interpolation process is:
[0058]
[0059] Among them, P(t) represents the joint coordinates of the t-th frame, t1 and t2 are the time points of the adjacent frames before and after the missing frame, and are the joint coordinate values of the corresponding frames respectively.
[0060] S3. Process the preprocessed skeleton point data through a two-stream spatiotemporal graph convolutional network, where each stream of the two-stream network contains an improved spatiotemporal graph convolutional layer to generate initial spatiotemporal features.
[0061] The first stream of the two-stream spatiotemporal graph convolutional network includes the following processing steps:
[0062] First, the input data is a human skeleton sequence tensor with the dimension Where N is the batch size, C in The model is constructed with n being the number of input channels, T being the time step, and V being the number of joint points. Subsequently, the input data is normalized via a batch normalization layer to eliminate data distribution differences and accelerate model convergence. A 10-layer cascaded spatiotemporal graph convolution module is then employed, each layer comprising an improved adjacency matrix modeling and an adaptive edge weight learning mechanism. After layer-by-layer processing, a high-order spatiotemporal feature map is output. On this basis, the feature map is compressed via a spatiotemporal joint global average pooling layer to generate an aggregated feature vector, which is then mapped to an unnormalized prediction score for category prediction via a fully connected layer. Finally, an enhanced feature map is generated via a class activation mapping mechanism to locate discriminative spatiotemporal regions, providing fine-grained feature support for the auxiliary branches.
[0063] The basic branch predicts the output category, and the specific calculation formula is:
[0064]
[0065] Where T' (or t) and V' (or v) represent the time dimension and spatial dimension after temporal convolution downsampling, N is the batch size, and K is the number of action categories.
[0066] The second stream of the two-stream spatiotemporal graph convolutional network includes the following processing steps:
[0067] First, the input data is the enhanced feature map generated by the base branch, which is adjusted to Among them C aux =256 is the fixed number of channels, and M is the multi-scale feature dimension. Next, a multi-layer spatiotemporal graph convolution module similar to the base branch is used, but each layer additionally incorporates the following mechanisms: frequency weighting and channel feature importance attention mechanism to extract local enhanced features. After layer-by-layer optimization, an enhanced feature map integrating multi-granular spatiotemporal information is output. On this basis, the feature dimension is compressed through a spatiotemporal joint global average pooling layer to generate an aggregated feature vector, which is then mapped to the unnormalized prediction score of the auxiliary branch via a fully connected layer.
[0068] The auxiliary branch predicts the output category, and the specific calculation formula is:
[0069]
[0070] in, C aux = 256 is the fixed number of channels in the auxiliary branch, N is the batch size, T is the number of time frames, V is the number of nodes, M is the number of different modalities, and T' (or t) and V' (or v) represent the time and spatial dimensions after downsampling by temporal convolution, respectively. Finally, the prediction scores of the base branch and the auxiliary branch are fused using learnable weight coefficients to generate more robust behavior prediction results.
[0071] S4. Process the initial spatiotemporal features through the multi-space category modeling graph convolution module to construct multi-scale joint point spatial dependencies to capture the short-range and long-range spatiotemporal features of the skeleton nodes. The specific calculation process of the spatial category modeling graph convolution module is as follows:
[0072] Input features (where N is the batch size, T is the number of time frames, and V is the number of nodes) a preliminary feature representation is generated by interacting with the weight matrix and bias vector through a two-dimensional convolution operation. The specific calculation formula is:
[0073]
[0074] Where W is the weight matrix and b is the bias vector. Then, the expanded features are split into multiple subsets along the channel dimension. The specific calculation formula is:
[0075]
[0076] Among them, s is the number of feature subsets, C out is the number of output feature channels. Through channel splitting, each spatial category obtains an independent feature subset x" :,k,:,:,: ](k=1,2,…,s), thereby achieving the decoupling of multi-scale spatial relationships. Then, the model uses the adjacency matrix A∈R s×V×V The first s subsets of , perform graph convolution operation on the split feature x', the specific calculation formula is:
[0077]
[0078] Where s_kernel_size is the size of the spatial convolution kernel. This operation captures the joint dependencies between spatial categories through weighted aggregation of A, ultimately fusing them into a unified feature representation. To efficiently implement multi-category graph convolution, the multi-spatial category modeling graph convolution module uses the Einstein sum function and combines adjacency matrix subsets to perform parallel aggregation for each spatial category.
[0079] S5. Process the output features of the first stream through the frequency weighting and channel feature importance attention mechanism, dynamically adjust the feature weights of high-frequency and low-frequency channels, and generate enhanced spatiotemporal features.
[0080] The frequency weighting and channel feature importance attention mechanism is implemented in the following processing steps:
[0081] The input feature map is subjected to average pooling and maximum pooling operations in the spatial dimension to extract low-frequency response features that characterize the law of motion and high-frequency detail features that contain detail changes. The low-frequency response features and high-frequency detail features are concatenated in the channel dimension and sequentially passed through the nonlinear transformation module composed of the fully connected layer and the Sigmoid activation function to generate the frequency attention weight matrix. The specific calculation formula is:
[0082] A freq =σ(FC(α·A avg +β·A max ))
[0083] Among them, α and β are learnable parameters, A freq is the frequency attention weight matrix, A avg and A maxThe frequency attention weight matrix, channel importance weight and the original input feature map are multiplied channel by channel to output an optimized feature map with frequency selection enhancement characteristics. The specific calculation formula is:
[0084] X final =X·A freq ·S channel
[0085] Among them, X freq is the frequency weighted feature, S channel is the channel importance weight.
[0086] S6. Fuse the output features of the dual-stream network to generate the final behavior prediction result and realize behavior prediction.
[0087] The fusion of the two-stream network output features includes:
[0088] First, the unnormalized prediction scores for the first stream output Unnormalized prediction scores from the second stream output A weighted fusion of learnable weight coefficients is performed. The fusion results are then normalized using a softmax function to generate a probability distribution for the final behavior prediction. The fusion weights are adaptively optimized through an end-to-end training process, dynamically balancing the contribution of dual-stream features based on a gradient backpropagation mechanism to maximize the model's generalization capabilities for complex motion patterns.
[0089] In the above embodiment, the present invention provides a method for predicting human behavior based on a dual-stream spatiotemporal graph convolutional network. The present invention further proposes a dual-stream spatiotemporal graph convolutional network for multi-space category modeling through a multi-space category modeling graph convolution module and a frequency weighting and channel feature importance attention mechanism, which significantly improves the accuracy of behavior prediction; the MSC-GC module effectively captures short-range and long-range joint dependencies by decomposing node neighborhood relationships into multiple spatial categories (such as physical connections, dynamic semantic connections, and global statistical connections), overcoming the limitations of traditional GCN in the spatial domain. The frequency weighting mechanism dynamically adjusts the weights of high and low frequency features, enhancing the model's frequency information modeling capabilities and feature channel importance allocation capabilities in the time domain. This multi-scale, multi-frequency feature extraction strategy enables the model to remain robust under sparse spatiotemporal clues and fuzzy initial patterns, and adapt to complex scenes and diverse actions.
[0090] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the concept and scope of the present invention. Any modifications and improvements made to the technical solution of the present invention by a person of ordinary skill in the art without departing from the design concept of the present invention shall fall within the scope of protection of the present invention. The technical content for which protection is sought in the present invention is fully set forth in the claims.
Claims
1. A human behavior prediction method based on a two-stream spatiotemporal graph convolutional network, characterized by: The following steps are included: S1. Obtain a human skeleton dataset, input a sequence of skeleton point data representing a single behavior, and construct a spatiotemporal skeleton graph based on the natural connections in the human skeleton space and the connections of the same nodes between adjacent frames to obtain input data; S2. Preprocess the skeleton point data, construct a spatiotemporal graph and fill in the number of frames for the skeleton point data to generate a standardized spatiotemporal skeleton graph; S3. Processing the preprocessed skeleton point data through a two-stream spatiotemporal graph convolutional network; each stream of the two-stream network contains an improved spatiotemporal graph convolutional layer to generate initial spatiotemporal features; S4, through the multi-space category modeling graph convolution module to process the initial spatiotemporal features, construct multi-scale joint point spatial dependencies, and capture the short-range and long-range spatiotemporal features of the skeleton nodes; S5. Process the output features of the first stream through the frequency weighting and channel feature importance attention mechanism, dynamically adjust the feature weights of high-frequency and low-frequency channels, and generate enhanced spatiotemporal features; S6. Fuse the output features of the dual-stream network to generate the final behavior prediction results and realize human behavior prediction.
2. The method for predicting human behavior based on a dual-stream spatiotemporal graph convolutional network according to claim 1, characterized in that: The spatiotemporal skeleton graph is formed in step S1, specifically, Original 3D joint coordinate X S , directly represents the absolute spatial position of each joint point; the relative coordinate X R , obtained by calculating the vector difference between each joint and the central node of the human body; motion information X T , obtained by calculating the temporal difference of joint coordinates between adjacent frames; the bone vector X K , generating bone connection features with the joints close to the center of the skeleton as source nodes; Then, X is transformed along the feature dimension S 、X R 、X T and X K Multimodal splicing is performed to form a fused feature tensor to achieve cross-modal information complementarity; finally, a spatiotemporal skeleton graph is formed based on the natural connection of the human skeleton and the connection between adjacent frames.
3. The method for predicting human behavior based on a dual-stream spatiotemporal graph convolutional network according to claim 1, characterized in that: The skeleton point data is pre-processed in step S2, specifically, Linear interpolation is used to fill in samples with insufficient frame numbers, and random interception is performed on samples with excessive frame numbers to ensure the consistency of the frame number of the input data. The specific calculation formula of the interpolation process is: Among them, P(t) represents the joint coordinates of the t-th frame, t1 and t2 are the time points of the adjacent frames before and after the missing frame, and are the joint coordinate values of the corresponding frames respectively.
4. The method for predicting human behavior based on a dual-stream spatiotemporal graph convolutional network according to claim 1, characterized in that: The first stream of the dual-stream spatiotemporal graph convolutional network in step S3 includes the following processing flow: First, the input data is a human skeleton sequence tensor with the dimension Where N is the batch size, C in is the number of input channels, T is the time step, and V is the number of joint points; then, the input data is standardized through the batch normalization layer to eliminate data distribution differences and accelerate model convergence; then, a 10-layer cascaded spatiotemporal graph convolution module is adopted, each layer contains: improved adjacency matrix modeling and adaptive edge weight learning mechanism, after layer-by-layer processing, high-order spatiotemporal feature maps are output; on this basis, the feature map is compressed through the spatiotemporal joint global average pooling layer to generate an aggregated feature vector, which is mapped to the unnormalized prediction score for category prediction through the fully connected layer; finally, the enhanced feature map is generated through the class activation mapping mechanism to locate the discriminative spatiotemporal regions and provide fine-grained feature support for the auxiliary branches.
5. The method for predicting human behavior based on a dual-stream spatiotemporal graph convolutional network according to claim 4, characterized in that: The basic branch predicts the output category, and the specific calculation formula is: Where T' (or t) and V' (or v) represent the time dimension and spatial dimension after temporal convolution downsampling, N is the batch size, and K is the number of action categories.
6. The method for predicting human behavior based on a dual-stream spatiotemporal graph convolutional network according to claim 1, characterized in that: The second stream of the dual-stream spatiotemporal graph convolutional network in step S3 includes the following processing flow: First, the input data is the enhanced feature map generated by the base branch, which is adjusted to Among them C aux =256 is a fixed number of channels, and M is the multi-scale feature dimension. Next, a multi-layer spatiotemporal graph convolution module similar to the basic branch is used, but each layer additionally integrates the following mechanisms: frequency weighting and channel feature importance attention mechanism to extract local enhanced features. After layer-by-layer optimization, an enhanced feature map that integrates multi-granularity spatiotemporal information is output; on this basis, the feature dimension is compressed through the spatiotemporal joint global average pooling layer to generate an aggregated feature vector, which is mapped to the unnormalized prediction score of the auxiliary branch through the fully connected layer.
7. The method for predicting human behavior based on a dual-stream spatiotemporal graph convolutional network according to claim 6, characterized in that: The auxiliary branch predicts the output category, and the specific calculation formula is: in, C aux =256 is fixed as the number of channels of the auxiliary branch, N is the batch size, T is the number of time frames, V is the number of nodes, M is the number of different modalities, T' (or t) and V' (or v) represent the time dimension and spatial dimension after temporal convolution downsampling, respectively; finally, the prediction scores of the basic branch and the auxiliary branch are fused through the learnable weight coefficient to generate a more robust behavior prediction result.
8. The method for predicting human behavior based on a dual-stream spatiotemporal graph convolutional network according to claim 1, characterized in that: In step S4, the specific calculation process of the multi-space category modeling graph convolution module is: Input features The initial feature representation is generated by interacting with the weight matrix and bias vector through a two-dimensional convolution operation, where N is the batch size, T is the number of time frames, and V is the number of nodes. The specific calculation formula is: Where W is the weight matrix and b is the bias vector; then, the expanded features are split into multiple subsets along the channel dimension. The specific calculation formula is: Where s is the number of feature subsets, C out is the number of output feature channels; through channel splitting, each spatial category obtains an independent feature subset x" [:,k,:,:,:] (k=1,2,…,s), thereby achieving the decoupling of multi-scale spatial relationships; then, the model uses the adjacency matrix A∈R s×V×V The first s subsets of , perform graph convolution operation on the split feature x', the specific calculation formula is: Among them, s_kernel_size is the size of the spatial convolution kernel.
9. The method for predicting human behavior based on a dual-stream spatiotemporal graph convolutional network according to claim 1, characterized in that: The frequency weighting and channel feature importance attention mechanism in step S5 is implemented according to the following processing steps: The input feature map is subjected to average pooling and maximum pooling operations in the spatial dimension to extract low-frequency response features that characterize the law of motion and high-frequency detail features that contain detail changes. The low-frequency response features and high-frequency detail features are concatenated in the channel dimension and sequentially passed through the nonlinear transformation module composed of the fully connected layer and the Sigmoid activation function to generate the frequency attention weight matrix. The specific calculation formula is: And freq =σ(FC(α·A avg +β·A max )) Among them, α and β are learnable parameters, A freq is the frequency attention weight matrix, A avg and A max are low-frequency features and high-frequency features respectively; the frequency attention weight matrix, channel importance weight and the original input feature map are multiplied channel by channel to output an optimized feature map with frequency selection enhancement characteristics. The specific calculation formula is: X final =X·A freq ·S channel Among them, X freq is the frequency weighted feature, S channel is the channel importance weight.
10. The method for predicting human behavior based on a dual-stream spatiotemporal graph convolutional network according to claim 1, characterized in that: The fusion of the dual-stream network output features in step S6 includes: First, the unnormalized prediction scores for the first stream output Unnormalized prediction scores from the second stream output Perform weighted fusion of learnable weight coefficients; then, normalize the fusion results through the softmax function to generate the probability distribution of the final behavior prediction; the fusion weights are adaptively optimized through an end-to-end training process, and the contribution ratio of the dual-stream features is dynamically balanced based on the gradient backpropagation mechanism, thereby maximizing the model's generalization ability for complex action patterns.
Citation Information
Cited By
Human activity intensity prediction method based on generalized spatial heterogeneity learning
CN121436040A