A Video-based Multi-person Behavior Analysis Method
By adopting a weighted fusion scene information and attention mechanism in behavior recognition and combining with graph convolutional networks for relationship reasoning, the problem of underutilization of scene information in the prior art is solved, and the accuracy and efficiency of behavior recognition are significantly improved.
Patent Information
- Application Number
- CN202110814083.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-07-19
AI Technical Summary
Existing behavior recognition methods rarely consider scene information, resulting in insufficient accuracy of individual and group behavior recognition.
Weighted fusion method is used to fuse scene information, and an attention mechanism module is introduced, and relationship reasoning is performed in combination with graph convolution networks to improve the accuracy of behavior recognition.
By integrating scene information and attention mechanism, the accuracy of individual and group behavior recognition is significantly improved, saving training time and computing resources.
Smart Images

Figure CN115641525B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the problem of individual behavior and group behavior recognition in the field of deep learning, and particularly to a method for multi-person behavior analysis based on video. Background Art
[0002] With the booming development of the computer vision field, more and more scholars have carried out relevant research. And human behavior recognition, as a hot topic among them, has also received great attention. Most of the existing research conducts behavior recognition by extracting relevant features such as the appearance information of RGB images, the motion information of optical flow images, or the human skeleton information, and all have achieved good results. In recent years, with the development of graph neural networks, the behavior recognition field has also begun to introduce relevant ideas to simulate the human brain's reasoning about the interaction between humans, so as to improve the accuracy of behavior recognition. Currently, multi-person behavior analysis based on video has profound research significance and broad application prospects in fields such as intelligent video surveillance, video retrieval, and intelligent driving.
[0003] Existing behavior recognition methods rarely consider the role of scene information. In fact, both individual and group behaviors are related to the scenes they are in. Therefore, based on the analysis of human appearance information and position information, this patent uses a weighted fusion method to fuse scene information and adds an attention mechanism module to improve the accuracy of behavior recognition. First, two channels are used to extract individual features and scene features respectively: for the individual channel, the Inception-v3 network is used for global feature extraction, and then the RoIAlign module and the attention mechanism are used to extract the features of each individual; for the scene channel, the ResNet-50 pre-trained on the place365 dataset is used for scene feature extraction. Secondly, the extracted individual appearance features and position features are input into a graph convolutional network for further relationship reasoning to learn the interaction between individuals. Among them, the cosine similarity method is used to judge the similarity of appearance between individuals. Finally, the scene features are resized and weighted-fused with the features inferred by the graph convolutional network, and after being input into the classifier, the group behavior and the behavior recognition results of all individuals are obtained. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for multi-person behavior analysis based on video, which takes into account the features related to behavior recognition, adopts a weighted fusion method to fuse scene features, introduces an attention mechanism module to assign different weights to different positions or channels, and uses a graph convolutional network for relationship reasoning, solving the problem of insufficient feature extraction in behavior recognition.
[0005] For the convenience of description, the following concepts are introduced first:
[0006] Graph: A graph structure is composed of nodes and edges that connect the nodes.
[0007] Graph Convolutional Network (GCN): The purpose of GCN is to extract the spatial features of a topological graph. The data it processes is in graph structure, i.e., non-Euclidean structure.
[0008] Inception Network: It is a multi-branch structure that uses convolutional kernels of different sizes to capture features at different depths, and finally concatenates them to obtain multi-scale features. Common versions include Inception-v1, Inception-v2, Inception-v3, Inception-v4, etc.
[0009] RoIAlign Module: It uses the method of bilinear interpolation to calculate the accurate values of features in each Region of Interest (RoI), and uses pooling operation to obtain the output result, solving the problem of quantization error in RoI-Pooling.
[0010] Residual Networks (ResNets): By introducing the identity mapping to solve the problem of network degradation. Common network types include ResNet-18, ResNet-34, ResNet-50, ResNet-101, and ResNet-152.
[0011] Squeeze-and-excitation Network: It is a channel attention mechanism that can automatically capture the importance of each channel, so as to amplify useful features and suppress useless features.
[0012] Transfer Learning: It refers to using a pre-trained model in another task, and making some modifications to assist in training a new model, which greatly improves the efficiency of model training.
[0013] The present invention specifically adopts the following technical solutions:
[0014] A method for multi-person behavior analysis based on video, characterized in that:
[0015] a. Extract features related to individual behavior and group behavior recognition through a convolutional neural network, a fully connected network, and a graph convolutional network;
[0016] b. Use the attention mechanism Squeeze-and-excitation Networks to give different degrees of attention to different channels, so as to recalibrate the behavior features;
[0017] c. On the basis of constructing a graph structure using appearance information and location information, a cosine similarity judgment method is adopted to calculate the appearance similarity between individuals;
[0018] d. The scene information is fused in a weighted fusion manner;
[0019] This method mainly includes the following steps:
[0020] (1) Data preprocessing: Extract frames from continuous video frames and directly input the sampled video frames into the network;
[0021] (2) Feature extraction: The Inception-v3 network and ResNet-50 network are used to extract the global features of individuals and scene features. Then, through the RoIAlign module and the attention mechanism module, the feature information of each individual is obtained, and the fully connected layer is used to synthesize the obtained information;
[0022] (3) Node generation: Combine the appearance features extracted in step (2) with the location features of each individual to form behavior feature nodes;
[0023] (4) Graph construction and reasoning: Construct an individual behavior relationship graph based on the behavior feature nodes generated in step (3), and perform reasoning on the graph through a graph convolutional network to fully explore the interaction of behaviors between individuals;
[0024] (5) Fuse scene information: Perform weighted fusion on the feature map inferred in step (4) and the scene feature map;
[0025] (6) Classification of individual behaviors and group behaviors: Input the features after fusing the scene into the classification layer to obtain the final categories of individual and group behaviors;
[0026] (7) Model training: The model training constructed through (2)-(6) is divided into two steps. In the first step, the network for generating behavior appearance feature nodes is trained as a whole. After saving the model parameters, it is input into the network model of the second step. In the second step, the location information of the individual is added, and then combined with the appearance feature nodes of the first step to construct a graph convolutional network and perform reasoning on it. Then, through the method of weighted fusion of scene features, the final classification result is obtained.
[0027] The beneficial effects of the present invention are:
[0028] (1) The scene features are fused in a weighted fusion manner, improving the accuracy of group and individual behavior recognition.
[0029] (2) Using the method of transfer learning for scene feature extraction saves a large amount of training time and computing resources.
[0030] (3) The graph convolutional network is adopted to infer the interaction between individuals, and the method of cosine similarity is used to calculate the appearance similarity between individuals, so as to better judge the individual behavior and group behavior.
[0031] (4) The channel attention mechanism is used to recalibrate all features, amplify important features and suppress less useful features. Description of the Drawings
[0032] Figure 1 It is a weighted fusion framework diagram. Specific Embodiments
[0033] The present invention will be further described in detail below in conjunction with the drawings and embodiments. It is necessary to point out that the following embodiments are only used to further illustrate the present invention and cannot be construed as limiting the protection scope of the present invention. Those skilled in the art can make some non-essential improvements and adjustments to the present invention according to the above-mentioned inventive content and implement it specifically, which should still fall within the protection scope of the present invention.
[0034] The multi-person behavior analysis method based on video specifically includes the following steps:
[0035] (1) Data preprocessing
[0036] The input continuous video frames are sampled, and the sampled video frames are directly input into the network for training.
[0037] (2) Feature extraction
[0038] The weighted fusion method is as Figure 1 shown. The individual channels adopt the Inception-v3 network to extract individual global features, and then obtain the behavioral appearance features of each individual through the RoIAlign module and the attention mechanism module. Finally, all information of the individual is integrated through the fully connected layer; the scene channels adopt the ResNet-50 network pre-trained on the place365 dataset to extract scene features.
[0039] (3) Node generation
[0040] According to the bounding box information of each individual annotated in the dataset, the coordinates of the center point of the bounding box are calculated to obtain the position features corresponding to each individual, and then combined with the appearance features extracted in step (2) to form behavioral feature nodes. Among them, the dimension of the appearance features is 1024 dimensions.
[0041] (4) Graph construction and inference
[0042] Connect the behavior feature nodes generated in step (3) in a fully connected manner to construct an individual behavior relationship graph, and use a graph convolutional network to reason about the graph, so as to fully explore the interactivity and connection of behaviors between individuals, and thus better classify the behavior categories. The specific method is as follows:
[0043] Assume that the individuals constituting the GCN nodes are defined as N represents the number of people in the scene, is the appearance feature of the i-th individual, is the center point coordinate of the bounding box of the i-th individual. Construct a graph to represent the relationship between each individual. G ij represents the importance of the j-th individual feature to the i-th individual, as shown in equation (1).
[0044]
[0045] Among them, the cosine similarity method is used to calculate the similarity of appearances between individuals, as shown in equation (2). and are learnable linear transformations. and are weight matrices, and are weight vectors.
[0046]
[0047] The calculation of the positional relationship is shown in equation (3). When the distance between two individuals exceeds the set threshold μ, G ij is 0. represents the Euclidean distance between two individuals, and g(·) is the indicator function.
[0048]
[0049] The reasoning of GCN is shown in equation (4). Among them, is the matrix representation of the graph, is the feature representation of the nodes in the l-th layer, Z 0 = X. is a learnable weight matrix, σ(·) represents the ReLU activation function, and only one layer of GCN is used for reasoning.
[0050] Z (l+1) = σ(GZ (l) W (l) ) (4)
[0051] (5) Integrate scene information
[0052] The feature map inferred through step (4) is subjected to max pooling operation to obtain the behavior features at the group level, and then weighted fusion is performed after normalization with the scene feature map through the softmax function to obtain the individual and group behavior features integrating scene information.
[0053] (6) Classification of individual behaviors and group behaviors
[0054] The behavior features of the integrated scene are input into the classification layer to finally achieve the classification of individual behavior and group behavior categories.
[0055] (7) Model training
[0056] The training of the model constructed through (2)-(6) is divided into two steps. In the first step, the network for generating behavior appearance feature nodes is trained as a whole, and after saving the model parameters, it is input into the network model of the second step; in the second step, the position information of the individual is added, and then combined with the appearance feature nodes of the first step, a graph convolutional network is constructed and inferred, and the scene features are fused through weighted fusion, and finally the classification results of group behaviors and all individual behaviors are obtained.
Claims
1. A video-based behavior analysis method, characterized in that: a. Extract features related to individual behavior and group behavior recognition through convolutional neural networks, fully connected networks, and graph convolutional networks; b. Use the attention mechanism Squeeze-and-excitation Networks to give different degrees of attention to different channels to recalibrate behavior features; c. On the basis of constructing a graph structure using appearance information and position information, adopt a cosine similarity judgment method to calculate the appearance similarity between individuals; d. Adopt a weighted fusion method to fuse scene information; This method mainly includes the following steps: (1) Data preprocessing: Extract frames from continuous video frames and directly input the sampled video frames into the network; (2) Feature extraction: Use Inception-v3 network and ResNet-50 network to extract individual global features and scene features, then obtain the feature information of each individual through the RoIAlign module and the attention mechanism module, and use the fully connected layer to synthesize the obtained information; (3) Node generation: Combine the appearance features extracted in step (2) with the position features of each individual to form behavior feature nodes; (4) Graph construction and reasoning: Construct an individual behavior relationship graph based on the behavior feature nodes generated in step (3), and perform reasoning on the graph through a graph convolutional network to fully explore the interactivity of behaviors between individuals; (5) Fuse scene information: Perform weighted fusion on the feature map inferred in step (4) and the scene feature map; (6) Classification of individual behavior and group behavior: Input the features after fusing the scene into the classification layer to obtain the final categories of individual and group behaviors; (7) Model training: The model training constructed through (2)-(6) is divided into two steps. In the first step, the network for generating behavior appearance feature nodes is trained as a whole, and after saving the model parameters, it is input into the network model of the second step; In the second step, the position information of the individual is added, and then combined with the appearance feature nodes of the first step, a graph convolutional network is constructed and reasoned, and finally, through the method of weighted fusion of scene features, the final classification result is obtained.
2. The video-based multi-person behavior analysis method according to claim 1, wherein In steps (2) and (4), feature extraction is divided into two channels: the individual channel and the scene channel; for the individual channel, use the Inception-v3 network to extract global features, and then input the features passing through the RoIAlign module into the fully connected layer to obtain the initial individual appearance features. In addition, combine the position relationship of the individual bounding box to form a graph to establish the mutual connection between individuals.
3. The method for multi-person behavior analysis based on video according to claim 1, wherein In steps (2) and (5): In the scene channel, adopt the method of transfer learning, use the ResNet-50 model pre-trained on the place365 dataset to extract scene features, and perform weighted fusion on the extracted scene features and the features inferred through the graph convolutional network after normalization through the softmax function to explore the behavior features containing scene information.
4. The method for multi-person behavior analysis based on video according to claim 1, wherein In step (2), the channel attention mechanism Squeeze-and-excitation Networks is adopted to give different degrees of attention to different channels, so as to better allocate greater weights to the channels that contribute more to action recognition and suppress the channels with little effect.
5. The method for multi-person behavior analysis based on video according to claim 1, wherein In step (4), in the inference of the graph convolutional network, the method of cosine similarity is adopted to calculate the similarity of appearance between individuals. This method can better characterize the similarity between every two individuals, thus better mining the interactivity.
6. The video-based multi-person behavior analysis method according to claim 1, wherein In steps (5) and (6), the individual features extracted by the convolutional neural network, fully connected network and graph convolutional network are obtained as high-level group features through the operation of max pooling. The individual and group features are weighted and fused with the scene features respectively, and finally the categories of individual and group behaviors are obtained through the classification layer; the pooling method can well synthesize the features of all individuals to obtain group-level features, making the final classification more reasonable.