Group Behavior Recognition Method Based on Multimodal Fusion and Implicit Interaction Relationship Learning
Through the methods of adaptive multimodal fusion and implicit interactive relationship learning, the problems of multimodal feature redundancy and complex calculations in group behavior recognition are solved, and efficient and accurate group behavior recognition is achieved.
Patent Information
- Application Number
- CN202211365228.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-03
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-03
AI Technical Summary
The prior art has redundant multimodal individual member feature information in group behavior recognition, which makes it difficult to highlight significant information through operations such as cascade addition, and explicit modeling methods are complex and huge in calculations, making it difficult to accurately identify the behavior and group behavior of each individual in the group.
Using the method based on adaptive multimodal fusion and implicit interactive relationship learning, group behavior recognition is achieved through dynamic and static dual-stream character feature extraction, multimodal feature fusion, member interaction relationship learning and global feature extraction.
It improves the accuracy of group behavior recognition, reduces the amount of calculation, enhances the modeling ability of group members' interaction relationships, reduces recognition errors, and improves recognition accuracy.
Smart Images

Figure CN115719510B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of group behavior recognition, and particularly relates to a group behavior recognition method based on adaptive multi-modal fusion and implicit interaction relationship learning. Background Art
[0002] In recent years, human behavior recognition in videos has achieved remarkable results in the field of computer vision, and human body behavior recognition has also been widely applied in real life, such as intelligent video surveillance, abnormal event detection, sports analysis, understanding social behavior, etc. These applications make group behavior recognition have important scientific practicality and great economic value.
[0003] The solution proposed in “A Multi-Stream Convolutional Neural Network Framework for Group Activity Recognition” (Document 1) published by Sina Mokhtarzadeh Azar et al. is a multi-stream convolutional neural network group behavior recognition framework. Different CNN streams are trained separately, and finally decision-level fusion is performed to predict the final group behavior. In addition, “Empowering Relational Network by Self-Attention Augmented Conditional Random Fields for Group Activity Recognition” (Document 2), and “Deep bilinear learning for rgb-d action recognition” (Document 3) published by Hu et al. also disclose corresponding group behavior recognition methods. However, for the algorithm framework disclosed in Document 1, although it combines multi-modal feature representations and enriches the extracted information, only late decision-level fusion is adopted, and feature multi-modal fusion is not considered, which will inevitably lead to the problem of redundant multi-modal feature information; although Document 2 aggregates pose, spatial position, and appearance features in a cascaded manner during the member feature extraction stage, such simple operations usually have few or no associated parameters and cannot learn the interaction of complementary information and the saliency representation of each modality; while Document 3 constructs a tensor structure cube of skeleton and RGB features for multi-modal fusion to combine human action features, enabling the features of each modality to complement each other in learning, but it cannot avoid its huge computational complexity.
[0004] In recent years, in the part of member interaction relationship reasoning, "Learning actor relation graphs for group activity recognition." published by Wu et al., and "Convolutional relational machine for group activity recognition." published by Azar et al. use graph convolutional networks and intermediate representations such as graph structures to establish a topological graph of human relationships, capture the appearance and position relationships among members, and perform relationship reasoning. Although the above methods have achieved good prediction results, in order to extract spatial features, these explicitly modeled methods require explicit human position nodes to establish a topological graph structure, and use convolutional neural networks as basic building blocks to calculate all input hidden representations and output positions in parallel, associate signals from two arbitrary input or output positions, and the number of operations required increases with the distance between positions. The characteristic of multiple repeated iterations makes the network calculation complex and huge. Summary of the Invention
[0005] Aiming at the problem of redundant multi-modal individual member feature information in the prior art group behavior recognition and the difficulty in highlighting significant information caused only by operations such as cascading addition, the present invention proposes a group behavior recognition method based on adaptive multi-modal fusion and implicit interaction learning to accurately identify the behavior of each individual in the group and infer the group behavior using the individual and their interaction features.
[0006] The present invention is implemented by the following technical solutions: A group behavior recognition method based on multi-modal fusion and implicit interaction relationship learning, comprising the following steps:
[0007] Step A, dynamic and static dual-stream human feature extraction: Extract human static pose features and dynamic optical flow features based on the human-level feature extraction module;
[0008] Step B, multi-modal feature fusion: After unimodal connection of the static pose features and the dynamic optical flow features, perform convolution compression to obtain a latent vector of significant information, and then obtain a feature information representation that contains both optical flow and the most representative features of each modality of fine-grained pose after fusion;
[0009] Step C, member interaction relationship learning: Using the fused feature information representation obtained in Step B, based on the self-attention mechanism, calculate the appearance similarity of paired human features through the association strength, and selectively extract the humans important for behavior recognition to obtain an implicit vector representation between group members calculated in the form of the sum of attention weights;
[0010] Step D, global feature extraction: Based on the global feature extraction module, extract global feature information including background information for the input video frame;
[0011] Step E: Based on Step C and Step D, achieve the recognition of group behavior.
[0012] Furthermore, the specific steps of Step B are as follows:
[0013] Step B1: First, connect the single unimodal features and reduce the number of channels through the encoder convolutional network to obtain a self-fused latent vector.
[0014] Step B2: Reconstruct the initially connected vector from the self-fused latent vector.
[0015] Step B3: Minimize the Euclidean distance between the original and reconstructed concatenated vectors, and use the intermediate vector as the fused multi-modal feature information representation.
[0016] Furthermore, Step B1 is specifically implemented in the following way:
[0017] (1) Embed the static pose feature and dynamic optical flow feature of the person obtained by the person-level feature extraction module into a vector space with the same dimension through Embedding linear mapping, and use them as unimodal inputs respectively.
[0018] (2) Given n d-dimensional multi-modal latent vectors, n ≤ 3, let The two modalities represent the static pose feature and dynamic optical flow feature vectors of the person respectively. First, perform a concatenation operation to obtain where
[0019] (3) Then obtain through the encoding part and reduce its dimension to t:
[0020] In the encoding part, first compress the dimension of the multi-modal concatenation through the Linear layer to become the dimension of the unimodal initialization, and then perform a non-linear mapping through the Tanh activation function; then continue to perform a second Linear linear transformation to compress the feature dimension, and then perform activation through the Relu function. At this time, is called the fused latent feature.
[0021] Furthermore, Step B2 is specifically implemented in the following way:
[0022] Reconstruct the initially connected vector from the fused latent feature through the decoding transformation part to obtain Calculate the loss F between and and to guide the iterative optimization of the network, so that the learned latent feature representation can best represent the significant information of each modality. tr
[0023] Further, in step B3, the MSE loss function is used to guide the learning of the fusion network, and the intermediate vector is used as the fused multi-modal feature information representation.
[0024] Further, step C specifically includes the following steps:
[0025] The first stage: Calculate the scores of the association degrees of each person with other participants by matching the query Q with the key-value set K. All three representations (Q, K, V) are calculated from the input sequence S through linear projection. S is a set of person features S = {s i | i = 1, …, N} obtained after multi-modal adaptive fusion by the person feature extractor, and A(S) = A(Q(S), K(S), V(S));
[0026] The second stage: Normalize the results of the association degrees of each person with other participants obtained by calculating the dot product of the query Q and K to obtain a similarity set a n , n = 1, 2, …, n persons, and the sum of them is 1;
[0027] The third stage: Multiply the similarity vectors obtained by normalization in the second stage by V respectively to obtain the final weighted sum attention matrix for final classification and recognition.
[0028] Further, in step D, I3D is used as the backbone network, and an RGB video clip is used as the input. Select T frames centered on the annotated frame, and use the deep spatio-temporal feature map extracted from the final convolutional layer as the rich semantic representation describing the entire video clip.
[0029] Further, in step E, the output of step D is fed to a fully connected layer. The output of this fully connected layer is connected to the output of the member interaction learning module in step C and is passed as input to a fully connected layer with a classification layer. The characteristics of the difference in scene-level information in the global view are used to assist the person-level features to jointly perform group behavior recognition;
[0030] When performing recognition and classification, two classifiers are set respectively to generate group behavior category scores and individual action category scores. The final recognition and classification results of group behavior and individual actions are obtained through network learning. The overall network training process selects the cross-entropy loss to guide the optimization.
[0031] Furthermore, for the input video frames, they are divided into three branches, including a static branch, a dynamic branch, and a global branch; the static branch and the dynamic branch are connected to the person-level feature extraction module, and the global branch is connected to the global feature extraction module. In step A, a two-stream feature extraction method is used to enrich the features of individuals, which specifically includes:
[0032] (1) The static branch backbone network extracts the static pose features of the person: using the bounding box around the person in the video frame as the input of the pose estimation model HRNet to predict the key nodes;
[0033] (2) The dynamic branch backbone network extracts the optical flow dynamic features: adopting the optical flow representation in the I3D network, first converting the input sequence of frames into a continuous sequence of optical flow frames, then processing the stacked frame sequence through an inflated 3D convolutional network, and adding an ROIAlign layer to project the coordinates onto the feature map of the frame, so as to extract the optical flow features of each person's bounding box in the input frame.
[0034] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0035] First, this solution extracts the individual behaviors in the group, models and reasons the interaction relationships among the group members, and finally achieves the purpose of predicting the group behavior. By designing a dynamic and static multi-modal feature extraction module (task-level feature extraction module), compared with the prior art, it not only extracts the optical flow in 3DCNN to represent the dynamic features of the person, but also considers that in group behaviors, the actions of most people are highly correlated with the positions and movements of body joints. The member pose features are extracted through HRNet to refine the subtle representation of the person's actions. Through the fusion of the 2D pose features and the 3D optical flow dynamic features, the person feature information is enriched;
[0036] Second, in the multi-modal feature fusion stage, a lightweight component design is adopted. The adaptive multi-modal fusion module is used to learn the collaborative shared representation of each modality, effectively combine the multi-modal data, and at the same time, the lightweight design also solves the problem of huge computational complexity in previous work. Compared with the simple fusion methods of addition and concatenation, the average recognition accuracy in the volleyball dataset has been improved by 4.6% and 3.7% respectively in the experiments;
[0037] Third, in the member interaction relationship reasoning module, the Transformer network relies on the self-attention mechanism in NLP tasks, which can well model the dependencies between words and has the advantages of not requiring repetition or recursion. From the perspective of model scalability, a standard Transformer encoder is directly applied in the group behavior recognition interaction reasoning module, acting on the participant interactions without any specific changes for visual tasks, to exert its effects in the visual field; experimental results show that compared with models that require explicit spatial and temporal constraints, this solution better models the relationships between people and combines person-level information for group behavior recognition;
[0038] Fourth, based on the global feature extraction module, the problem of confusion between similar group behaviors is solved, and the recognition error is reduced. For example, the two behaviors of "walking" and "crossing" in the Collective Activity Datasets CAD usually occur in open-air places, and the actions of people are essentially walking. It is difficult to distinguish the categories of group behaviors through individual motion features, and confusion often occurs during recognition. However, "crossing" mostly occurs at intersections with important scene information such as zebra crossings and traffic lights, while the global scenes of "walking" are mostly park paths, etc. Therefore, this solution designs a global feature extraction module to assist recognition through the scene context features of different behaviors. Experiments prove that adding this module improves the average recognition accuracy by 2.3% in the CAD ablation experiment. Description of the Drawings
[0039] Figure 1 Schematic flow chart of the group behavior recognition method described in the embodiment of the present invention;
[0040] Figure 2 Schematic diagram of the principle of the adaptive multi-modal feature fusion module described in the embodiment of the present invention;
[0041] Figure 3 Schematic diagram of the principle of the member interaction relationship thrust module described in the embodiment of the present invention;
[0042] Figure 4 Schematic diagram of the principle of the attention mechanism described in the embodiment of the present invention. Detailed Embodiment
[0043] In order to more clearly understand the above-mentioned objects, features, and advantages of the present invention, the present invention will be further described below with reference to the drawings and embodiments. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0044] Embodiment: This embodiment proposes a group behavior recognition method based on adaptive multimodal fusion and implicit interactive relationship learning. The schematic diagram of its principle framework is shown in FIG. Figure 1 As shown, the design architecture includes a character-level feature extraction module, a multimodal feature fusion module, a member interaction relationship learning module, and a global feature extraction module. The method specifically includes the following steps:
[0045] Step A: Extracting static and dynamic character features in two streams: extracting static character posture features and dynamic optical flow features based on a character-level feature extraction module;
[0046] Step B, multimodal feature fusion: static posture features and dynamic optical flow features are concatenated and then convolutionally compressed to obtain a latent vector of significant information. Under the guidance of the latent vector and the original feature information loss function, iterative calculations are performed to strengthen the significant features, and the most representative feature information representation of each modality containing both optical flow and fine posture is obtained after fusion. At the same time, the redundant information and useless feature information that may exist in the two modalities are removed, the dimension of the fused feature vector is reduced, and the computational consumption for the learning of downstream interactive relationship modeling is reduced;
[0047] Step C, member interaction relationship learning: Based on the member feature vector obtained in step B that aggregates multimodal information after fusion, the self-attention mechanism in the Transformer encoder network is used to calculate the appearance similarity of paired character features through the correlation strength. The assigned similarity calculation score can spontaneously find the relationship between characters. This process can be regarded as an implicit interaction relationship learning process, which selectively extracts characters that are important for behavior recognition and better models the relationship between characters. The implicit vector representation between group members calculated in the form of the sum of attention weights is obtained;
[0048] Step D, global feature extraction: Based on the global feature extraction module, the global feature information including background information is extracted through the 3DCNN network;
[0049] Step E: Based on step C and step D, the output of step D is fed to a fully connected layer with multiple neurons and a tanh activation function. The output of this layer is connected to the Transformer encoder output of step C and passed as input to a fully connected layer with a Softmax classification layer. The features of the scene-level information differences in the global field of view are used to assist the person-level features in jointly identifying group behaviors.
[0050] The following is a detailed description of the embodiments of the present invention in conjunction with the specific operation mode:
[0051] 1. Dynamic and static dual-stream character feature extraction;
[0052] In this embodiment, a human-level feature extraction module is designed, and the dual-stream feature extraction method is adopted to enrich the features of individuals. For example, Figure 1 As shown, for the input video frames, they are divided into three branches, including a static branch, a dynamic branch, and a global branch. The static branch and the dynamic branch are connected to the human-level feature extraction module, and the global branch is connected to the global feature extraction module:
[0053] (1) The main network of the static branch is used to extract the 2D pose features of the human:
[0054] The actions of the human in the video all involve the movement of body joints, such as hands, arms, and legs. Human pose prediction is applicable not only to the fine action recognition performed in sports behaviors (such as spike and catch in the Volleyball dataset), but also to daily actions (such as walking and talking in CAD). The features of the human in the video not only need to capture the positions of the joints, but also the temporal dynamics of the joints. To obtain the joint positions of the human, this embodiment adopts the pose estimation model HRNet, which is used to accept the bounding box around the human in the video frame as the input and predict the key node positions. For example, in the experiment, the last layer feature map of this network can be used, and the smallest network pose_hrnet_w32 trained on the COCO dataset key points can be used to achieve good performance.
[0055] (2) The main network of the dynamic branch is responsible for modeling the 3D time dynamics:
[0056] Research shows that 3DCNNs with sufficient available data for training can build strong spatio-temporal representations for action recognition. Since the individual pose network in the static branch cannot capture the movement of the joints from a single frame, this embodiment adopts the I3D network and uses the optical flow representation in the I3D network. First, the input sequence of frames is converted into a sequence of continuous optical flow frames, and then the stacked frame sequence Ft, t = 1, …, T frames is processed by an inflated 3D convolutional network. Since the 3DCNN has a high computational cost, this embodiment adds an ROIAlign layer to project the coordinates onto the feature map of the frame, and then extracts the optical flow features of each human bounding box in the input frame.
[0057] In this embodiment, the 2D pose network and the 3DCNN main network are used to extract the features of the human in the continuous video frames respectively. By designing the human-level feature extraction module, not only the optical flow representing the dynamic features of the human in the 3DCNN is extracted, but also it is considered that in group behaviors, the actions of most people are highly related to the positions and movements of the body joints. The member pose features are extracted through HRNet to refine the subtle representation of the human actions. Through the fusion of the 2D pose features and the 3D optical flow dynamic feature dual-stream representation, rich static and dynamic human feature information is expressed.
[0058] II. Multi-modal Feature Fusion
[0059] Most existing fusion techniques, such as cascading and TFN, involve deterministic operations for constructing joint multimodal representations. For example, TFN uses the 3-fold Cartesian product of unimodal features for prediction. This method focuses more on learning rich unimodal features, resulting in the inability to effectively utilize multimodal information.
[0060] This embodiment designs a lightweight adaptive fusion technique. The basic principle of fusion is as follows: Using unimodal embedding cascading as the initial step, instead of using the cascaded vector as the final fusion result, by reconstructing the input information, extracting the most representative key information of multimodality, alleviating the "static nature" of existing fusion methods, effectively combining multimodal inputs, alleviating the "static nature", superficiality, and overhead problems of existing fusion methods.
[0061] As Figure 2 shown, the adaptive multimodal feature fusion module extracts multimodal features by maximizing the correlation between multimodal inputs, designs an encoder and a decoder. By learning to "compress" the information of different modalities, it restores the original data with a vector that has less information but contains all key information, and performs backpropagation through the prediction error to gradually improve the accuracy. Specifically, it includes the following steps:
[0062] 1. First, connect the single unimodal features and reduce the number of channels through the encoder convolutional network to achieve the effect of compressing the vector dimension, obtaining the self-fused latent vector;
[0063] (1) Embed the static pose feature and dynamic optical flow feature obtained by the person-level feature extraction module into a vector space with the same dimension through the Embedding linear mapping operation, and use them as unimodal inputs respectively. In the specific operation, set the dimension d = 256;
[0064] (2) Given n (n ≤ 3) d-dimensional multimodal latent vectors, where the two modalities represent the static pose feature and the dynamic optical flow feature vector respectively. First, perform a cascading operation to obtain where the dimension after cascading becomes 512;
[0065] (3) Then pass through the encoding part T to obtain reduce its dimension to t:
[0066] In the encoding part, first pass through the Linear layer to compress the dimension of the multimodal cascade to the dimension d = 256 of the unimodal initialization, and then perform a non-linear mapping through the Tanh activation function; then continue to perform a second Linear linear transformation to compress the feature dimension to t = 128, and then perform activation through the Relu function. At this time, It is called the fused latent feature (the most representative feature information representation). In this embodiment, by restricting the dimension to become smaller, learning this under-complete fused representation will enable the encoder to capture the most significant features in the training data while removing useless redundant information.
[0067] 2. Then try to reconstruct the initially connected vector from the self-fused latent vector;
[0068] During the specific operation, it is necessary to reproduce the input signal as much as possible. In order to achieve this reproduction, the network must automatically capture the most important factors that can represent the input data; in this embodiment, the fused latent features are reconstructed into the initially connected vector through the decoding transformation part F to obtain Calculate and the loss F between them tr to guide the iterative optimization of the network, so that the learned latent feature representation can best represent the significant information of each modality.
[0069] In the F transformation, the latent feature vector is first input After passing through a Linear linear transformation, the dimension is increased to 256. Subsequently, the Relu activation function is adopted, and then a second linear transformation is performed to change the dimension to the initially concatenated 512, and a reconstruction close to is obtained Here and are not exactly the same.
[0070] 3. Minimize the Euclidean distance between the original and reconstructed concatenated vectors. The training of this model can encourage the model to "compress" information without losing any basic clues, effectively increasing the correlation between the self-fused and concatenated latent vectors. Specifically:
[0071] The learning of the fusion network is guided by the loss function. In this embodiment, the mean squared error (MSE) loss function is adopted. The smaller the difference of the loss function, the better the reconstruction effect, which also means that the multi-modal latent features after compression and fusion The more significant the multi-modal information adaptively learned. The MSE loss is as follows:
[0072]
[0073] For the adaptive multi-modal fusion network, the intermediate vector is used as the fused multi-modal representation.
[0074] This embodiment proposes a lightweight multi-modal fusion strategy to learn and compress information of different modalities. By calculating the loss between the initial features and the approximated initial features reproduced through learning via an encoder and a decoder, it adaptively captures the effective information between 2D and 3D two-stream cross-modal human features, alleviating information redundancy.
[0075] III. Learning of Member Interaction Relationships
[0076] In the member interaction relationship learning module, the Transformer encoder architecture is applied to the challenging group behavior recognition of people in the video to learn and refine and aggregate person-level relationship features, as Figure 3 shown. The encoder receives the input sequence processed by a stack composed of the same layers, and this stack consists of a multi-head self-attention layer and a fully connected feed-forward network. Among them, the self-attention mechanism is an important part of the Transformer encoder network. In sequence modeling in the natural language field, this mechanism pays more attention to the most relevant words in the source sequence. Extended to group behavior recognition, this mechanism can be used to find the appearance similarity relationship between relevant people in group behaviors, enhance the feature information of each participant based on other participants in the video, so as to pay more attention to the key people in group behaviors, and no spatial constraints are required for the whole process.
[0077] In this example, we represent self-attention as A, that is, a function representing the weighted sum of values V. The specific process of learning group member relationships is as Figure 4 shown:
[0078] First, in the first stage, the score of the association degree of each person with other participants is calculated by querying Q to match the key-value set K. All three representations (Q, K, V) are calculated from the input sequence S through linear projection. S is a set of human features S = {s i | i = 1,…, N} obtained after multi-modal adaptive fusion by the human feature extractor, and A(S) = A(Q(S), K(S), V(S)). Q and each K value set represent the feature vectors of each participant. F is a matching function ( Figure 4 ) for calculating the appearance similarity of pairwise human features. This function can have different forms. In this embodiment, the dot product is used, which represents the projection length of one vector on another vector and can reflect the similarity between the two. The higher the similarity and association degree between human feature vectors, the higher the calculated score. And the dot product calculation is faster and more space-saving in practice and can be implemented using highly optimized matrix multiplication code. In form, the attention with the dot product matching function can be written as:
[0079]
[0080] d is the dimension of the query Q and the key value K.
[0081] In the second stage, after calculating the dot product of Q and K, the result of the correlation degree of each person with other participants is subjected to Softmax normalization to obtain a similarity set a n , n = 1, 2, …, n persons, and their sum is 1. The purpose is to convert the learned correlation degree values of each person with respect to the remaining participants into 0 - 1, representing the strength of the correlation between the person and other participants in the form of probability.
[0082] In the third stage, the similarity vectors obtained by Softmax in the second stage are multiplied by V respectively to obtain the final weighted sum attention matrix. This matrix can be regarded as the interaction relationship between group members implicitly learned through the self - attention mechanism and is used for the final classification and recognition.
[0083] In this embodiment, since the feature s i does not follow any specific order, the self - attention mechanism is more suitable for the refinement and aggregation of these features than RNN and CNN. The Transformer encoder can implicitly utilize the spatial relationship between persons through the positional encoding of s i . Relying only on the self - attention mechanism alleviates the problem of explicit modeling. The model uses the center point (x i , y i ) to represent each bounding box b i of each person feature s i , and encodes the center point using the same function PE (Position Encoding) as in the literature Attention is all you need.
[0084] To sum up, in view of the problems of inaccurate description of long - distance member interaction relationships and finding key person modeling, this embodiment proposes an implicit interaction relationship reasoning module without any explicit spatial and time modeling. By using the self - attention mechanism in the Transformer encoder to associate different positions of a single sequence to calculate the sequence representation, it learns the interaction between group members by calculating the appearance similarity of paired feature vectors through the association strength, that is, learning the appearance similarity of feature vectors between persons and the participant spatial structure information captured by positional encoding, assigning similarity calculation scores to selectively extract key person information important for behavior recognition, simulating and inferring the dependence relationship between persons, selectively paying attention to participants containing important information for behavior recognition, and at the same time being able to implicitly model the appearance and position relationships between persons without relying on any prior spatial or time structure.
[0085] IV. Global Feature Extraction
[0086] It is the same as the dynamic feature extraction mentioned in the feature extraction section (when extracting global features, there is no RoIAlign module. When extracting dynamic and static dual-stream human features, for the human bounding boxes in the dataset, only human features are focused on through RoIAlign, similar to image matting. Global features are the overall data features including background information). In this embodiment, I3D pre-trained on Kinetics is used as the backbone network, and RGB video clips are used as inputs. T frames centered on the annotated frame are selected, and the deep spatio-temporal feature maps extracted from the final convolutional layer are used as rich semantic representations describing the entire video clip. These deeper features provide low-resolution but high-level representations and are regarded as the scene features of the entire video clip.
[0087] V. Group behavior recognition:
[0088] To integrate the models (human-level model and global scene model), the output of the global feature extraction module is fed into a fully connected layer with multiple neurons and a tanh activation function. The output of this layer is connected to the Transformer output and passed as input to a fully connected layer with a Softmax classification layer for recognition.
[0089] In this embodiment, two classifiers are set during recognition and classification to generate group behavior category scores and individual action category scores respectively. The final recognition and classification results of group behavior and individual action are obtained through network learning. The overall network training process selects cross-entropy loss to guide optimization:
[0090]
[0091] where and represents the cross-entropy loss, and are the group behavior scores and individual action scores, while y g and y a represent the true labels of the target group behavior and individual behavior, and λ is a hyperparameter that balances these two terms.
[0092] The above are only the preferred embodiments of the present invention, and are not limitations on the present invention in other forms. Any person skilled in the relevant art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A group behavior recognition method based on multimodal fusion and implicit interaction relationship learning, characterized in that It includes the following steps: Step A, extraction of static and dynamic dual-stream human features: Based on the human-level feature extraction module, extract the static pose features and dynamic optical flow features of the human. Specifically: For the input video frames, they are divided into three branches, including a static branch, a dynamic branch, and a global branch; the static branch and the dynamic branch are connected to the human-level feature extraction module, and the global branch is connected to the global feature extraction module. The dual-stream feature extraction method is used to enrich the features of individuals: (1) The main network of the static branch extracts the static pose features of the human: Using the bounding box around the human in the video frame as the input of the pose estimation model HRNet to predict the key nodes; (2) The main network of the dynamic branch extracts the optical flow dynamic features: Using the optical flow representation in the I3D network, first convert the input sequence of frames into a continuous sequence of optical flow frames, then process the stacked frame sequence through an inflated 3D convolutional network, and add a ROIAlign layer to project the coordinates onto the feature map of the frame, thereby extracting the optical flow features of each human bounding box in the input frame; Step B, multi-modal feature fusion: Through the adaptive multi-modal feature fusion module, the static pose features and dynamic optical flow features are connected in a single-modal manner and then compressed by convolution to obtain a latent vector of significant information. First, the single-modal features are connected, and their channel numbers are reduced through the encoder convolutional network to obtain the self-fused latent vector; the initially connected vector is reconstructed from the self-fused latent vector, then the Euclidean distance between the original and reconstructed concatenated vectors is minimized, and the intermediate vector is used as the fused multi-modal feature information representation; thereby obtaining the fused feature information representation that contains both optical flow and the most representative feature information of each modality of the fine pose; Step C, learning of member interaction relationships: Using the fused feature information representation obtained in Step B, based on the self-attention mechanism, calculate the appearance similarity of pairwise human features through the association strength, and selectively extract the humans important for behavior recognition to obtain the implicit vector representation between group members calculated in the form of the sum of attention weights; Step D, global feature extraction: Based on the global feature extraction module, extract the global feature information containing background information for the input video frames; Step E, based on Step C and Step D, realize the recognition of group behaviors.
2. The group behavior recognition method based on multi-modal fusion and implicit interaction relationship learning according to claim 1, wherein: The specific implementation of Step B1 is as follows: (1) Embed the static pose features and dynamic optical flow features of the human obtained by the human-level feature extraction module into a vector space with the same dimension through Embedding linear mapping, and use them as single-modal inputs respectively; (2) Given n d-dimensional multimodal latent vectors, where n ≤ 3, let Two modalities represent the static pose features and the dynamic optical flow feature vectors of a person respectively. First, perform a concatenation operation to obtain where (3) Then, through the encoding part, we get Reduce its dimension to t: In the encoding part, first, the dimensions of the multi-modal concatenation are compressed through a Linear layer to become the dimensions of the single-modal initialization, and then a non-linear mapping is performed through the Tanh activation function; then, a second Linear linear transformation is continued to compress the feature dimensions, and then the Relu function is used for activation. At this time, it is called the fused latent feature. 3. The group behavior recognition method based on multimodal fusion and implicit interaction relationship learning according to claim 2, wherein: The specific implementation of Step B2 is as follows: The fused latent features are reconstructed into the original connection vectors through the decoding transformation part, obtaining and calculating the loss F between and to guide the iterative optimization of the network, so that the learned latent feature representation can best represent the significant information of each modality. tr 4. The group behavior recognition method based on multi-modal fusion and implicit interaction relationship learning according to claim 3, wherein: In the step B3, the MSE loss function is used to guide the learning of the fusion network, and the intermediate vector is used as the fused multi-modal feature information representation.
5. The group behavior recognition method based on multi-modal fusion and implicit interaction relationship learning according to claim 1, characterized in that: The specific steps of Step C include the following steps: First stage: Calculate the scores of the association degrees of each character with other participants by matching the query Q with the key-value set K. All three representations (Q, K, V) are calculated from the input sequence S through linear projection. S is a set of character features obtained by the character feature extractor after multi-modal adaptive fusion, S = {s i | i = 1, …, N}, and there is A(S) = A(Q(S), K(S), V(S)); In the second stage, the result of the dot product calculation of the query Q and the K points, which is the correlation degree of each person with other participants, is normalized to obtain a similarity set a n , where n = 1, 2, …, n persons, and their sum is 1; The third stage, multiply the similarity vectors obtained by normalizing in the second stage by V respectively to obtain the final weighted sum attention matrix for the final classification and recognition.
6. The method for group behavior recognition based on multimodal fusion and implicit interaction relationship learning according to claim 1, wherein: In Step D, use I3D as the backbone network, take the RGB video clip as the input, select T frames centered on the annotated frame, and use the deep spatio-temporal feature map extracted from the final convolutional layer as the rich semantic representation describing the entire video clip.
7. The group behavior recognition method based on multi-modal fusion and implicit interaction relationship learning according to claim 1, wherein: In step E, the output of step D is fed to a fully connected layer. The output of this fully connected layer is connected to the output of the member interaction learning module in step C and is passed as an input to a fully connected layer with a classification layer. By utilizing the feature of the difference in scene-level information in the global view, it assists the person-level features to jointly perform group behavior recognition. When performing recognition and classification, two classifiers are set respectively to generate the group behavior category score and the individual action category score. The final recognition and classification results of the group behavior and the individual action are obtained through network learning. The overall network training process selects the cross-entropy loss to guide the optimization.
Citation Information
Patent Citations
Volleyball group behavior identification method based on multi-modal information fusion
CN111401174A