Group behavior recognition method based on multi-scale feature extraction
By adopting multi-scale feature extraction and interactive relationship analysis methods in group behavior recognition technology, the problem of extracting multi-scale features and molecular groups in the prior art is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510118192.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Existing group behavior recognition technologies are difficult to effectively extract multi-scale individual characteristics and relationship characteristics, and there are challenges in demarcating groups and eliminating interfering individuals in large-scale groups.
The group behavior recognition method based on multi-scale feature extraction is adopted to extract and analyze individual characteristics and relationship characteristics in sensor data by constructing an identification network model, including multi-scale feature extraction module, interactive relationship extraction module, subgroup cardinality prediction module, graph clustering module and group behavior classifier.
It improves the robustness and accuracy of group behavior recognition, can more effectively understand and identify group behavior patterns, and enhances the generalization ability of the model.
Smart Images

Figure CN120108034A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of behavior recognition technology, and specifically relates to a group behavior recognition method based on multi-scale feature extraction. Background Art
[0002] Group behavior recognition based on sensor data is a research direction that has attracted much attention and has broad prospects. This field aims to understand the overall behavior pattern formed by the interaction between multiple individuals. Its research results have broad application value in many fields such as urban planning, traffic management, and sociological research. Group behavior recognition in a complex environment is a very challenging task. This is mainly because group behavior is not a simple superposition of individual behavior, but requires a comprehensive analysis of individual behavior and the complex interaction relationship between individuals, so as to achieve bottom-up behavior pattern inference.
[0003] In recent years, with the rapid development of Internet of Things technology and the popularization of wireless sensor networks and wearable devices, the widespread application of various sensors has provided a rich source of data for group behavior recognition. At the same time, the widespread application of machine learning and deep learning technologies has significantly promoted the research on group behavior recognition based on sensor data. However, most of the current group behavior research is still concentrated in the fields of vision and image processing. In contrast, group behavior recognition based on sensor data has the advantages of low cost, no geographical restrictions, and strong privacy protection. Nowadays, smart terminal devices have generally integrated a variety of sensor modules, such as accelerometers, magnetometers, gyroscopes, and global positioning systems, which provides reliable technical support and practical feasibility for using smart terminals for group behavior recognition.
[0004] Although the widespread application of various sensor technologies and the introduction of deep learning methods have provided rich data resources and theoretical basis for group behavior recognition research, there are still many challenges that need to be solved in this field. First, the feature extraction of group behavior is still a complex problem. Existing feature extraction methods usually only extract individual features and relationship features between individuals at one scale. How to mine and extract individual features and relationship features at different scales, and design an effective model structure on this basis to improve the performance of group behavior recognition is an issue that needs to be explored. Secondly, large-scale groups usually contain several sub-groups and interfering individuals. How to reasonably divide sub-groups and exclude interfering individuals to help the model understand group behavior in large-scale scenarios is still a challenging task. Summary of the invention
[0005] The purpose of this application is to provide a group behavior recognition method based on multi-scale feature extraction to improve the robustness and accuracy of group behavior recognition.
[0006] To achieve the above purpose, the technical solution adopted in this application is:
[0007] A group behavior recognition method based on multi-scale feature extraction, comprising:
[0008] Constructing and training a recognition network model, wherein the recognition network model includes a multi-scale feature extraction module, an interaction relationship extraction module, a subgroup cardinality prediction module, a graph clustering module, and a group behavior classifier;
[0009] The collected sensor data is preprocessed, and the preprocessed sensor data is used to obtain individual refined features through a multi-scale feature extraction module;
[0010] The refined features of individuals are input into the interaction relationship extraction module to capture the interaction relationship between individuals and obtain a refined feature vector;
[0011] The number of subgroups in the refined feature vector is extracted through the subgroup cardinality prediction module, and the feature adjacency matrix of the refined feature vector is constructed. The matrix is input into the graph clustering module for subgroup division and the global features of the subgroup are generated. The recognition result is then obtained through the group behavior classifier.
[0012] Wherein, the multi-scale feature extraction module performs the following operations:
[0013] Perform channel upsampling on the preprocessed sensor data;
[0014] The upsampled sensor data is evenly divided in the channel dimension according to the preset scale number to obtain channel features;
[0015] Scale features are extracted for each channel feature to obtain the scale features corresponding to each channel feature, and then spliced to obtain the target scale features;
[0016] The target scale features are input into an efficient channel attention network to obtain individual refined features.
[0017] Furthermore, the number of the channel features is four, and the extraction of different scale features is performed on each channel feature, and then the features are spliced to obtain the target scale feature, including:
[0018] The first channel feature is directly used as part of the target scale feature through a long-skip connection;
[0019] The second channel feature undergoes a general convolution operation and passes through a linear layer to extract the second channel feature hidden feature. The first half of the dimension in the hidden feature will be merged with the third channel feature after convolution, and the information of the other half of the dimension is used as part of the target scale feature.
[0020] The third channel feature undergoes a convolution operation and is merged with the first half dimension of the second channel feature hidden feature, and then further extracted through the convolution layer and the linear layer to obtain the third channel feature hidden feature. The first half dimension of the third channel feature hidden feature will be merged with the next channel feature after convolution, and the second half dimension is part of the target scale feature;
[0021] The fourth channel feature undergoes a convolution operation and is merged with the first half dimension of the third channel feature hidden feature, and then further extracted through the convolution layer and the linear layer to obtain the fourth channel feature hidden feature as part of the target scale feature;
[0022] Finally, the target scale features are obtained by splicing.
[0023] Furthermore, the interaction relationship extraction module performs the following operations:
[0024] Each individual refined feature is evenly divided according to the preset scale number in the channel dimension to obtain the scale feature;
[0025] With individuals as nodes and similarity as edge weights, a graph structure is constructed for each scale feature;
[0026] The graph structure is input into the graph transformer module to update the node features, and then input into the feedforward neural network for nonlinear transformation to obtain scale-enhanced features;
[0027] The scale-enhanced features are recombined to obtain individual-enhanced features;
[0028] The individual enhanced features are input into the graph Transformer module to further extract the deep relationship between features of different scales, and then normalized to obtain a refined feature vector.
[0029] Furthermore, the subgroup cardinality prediction module performs the following operations:
[0030] The input features first undergo a linear transformation, then the ReLU activation function is used to increase the model's expressiveness, and finally another layer of linear transformation is performed to output the predicted value of the number of subgroups.
[0031] Furthermore, the graph clustering module performs the following operations:
[0032] Based on the identity matrix and the characteristic adjacency matrix, a Laplacian matrix is constructed, and then characteristic decomposition is performed to obtain the node embedding matrix;
[0033] Apply clustering algorithm to individuals in the node embedding matrix to complete subgroup division;
[0034] The individual refined feature vectors of each subgroup are aggregated through an average pooling operation to obtain the global feature vector of each subgroup.
[0035] Furthermore, the loss functions used in the training recognition network model include: subgroup cardinality loss function, subgroup member loss function and group behavior recognition loss function.
[0036] The present application proposes a group behavior recognition method based on multi-scale feature extraction. In the process of group behavior recognition, a multi-scale feature extraction module is used to extract and expand individual features in sensor data to obtain individual behavior information, and an interactive relationship extraction module is used to fully mine the relationship characteristics between individuals. The potential behavior patterns in the sensor data are mined in combination with the subgroup division task, which enhances the generalization ability of the model, improves the robustness and accuracy of group behavior recognition, and realizes the effective recognition of group behavior by the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a flow chart of the group behavior recognition method based on multi-scale feature extraction in this application.
[0038] Figure 2 A schematic diagram of the network model structure is identified for this application.
[0039] Figure 3 Schematic diagram of the structure of the multi-scale feature extraction module of the embodiment of the present application.
[0040] Figure 4 This is a structural diagram of the interactive relationship extraction module of an embodiment of the present application. DETAILED DESCRIPTION
[0041] The technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all of the embodiments. Based on the embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0043] An embodiment of the present application, such as Figure 1 As shown, a group behavior recognition method based on multi-scale feature extraction is provided, including:
[0044] Step S1, constructing and training a recognition network model, wherein the recognition network model includes a multi-scale feature extraction module, an interaction relationship extraction module, a subgroup cardinality prediction module, a graph clustering module and a group behavior classifier.
[0045] The recognition network model constructed in this embodiment is as follows Figure 2 As shown, it includes: a multi-scale feature extraction module, an interaction relationship extraction module, a subgroup cardinality prediction module, a graph clustering module and a group behavior classifier.
[0046] Step S2: preprocessing the collected sensor data, and obtaining individual refined features from the preprocessed sensor data through a multi-scale feature extraction module.
[0047] The embodiment preprocesses the collected sensor data by using a sliding window segmentation method to divide a longer time series sample into a plurality of shorter time series segments.
[0048] In this embodiment, the sensor needs to be worn on the wrist of the experimental individual. The sensor includes an accelerometer and a gyroscope, which can record the acceleration and angular velocity data of the x, y, and z axes respectively. In this embodiment, 50 Hz is selected as the data acquisition frequency, and a complete set of data is collected every 4 seconds. For the collected data, the sliding window segmentation method is used for processing. The sliding window length is set to 200, and the overlap rate is 50%. The sliding window starts from the starting position of the time series and slides backward step by step, generating a subsequence of length 200 each time, until the long time series sample is divided into multiple short time series fragments, and each short time series fragment is counted as a data X.
[0049] The preprocessed sensor data is input into the multi-scale feature extraction module to obtain individual refined features. The multi-scale feature extraction module mainly includes channel upsampling, channel partitioning, multi-scale convolution and efficient channel attention network ECANet.
[0050] In this embodiment, the multi-scale feature extraction module is as follows: Figure 3 As shown, perform the following operations:
[0051] Step 2.1: Perform channel upsampling on the preprocessed sensor data.
[0052] In order to apply 2D convolution to sensor data, the sensor input data is first upsampled in the channel dimension, converting (1,T,C 0 )-dimensional input data X is expanded to (H, W, C)-dimensional data X′, where T is the time span and C is the 0is the initial number of channels, C is the number of channels after expansion, H and W are the height and width after the sensor data is expanded to two dimensions. Through channel upsampling operations, the feature expression ability of the model can be enhanced, and the embedding enhancement of low-dimensional features can be achieved. The formula can be expressed as:
[0053]
[0054] Where UpSampling represents the upsampling operation.
[0055] Step 2.2: Evenly divide the upsampled sensor data in the channel dimension according to the preset scale number to obtain channel features.
[0056] In order to perform multi-scale feature extraction in the following step, the present embodiment first presets the number of scales n to be divided, that is, the sensor data after upsampling is evenly divided into n parts according to the number of scales n in the channel dimension. The formula can be expressed as:
[0057]
[0058] Where ChannelSplit represents a channel splitting operation. In this embodiment, the number of scales n is 4, the value of i is 0, 1, 2 or 3, and the initial number of channels of each scale is 8.
[0059] Step 2.3: Extract the scale feature of each channel feature separately to obtain the scale feature corresponding to each channel feature, and then splice them to obtain the target scale feature.
[0060] Each channel features Scale features are extracted one by one, s i Represents scale information.
[0061] Take four channel features as an example, Figure 3 The first channel feature Without any processing, it is directly used as part of the target scale feature Y through the far jump link The long-jump link operation solves the gradient vanishing problem in model training, allowing the model to have a deeper network structure, retaining important information in the initial features, and accelerating the convergence of the model.
[0062] The second channel feature undergoes a general convolution operation and passes through a linear layer Thus, the hidden features of the second channel features are extracted The first half of the dimension in will be merged with the third channel feature after convolution, and the information of the other half of the dimension will be used as part of the target scale feature Y In this process, the convolution operation is used to extract the hidden information in the current channel features, and the linear layer In essence, it is a weight matrix, which will be updated according to the gradient of the target loss function during the back-propagation process, and its parameters will be automatically adjusted to select the feature dimension that has a strong information contribution to the next channel feature, and realize the weight update of each dimension in the current channel feature.
[0063] The third channel feature After a convolution operation, and The first half of the dimension is merged and then passed through the convolutional neural network and linear layer Further extraction to obtain hidden features Hidden Features The first half of the dimension in will be merged with the next channel feature after convolution, and the second half of the dimension will be part of the target scale feature Y In this process, the linear layer Its parameters are automatically adjusted to select feature dimensions that have a strong information contribution to the next channel features.
[0064] The fourth channel feature After a convolution operation, the features are hidden with the third channel features. The first half dimension in is merged, and then further extracted through the convolution layer and the linear layer to obtain the fourth channel feature hidden feature As part of the target scale feature
[0065] Above and That is, the scale feature corresponding to each channel feature.
[0066] It should be noted that if the number of scales is larger, the steps of subsequent channels are similar to the third channel feature. No longer segmented, but all as part of the target scale feature Y The process of scale feature extraction can be expressed by the following formula:
[0067]
[0068]
[0069] Among them, Conv is a 2D convolution, and the convolution kernel size is set to 1. and To take the first and second halves of the tensor dimension, Concat represents the tensor concatenation operation. is a linear layer, and Y represents the target scale feature. The target scale feature Y contains features of different scales, which enhances the expressiveness of the feature.
[0070] It should be noted that in the multi-scale feature extraction module, each branch corresponding to each channel feature will select a convolution kernel of different sizes. The first part of the feature is linked and directly used as part of the target feature. Each subsequent one will pass through a convolution layer, and the convolution kernel size will increase successively. The features obtained will have receptive fields of different sizes, that is, scale features of different sizes.
[0071] Step 2.4: Input the target scale features into the efficient channel attention network to obtain individual refined features.
[0072] Then the target scale feature Y is input into the efficient channel attention network ECANet to refine the channel features. The main process of ECANet is as follows: First, global average pooling is performed on each channel to generate a channel description vector h t This process compresses the spatial dimension and only retains the channel information. t Input a one-dimensional convolution layer to capture the local interaction relationship between channels. The adaptive selection of the convolution kernel size k ensures that the local interaction range is moderate. In this embodiment, the convolution kernel size k is selected to be 3. The output of the one-dimensional convolution is then mapped to the [0,1] interval using the sigmoid activation function to generate channel attention weights. The attention weights are applied to the target scale feature Y to highlight important channels and suppress irrelevant channels. The formula of ECANet can be expressed as:
[0073] h t =AvgPooling(Y) (5)
[0074]
[0075]
[0076] Where AvgPooling is the average pooling operation. Conv1D is a 1D convolutional neural network with a convolution kernel size of j. σ is the sigmoid function. The individual refined feature h is the final output result of the multi-scale feature extraction module.
[0077] Step S3: Input the refined individual features into the interaction relationship extraction module to capture the interaction relationship between individuals and obtain a refined feature vector.
[0078] There may be complex interdependencies between information at different scales. If they are analyzed as a whole, important information may be diluted. Figure 4As shown in the figure, the interactive relationships of multiple individual refined features input are extracted through the interactive relationship extraction module, so that the model can more accurately capture the features of a specific scale, enhance the flexibility and controllability of the model, and enhance the interpretability of the model.
[0079] The interactive relationship extraction module of this embodiment performs the following operations:
[0080] Step 3.1: Segment each individual refined feature in the channel dimension according to the preset scale number to obtain scale features.
[0081] In this step, multiple individual refined features of the input are divided into n parts in the channel dimension to obtain scale features Each of these features m represents the individual number, the total number of individuals is N, s i Represents scale information, scale characteristics The channel dimension and The channel dimension is consistent with Represents the scale information i This is because the localization operation adopted by ECANet when calculating the channel attention weight ensures that each scale information can be represented independently on the channel without being fused with other scale features, and is only used as a way to update the weight.
[0082] Step 3.2: Using individuals as nodes and similarity as edge weights, construct a graph structure for each scale feature.
[0083] In order to model the relationship between individuals and extract higher-level group features, this step constructs a graph-based network structure, where the nodes of the graph represent individuals in the group, and the node features are The edges represent the relationships between individuals. Figure 4 As shown in Figure 2, for each scale feature, the corresponding graph structure is constructed by calculating the similarity based on the individual relationship features. in is a set of nodes, each node corresponds to each individual in the group; is a set of edges used to represent the dependency relationship between individuals. The weight of the edge is obtained by calculating the similarity between the features of individuals of the same scale. The calculation formula is expressed as:
[0084]
[0085] Where ∈ is a positive number greater than 0, which is used to avoid division by zero errors in similarity calculation. Represent the characteristics of different individuals in the scale feature. Through the above method, a weighted undirected graph is constructed in is the edge weight matrix.
[0086] Step 3.3: Input the graph structure into the graph transformer module to update the node features, and then input it into the feedforward neural network for nonlinear transformation to obtain scale-enhanced features.
[0087] For each construction completed at different scales It is processed by two-layer graph transformer modules at different scales to update node features. Each graph transformer contains a multi-head attention mechanism to enhance the feature capture capability of the model. Multi-head attention processes the input node features through multiple parallel attention heads and maps them to the output space after splicing, thereby capturing different channel dependency patterns. Subsequently, the node features updated by the graph attention module are input into the feedforward neural network (FFN) for nonlinear transformation to obtain scale-enhanced features.
[0088] Step 3.4: Recombine the scale-enhanced features to obtain individual-enhanced features.
[0089] For each scale, after processing by graph Transformer and FFN, the updated scale-enhanced features are output, where the features of each individual are represented as p′ mi Recombine the individual into each individual enhanced feature, using p′ m express.
[0090] Step 3.5: Input the individual enhanced features into the Transformer module to further extract the deep relationship between features of different scales, and then perform normalization to obtain a refined feature vector.
[0091] Individual enhancement feature p′ m The global graph Transformer module is used to further extract the deep relationship between features of different scales, and then the LayerNorm normalization process is performed to obtain the final output refined feature vector Y of the model. norm (Also called normalized features). The LayerNorm normalization operation reduces the risk of gradient explosion or vanishing during training and accelerates model convergence by standardizing feature distribution.
[0092] Through the interaction relationship extraction module, efficient fusion of multi-scale features can be achieved, and the interaction relationship between individuals can be fully captured, thereby improving the model's feature expression ability and adaptability to complex scenarios.
[0093] It should be noted that the graph Transformer model, feedforward neural network and LayerNorm normalization operation are relatively mature technologies in this field and will not be described in detail here.
[0094] Step S4: extract the number of subgroups in the refined feature vector through the subgroup cardinality prediction module, construct the feature adjacency matrix of the refined feature vector, input it into the graph clustering module for subgroup division, generate the global features of the subgroup, and then obtain the recognition result through the group behavior classifier.
[0095] Refined feature vector Y norm The individual features and relationship features in the input features have been fully extracted. This step uses the refined feature vector Y norm Construct a feature adjacency matrix, determine the number of subgroups through the subgroup cardinality prediction module, use the graph clustering module to implement subgroup division, and finally obtain the group behavior recognition result through the group behavior classifier.
[0096] First, we use the group to refine the feature vector Y norm Construct the feature adjacency matrix A. The purpose of constructing the feature adjacency matrix is to express the relationship between individuals in the group. If the feature of each individual in the group refined feature vector is p′ m , a in the feature adjacency matrix ij It represents the relationship strength between the i-th and j-th individuals. The calculation formula of relationship strength is as follows:
[0097]
[0098] In order to improve numerical stability and adapt to subsequent graph operations, the adjacency matrix will be normalized to obtain a symmetric normalized adjacency matrix. The calculation formula is:
[0099]
[0100] Where D is the degree matrix, D ii =∑ j a ij The normalized adjacency matrix is used to represent the global relationship between individuals in a group.
[0101] This embodiment also refines the feature vector Y based on the group norm , the number of subgroups K in the input group is estimated through the subgroup cardinality prediction module. In the subgroup cardinality prediction module, the input features first undergo a linear transformation, then the ReLU activation function is used to increase the model's expressiveness, and finally another layer of linear transformation is performed to output the subgroup number prediction value K. The calculation formula is:
[0102] K=σ(W 2 ReLU(W 1 ·Ynorm +b 1 )+b 2 ) (11)
[0103] Among them, W 1 ,W 2 is the weight matrix, b 1 ,b 2 is the bias term, and σ is the sigmoid activation function.
[0104] Then, based on the predicted number of subgroups K and the adjacency matrix Graph clustering is performed to divide the subgroups in the sample and obtain the global characteristics of the subgroups.
[0105] In a specific embodiment, the graph clustering module performs the following operations:
[0106] Step 4.1: Based on the identity matrix and the characteristic adjacency matrix, construct the Laplacian matrix, and then perform eigendecomposition to obtain the node embedding matrix.
[0107] In this embodiment, graph clustering models the graph structure of the group through the Laplace matrix L, and obtains the first K eigenvectors by eigendecomposing the Laplace matrix. These eigenvectors constitute a low-dimensional node embedding matrix. It can be expressed as:
[0108]
[0109] U=eig(L,K) (13)
[0110] Among them, U is the node embedding matrix composed of eigenvectors, I is the identity matrix, and eig represents the eigenvalue decomposition of the matrix to calculate the eigenvalues and eigenvectors of the matrix.
[0111] Step 4.2: Apply the clustering algorithm to the individuals in the node embedding matrix to complete the subgroup division.
[0112] Then, the K-Means clustering algorithm is applied to the points (i.e., individuals) in the node embedding matrix to assign each individual to one of the K subgroups, thus completing the subgroup division.
[0113] Step 4.3: Aggregate the individual refined feature vectors of each subgroup through an average pooling operation to obtain the global feature vector of each subgroup.
[0114] After the subgroup division is completed, the network further integrates the division results to generate a global feature representation of the subgroup. The individual features within each subgroup are aggregated through an average pooling operation to obtain a global feature vector for each subgroup. The subgroup feature vector summarizes the behavior pattern or characteristics of the entire subgroup. The formula is expressed as:
[0115]
[0116] where c i is the subgroup label to which the ith individual belongs, k represents the index of the subgroup, is the refined feature vector of the i-th individual.
[0117] Finally, the global feature vector of the subgroup is fed into the group behavior classifier, which uses a fully connected layer plus Softmax structure to identify and classify the behavior of each subgroup. The classification result is the probability distribution of each subgroup in the behavior category.
[0118] Another embodiment of the present application, when training the recognition network model, the loss function in this embodiment consists of two parts. Since subgrouping can help the network learn the interaction characteristics between individuals from the spatial and relational levels, individuals belonging to the same subgroup often have similar behavior patterns or interaction methods. This belonging information can supplement the implicit semantics in behavior recognition. Therefore. The model will simultaneously learn the intrinsic patterns of subgrouping and group behavior, and use the characteristics of the subgrouping task as a shared representation of the group behavior recognition task, allowing the model to focus on more representative local relationships, thereby providing more focused contextual information for group behavior recognition. The subgrouping task predicts the subgroup cardinality of the group, and based on the subgroup cardinality, performs graph clustering on individual hidden features to obtain the final classification result. The loss function during the training of the subgrouping task is defined as follows:
[0119]
[0120]
[0121] in, is the subgroup cardinality loss function, using mean square error, is the predicted subgroup cardinality of the ith sample, K i is the true subgroup cardinality of the ith sample. is the subgroup membership loss function, using binary cross entropy loss, is the predicted membership relationship, with a value range of [0,1], M ij is the true membership relationship, with a value of 0 or 1, and N is the total number of samples.
[0122] In this embodiment, a group behavior recognition task is performed by a classifier, which is composed of a layer of fully connected neural network and a Softmax activation function. In this embodiment, the group behavior recognition loss function uses cross entropy loss to measure the difference between the predicted group behavior category distribution and the actual category distribution. The loss function can be defined as follows:
[0123]
[0124] in, is the predicted probability of the i-th sample in the c-th category, P ic is the true distribution of the i-th sample in the c-th category, C is the number of behavior categories, and N is the total number of samples.
[0125] The total loss function of this embodiment is It can be expressed as the following formula:
[0126]
[0127] Among them, λ 1 ,λ 2 Represents the weight coefficient, which can control the contribution of the three losses to model training, thereby optimizing the overall performance of the model.
[0128] In this application, the technical solution of this application is also verified through experiments. The experiment is based on the UT-Group dataset reconstructed from UT-Data and the WBSensor dataset produced by itself, and is compared with the current mainstream subgrouping algorithms and group behavior recognition algorithms. The results of the subgrouping task are measured using F1 score, precision, recall rate, mean average precision mAP, and IOU-AOC indicators, and the group behavior recognition task is measured using F1 score, precision, recall rate, and accuracy indicators. The experimental results are shown in Tables 1 and 2. Table 1 shows the experimental results of subgrouping, and Table 2 shows the experimental results of group behavior recognition:
[0129] Table 1
[0130]
[0131] Table 2
[0132]
[0133] The experimental results of the proposed method are divided into two parts, which evaluate the performance of the model in the subgroup division and group behavior recognition tasks. In the subgroup division task, the proposed method has an F1 value of 58.40%, a precision of 55.30%, a recall of 62.00%, a mean average precision of 77.60%, and an IOU-AOC of 38.30% on UT-Group; the F1 value on the WBSensor dataset is 64.16%, the precision is 62.12%, the recall is 66.35%, the mean average precision is 81.67%, and the IOU-AOC is 40.77%. The proposed method outperforms other methods in all indicators. In the group behavior recognition task, the F1 value of the proposed method on UT-Group is 84.44%, the precision is 86.12%, the recall is 82.83%, and the accuracy is 84.52%; the F1 value on the WBSensor dataset is 88.37%, the precision is 89.21%, the recall is 87.56%, and the accuracy is 93.86%. In all indicators of this task, the proposed method is also better than other methods.
[0134] In summary, the method of the present invention has certain advantages over other algorithms. The method of the present invention can more comprehensively understand the group behavior reflected by the sensor data, and efficiently extract group features from the sensor data, thereby significantly improving the accuracy and robustness of group behavior recognition.
[0135] The above-mentioned embodiments only express several implementation modes of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.
Claims
1. A group behavior recognition method based on multi-scale feature extraction, characterized in that: The group behavior recognition method based on multi-scale feature extraction includes: Constructing and training a recognition network model, wherein the recognition network model includes a multi-scale feature extraction module, an interaction relationship extraction module, a subgroup cardinality prediction module, a graph clustering module, and a group behavior classifier; The collected sensor data is preprocessed, and the preprocessed sensor data is used to obtain individual refined features through a multi-scale feature extraction module; The refined features of individuals are input into the interaction relationship extraction module to capture the interaction relationship between individuals and obtain a refined feature vector; The number of subgroups in the refined feature vector is extracted through the subgroup cardinality prediction module, and the feature adjacency matrix of the refined feature vector is constructed. The matrix is input into the graph clustering module for subgroup division and the global features of the subgroup are generated. The recognition result is then obtained through the group behavior classifier. Wherein, the multi-scale feature extraction module performs the following operations: Perform channel upsampling on the preprocessed sensor data; The upsampled sensor data is evenly divided in the channel dimension according to the preset scale number to obtain channel features; Scale features are extracted for each channel feature to obtain the scale features corresponding to each channel feature, and then spliced to obtain the target scale features; The target scale features are input into an efficient channel attention network to obtain individual refined features.
2. The method for group behavior recognition based on multi-scale feature extraction according to claim 2 is characterized in that: The number of the channel features is four, and each channel feature is respectively subjected to different scale feature extraction, and then spliced to obtain the target scale feature, including: The first channel feature is directly used as part of the target scale feature through a long-skip connection; The second channel feature undergoes a general convolution operation and passes through a linear layer to extract the second channel feature hidden feature. The first half of the dimension in the hidden feature will be merged with the third channel feature after convolution, and the information of the other half of the dimension is used as part of the target scale feature. The third channel feature undergoes a convolution operation and is merged with the first half dimension of the second channel feature hidden feature, and then further extracted through the convolution layer and the linear layer to obtain the third channel feature hidden feature. The first half dimension of the third channel feature hidden feature will be merged with the next channel feature after convolution, and the second half dimension is part of the target scale feature; The fourth channel feature undergoes a convolution operation and is merged with the first half dimension of the third channel feature hidden feature, and then further extracted through the convolution layer and the linear layer to obtain the fourth channel feature hidden feature as part of the target scale feature; Finally, the target scale features are obtained by splicing.
3. The method for group behavior recognition based on multi-scale feature extraction according to claim 1 is characterized in that: The interaction relationship extraction module performs the following operations: Each individual refined feature is divided according to the preset scale number in the channel dimension to obtain the scale feature; With individuals as nodes and similarity as edge weights, a graph structure is constructed for each scale feature; The graph structure is input into the graph transformer module to update the node features, and then input into the feedforward neural network for nonlinear transformation to obtain scale-enhanced features; The scale-enhanced features are recombined to obtain individual-enhanced features; The individual enhanced features are input into the graph Transformer module to further extract the deep relationship between features of different scales, and then normalized to obtain a refined feature vector.
4. The method for group behavior recognition based on multi-scale feature extraction according to claim 1 is characterized in that: The subgroup cardinality prediction module performs the following operations: The input features first undergo a linear transformation, then the ReLU activation function is used to increase the model's expressiveness, and finally another layer of linear transformation is performed to output the predicted value of the number of subgroups.
5. The method for group behavior recognition based on multi-scale feature extraction according to claim 1 is characterized in that: The graph clustering module performs the following operations: Based on the identity matrix and the characteristic adjacency matrix, a Laplacian matrix is constructed, and then characteristic decomposition is performed to obtain the node embedding matrix; Apply clustering algorithm to individuals in the node embedding matrix to complete subgroup division; The individual refined feature vectors of each subgroup are aggregated through an average pooling operation to obtain the global feature vector of each subgroup.
6. The method for group behavior recognition based on multi-scale feature extraction according to claim 1, characterized in that: The training recognition network model adopts loss functions including: subgroup cardinality loss function, subgroup member loss function and group behavior recognition loss function.
Citation Information
Patent Citations
Crowd grouping detection method based on split-merge strategy
CN104951806A
Handset sensor-based group behavior analysis method
CN106940805A
Group behavior recognition method based on graph convolutional network and group relationship modeling
CN114781638A
Sensor data group behavior recognition method based on multistage feature enhancement
CN117807494A
Swarm control systems and methods
US8379967B1
Cited By
Group behavior identification method and device based on multi-task self-supervision framework
CN120354157A
A method and device for group behavior recognition based on a multi-task self-supervised framework
CN120354157B