A method and device for group behavior recognition based on a multi-task self-supervised framework
Through the group behavior recognition method of the multi-task self-supervising framework, individual behavior characteristics are extracted from the multi-modal sensor data using the Transformer encoder and task head, and attention is calculated inside and outside the cluster through the global feature supplement module, the accuracy and robustness of group behavior recognition under self-supervised learning are solved, and more efficient group segmentation and behavior recognition are achieved.
Patent Information
- Application Number
- CN202510858515.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-06-25
AI Technical Summary
In large-scale group scenarios, how to adaptively divide groups and effectively filter noise individuals to improve the model's understanding of complex scenarios, especially the accuracy and robustness of group behavior recognition under the self-supervised learning framework.
The multi-task self-supervision framework is adopted, and an individual behavior understanding module is formed through a Transformer encoder and multiple task heads. Individual behavior characteristics are extracted from the multi-modal sensor data, and attention is calculated within and between clusters through the global feature supplement module, reducing the computational complexity and capturing global information between individuals.
It significantly improves the accuracy and robustness of group segmentation and behavior recognition, reduces the computational complexity, and can capture global information between individuals, which is highly explanatory.
Smart Images

Figure CN120354157B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of behavior recognition technology, and specifically relates to a group behavior recognition method and device based on a multi-task self-supervision framework. Background Art
[0002] Crowd behavior analysis based on sensor data has become a research hotspot at the intersection of artificial intelligence and the Internet of Things. This technology, by capturing interaction patterns among multiple individuals and revealing the dynamic characteristics of group behavior, has important applications in smart cities, intelligent transportation, social computing, and other fields. However, analyzing crowd behavior in real-world scenarios still faces significant challenges. This is primarily because crowd behavior is not a simple summation of individual behaviors but rather requires the integration of individual behavior data with complex interactions, requiring a modeling framework that spans from the local to the global level.
[0003] In recent years, with the widespread adoption of the Internet of Things (IoT), particularly wireless sensor networks and wearable devices, massive amounts of sensor data have provided a rich source of information for group behavior research. Furthermore, advances in deep learning have further fueled the field. Compared with traditional vision-based approaches, sensor data solutions offer advantages such as low cost, wide applicability, and strong privacy protection. Furthermore, the widespread integration of multimodal sensors in modern smart terminals provides reliable technical support for real-time monitoring and analysis of group behavior, significantly enhancing the practicality and scalability of this research.
[0004] Although group behavior analysis based on sensor data shows great promise, the field still faces many challenges. First, sensor data is usually huge in volume, and much of the data lacks clear labels, making traditional supervised learning methods difficult to apply directly. Therefore, self-supervised learning frameworks have gradually become a research hotspot. They learn universal representations from unlabeled data by designing pre-training tasks, and then fine-tune them in downstream tasks to complete the given task. This approach effectively alleviates the problem of insufficient labeled data. In addition, in large-scale group scenarios, the group structure is often more complex. This is because large-scale groups usually contain multiple dynamically changing subgroups, as well as irrelevant interfering individuals. How to adaptively divide subgroups and effectively filter out noisy individuals to improve the model's ability to understand complex scenarios remains one of the difficulties in current research. Summary of the Invention
[0005] The purpose of this application is to provide a group behavior recognition method based on a multi-task self-supervised framework to improve the accuracy and robustness of group segmentation and behavior recognition under self-supervised conditions.
[0006] In order to achieve the above objectives, the technical solutions of this application are as follows:
[0007] A group behavior recognition method based on a multi-task self-supervised framework includes:
[0008] Obtaining individual behavioral data to generate a first training data set and a second training data set;
[0009] Constructing a first network model including an individual feature extraction module, a Transformer encoder, and a task head, and training the first network model using a first training dataset. During training, the individual embedding features extracted by the individual feature extraction module are input into the Transformer encoder to extract the encoded features, which are then input into each task head. The loss corresponding to each task head is calculated to complete the training of the first network model.
[0010] Constructing a second network model including an individual feature extraction module, a Transformer encoder, a global feature supplementation module, and an inference module, freezing the parameters of the Transformer encoder, and training the second network model using the second training dataset. During training, the individual embedding features extracted by the individual feature extraction module are input into the global feature supplementation module and the Transformer encoder. The global features and encoded features obtained are fused and input into the inference module. The loss is calculated to complete the training of the second network model.
[0011] The collected individual behavior data is input into the trained second network model to obtain the group behavior recognition results.
[0012] Furthermore, the task head includes a contrast learning task head, a behavior prediction task head and a speed prediction task head;
[0013] The contrastive learning task head includes a first multi-layer perceptron, a second multi-layer perceptron, and a normalization layer connected in sequence. The encoded features are input into the contrastive learning task head to obtain contrastive learning features. The loss corresponding to the contrastive learning task head is InfoNCE loss.
[0014] The behavior prediction task head adopts a long short-term memory network, the encoding feature is input into the behavior prediction task head to obtain a behavior prediction value, and the loss corresponding to the behavior prediction task head is the mean square error loss between the behavior prediction value and the true value;
[0015] The speed transformation task head includes a mean pooling layer, a fully connected layer and a softmax function. The encoded features are input into the behavior speed transformation task head to obtain the probability distribution of the predicted transformation ratio. The loss corresponding to the speed transformation task head is the cross entropy loss between the probability distribution of the predicted transformation ratio and the actual ratio.
[0016] Furthermore, the total loss function used in training the first network model is the sum of the losses corresponding to each task head.
[0017] Furthermore, the global feature supplementation module performs the following operations:
[0018] Cluster the group feature matrix formed by splicing individual embedded features into clusters;
[0019] For each cluster, calculate the intra-cluster attention feature;
[0020] The center of each cluster is obtained by taking the mean of the vectors within the cluster;
[0021] All cluster centers are stacked into a cluster center matrix to calculate the inter-cluster attention features;
[0022] The intra-cluster attention features and inter-cluster attention features are concatenated after linear transformation, and then the dynamic weight value is obtained through the activation function;
[0023] Dynamically weight the intra-cluster attention features and inter-cluster attention features using dynamic weight values to obtain weighted attention features;
[0024] After linearly transforming the individual embedding features of each input, the value vector is obtained, which is then multiplied by the weighted attention feature to obtain the global feature corresponding to the individual embedding feature;
[0025] The global features corresponding to all the input individual embedding features are stacked to obtain the output global features.
[0026] Furthermore, the reasoning module performs the following operations:
[0027] The fused features obtained by fusing the global features and the encoding features are input into the feedforward neural network to predict the subgroup cardinality and obtain the predicted subgroup cardinality;
[0028] Based on the fusion features, the feature similarity between individuals is calculated and the feature adjacency matrix is constructed;
[0029] According to the subgroup cardinality, the feature adjacency matrix is subjected to graph clustering operation to obtain the subgroup segmentation result;
[0030] The subgroup segmentation results are input into the group behavior classifier to obtain the group behavior recognition results.
[0031] Furthermore, the second network model is trained, and the loss functions adopted include subgroup cardinality loss, subgroup segmentation loss, and group behavior recognition loss.
[0032] The present application also proposes a group behavior recognition device based on a multi-task self-supervision framework, which includes a processor and a memory storing a plurality of computer instructions. When the computer instructions are executed by the processor, the steps of the above method are implemented.
[0033] This application proposes a method and device for group behavior recognition based on a multi-task self-supervised framework. By using an individual behavior understanding module composed of a Transformer encoder and multiple task heads, it extracts rich individual behavior features from multimodal sensor data, improving the accuracy and robustness of group segmentation and behavior recognition under self-supervision. In addition, this application significantly reduces computational complexity by grouping input data and calculating attention separately within and between clusters. The global feature supplementation module adopted not only reduces the model's computational complexity, but also captures global information between individuals and has high interpretability. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is a flowchart of the group behavior recognition method based on the multi-task self-supervision framework of this application.
[0035] Figure 2 This is a schematic diagram of the first network model structure of an embodiment of the present application.
[0036] Figure 3 This is a schematic diagram of the second network model structure of an embodiment of the present application.
[0037] Figure 4 This is a schematic diagram of the global feature supplement module structure of this application. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0039] One embodiment of the present application, such as Figure 1 As shown in the figure, a group behavior recognition method based on a multi-task self-supervised framework is proposed, including:
[0040] Step S1: Obtain individual behavior data to generate a first training data set and a second training data set.
[0041] For each participating individual, collect individual behavioral data. This data can be collected using a wrist sensor or a smart device carried by the individual. This data can include accelerometer and gyroscope data, or other sensor data.
[0042] We segmented the collected continuous time series data using a sliding window technique. Considering the continuity and periodicity of human behavior, we set the window length to 4 seconds and the overlap ratio to 50%. This ensures that each sample contains sufficient behavioral information while also ensuring continuity between samples. This segmentation method allows us to extract representative individual behavior segments from the original data stream.
[0043] However, in the framework of self-supervised learning, simply relying on sliding window technology to expand data is far from enough. In order to further improve the performance of the model, data enhancement must be performed on each sample.
[0044] To fully tap the potential of this data, the following common data augmentation strategies for time series data are introduced to further expand and enrich the training dataset.
[0045] Jitter: Adding random noise to sensor data simulates sensor errors in real scenarios. This method can effectively enhance the model's robustness to noise. The formula is as follows:
[0046]
[0047] in, is the original data point, It is subject to the mean of 0 and the variance of Normally distributed random noise.
[0048] Scaling: Scale the sensor data to simulate behavioral differences between individuals. Linearly scale the data within each time window using the formula:
[0049]
[0050] in, is the scaling factor, and are the upper and lower bounds of the scaling factor.
[0051] Amplitude distortion: Enhances the data by introducing nonlinear amplitude changes in the sensor data to simulate the intensity fluctuation of the sensor signal. The formula is:
[0052]
[0053] in, is the amplitude distortion function, is the distortion amplitude, is the frequency, It's the phase.
[0054] Time warping: Enhance the data by performing linear or nonlinear transformation on the time axis of sensor data to simulate the change in time scale. Linear change is to fast forward or slow down the original data. is the original time series, Indicates time, Indicates time The sensor observation value at . Its formula is:
[0055]
[0056] Among them, when When , the time series is accelerated, that is, the time axis is compressed and the speed of the sequence changes is accelerated. When , the time series is slowed down, that is, the time axis is stretched and the speed of the sequence changes is slowed down. The time series remains unchanged. The linear change of time is used to complete the speed conversion task in the upstream task, and the speed multiplier is selected as: .
[0057] The formula for nonlinear transformation is:
[0058]
[0059] in, is the time warp function, is the distortion amplitude, is the frequency, It's the phase.
[0060] When performing data augmentation in practice, 1 to 4 data augmentation methods are usually applied to each individual's data at the same time to achieve the purpose of data expansion.
[0061] Then, the preprocessed data set is divided, for example, 80% of it is used as the first training data set and the remaining 20% is used as the second training data set.
[0062] Step S2: construct a first network model including an individual feature extraction module, a Transformer encoder and a task head, and use the first training data set to train the first network model. During training, the individual embedding features extracted by the individual feature extraction module are input into the Transformer encoder to extract the encoding features, and then input into each task head, and the loss corresponding to each task head is calculated to complete the training of the first network model.
[0063] like Figure 2 As shown, this embodiment constructs a first network model including an individual feature extraction module, a Transformer encoder and a task head, which is used to train the individual feature extraction module and the Transformer encoder.
[0064] The first phase of training is performed by inputting the first training data set into the individual feature extraction module to obtain individual embedded features. The individual feature extraction module can use a feature extraction module of the existing technology to extract features, such as CNN, LSTM, etc.
[0065] After extracting the individual embedded features, they are input into the Transformer encoder to comprehensively and three-dimensionally extract the deep features of individual activities.
[0066] Traditional self-supervised upstream tasks only use a single upstream task, which can only capture specific patterns in the data related to the task, and may result in incomplete learned feature representations. This limitation directly affects the generalization ability of the model, resulting in poor performance on downstream tasks. This application presents a multi-task individual behavior understanding module, which consists of a Transformer encoder and a task head. The individual behavior understanding module is the core component of the entire self-supervised framework, designed to extract rich individual behavior features from multimodal sensor data.
[0067] Among them, the Transformer encoder is the basic architecture of this module, which is used to capture the temporal dependencies and cross-modal interaction characteristics in sensor data. Given the input sensor data ,in is the number of individuals, is the feature dimension. The Transformer encoder encodes the input data through a multi-head self-attention mechanism and a feedforward neural network.
[0068] The calculation process of the multi-head self-attention mechanism is as follows:
[0069]
[0070] in 、 and Represent the query (Query), key (Key) and value (Value) matrices respectively, is the dimension of the key vector. Multi-head attention enhances the expressive power of the model by projecting the input into multiple subspaces and calculating the attention weights in parallel. The formula is as follows:
[0071]
[0072] in , and are all learnable projection matrices, is the number of attention heads. In this embodiment, the number of attention heads is selected as 2. After the multi-head attention, the feedforward neural network performs a nonlinear transformation on the features of each time step:
[0073]
[0074] in 、 and 、 are learnable weights and biases. Finally, the Transformer encoder gradually extracts high-level behavioral feature representations by stacking two layers of the above modules to obtain encoded features.
[0075] After the Transformer encoder, this embodiment designs three task heads, one for contrastive learning, one for behavior prediction, and one for speed prediction. These task heads take the output of the Transformer encoder as input and share the output features of the Transformer encoder, but guide the model to learn diverse behavioral features through different objective functions. Multi-task learning can significantly improve the performance, generalization ability, and robustness of the model through shared feature representation and synergy between tasks. The following is a detailed description of the three task heads selected for use in this embodiment.
[0076] The contrastive learning task aims to learn discriminative feature representations by maximizing the similarity between positive pairs and minimizing the similarity between negative pairs. Contrastive learning can be used to narrow the feature distances between activities of the same category while simultaneously amplifying the distances between different categories, making the model more discriminative. Positive pairs are constructed by augmenting the input real data. Data of the same category can form positive pairs. Data from different categories, whether before or after augmentation, can form negative pairs. All real data must form at least one positive / negative pair, and augmented data must form at least one positive pair with its corresponding real data. Positive / negative pairs are randomly selected between augmented data.
[0077] The contrastive learning task head includes the first multi-layer perceptron, the second multi-layer perceptron and the normalization layer connected in sequence, which converts the encoded features output by the Transformer encoder into contrastive learning features. , used to calculate the objective function. The objective function is based on InfoNCE loss, and its formula is as follows:
[0078]
[0079] in, is the contrastive learning feature of the positive sample pair, is the contrastive learning feature of negative samples, is the cosine similarity function, is the temperature parameter. exp Functions are based on natural constantse An exponential function with base . is the number of samples in the current batch, For the current batch The total number of negative samples.
[0080] The behavior prediction task aims to predict sensor data within a certain time step in the future based on current sensor data. This task captures the temporal dynamics and physical laws in sensor data by modeling historical data, thereby predicting the individual's future behavior.
[0081] The core goal of the behavior prediction task is to Predicting the future Sensor data within a time step ,in is the historical time step, is the feature dimension of the sensor data. Given the output features of the Transformer encoder, the behavior prediction task head generates predictions for future time steps through a decoder module (long short-term memory network LSTM):
[0082]
[0083] in, is the decoder function, is the predicted future sensor data. In order to optimize the behavior prediction task, the mean square error (MSE) loss function is used, which is defined as follows:
[0084]
[0085] in, is the future time step No. The true value of each sensor channel, is the future time step predicted by the model No. The MSE loss function minimizes the difference between the predicted value and the true value, guiding the model to learn accurate time series dynamics and physical laws.
[0086] The speed transformation task aims to transform the time scale (speed up or slow down) of input time series sensor data and allow the model to select the correct transformation ratio from a predefined set of ratios. This task captures dynamic patterns in sensor data by modeling the time scale variation of time series data, thereby accurately classifying the speed transformation ratio.
[0087] The core goal of the speed conversion task is to transform the time scale of the input time series sensor data (accelerate or slow down) and let the model The speed conversion task extracts feature representations through the Transformer encoder and uses a classification module (including a mean pooling layer, a fully connected layer, and a softmax function) to predict the probability distribution of the conversion ratio. The formula is:
[0088]
[0089] in and are learnable weights and biases, is the size of the predefined magnification set (in this task ), is mean pooling, Represents the first Transformation ratio, The encoded features output by the Transformer encoder.
[0090] In order to optimize the speed transformation task, the cross entropy loss function is adopted, which is defined as follows:
[0091]
[0092] in is the one-hot encoding of the true ratio, The model predicts The cross entropy loss function guides the model to learn accurate time scale transformation characteristics by minimizing the difference between the true ratio and the predicted probability distribution.
[0093] pass 、 and Loss functions for each of the three tasks are calculated separately and summed during training to form the overall loss function. Through the initial phase of training on these three tasks, the model can mine information from different perspectives within the data. By simultaneously optimizing the three related tasks, the model's performance and generalization capabilities are significantly improved. This multi-faceted learning generates more comprehensive and richer feature representations, thereby improving the model's performance in downstream tasks.
[0094] Step S3: construct a second network model including an individual feature extraction module, a Transformer encoder, a global feature supplement module and an inference module, freeze the parameters of the Transformer encoder, and use the second training data set to train the second network model. During training, the individual embedding features extracted by the individual feature extraction module are input into the global feature supplement module and the Transformer encoder, and the global features and encoded features obtained are fused and input into the inference module, the loss is calculated, and the training of the second network model is completed.
[0095] The second network model of this embodiment is as follows Figure 3 As shown, it includes individual feature extraction module, Transformer encoder, global feature supplement module and reasoning module.
[0096] In the second phase of training, the remaining 20% of the training set data was used to train the second network model. During training, the Transformer encoder parameters in the individual behavior understanding module, obtained from the first phase of training, were frozen to ensure that the model still understood individual activities during feature extraction. The remaining labeled data was used to train other modules, such as the subgroup cardinality prediction FFN and the group behavior classifier, enabling the model to complete the two downstream tasks of group segmentation and group behavior recognition.
[0097] First, the labeled input data is passed through the individual feature extraction module to generate corresponding individual behavior embeddings. This process is similar to the first stage and aims to generate preliminary individual features. These embeddings are then input into the Transformer encoder with frozen parameters. This encoder has already fully captured the deep semantic information of individual behaviors through multi-task learning in the first stage. By freezing its parameters, further adjustments to the encoder are avoided in the second stage, thereby preserving its general feature representation capabilities learned on the augmented data. Simultaneously, the outputs of the individual feature extraction module are spliced into a group feature matrix, which is input into the subsequent lightweight global feature supplementation module.
[0098] The group feature matrix is input into the global feature supplementation module to model the interactions between individuals and fuse them with the encoded features output by the Transformer encoder. Traditional global feature extraction methods face high computational complexity when processing large-scale data. To address this issue, this application significantly reduces computational complexity by grouping the input data and calculating attention separately within and between clusters.
[0099] like Figure 4 As shown, the input of the global feature supplementation module is the group feature matrix, which is the set of all individual embedded features, including vectors, each with a dimension of , which can be expressed as a matrix The global feature supplementation module performs the following operations:
[0100] Step 3.1: Cluster the group feature matrix formed by concatenating individual embedded features into clusters.
[0101] In order to reduce the computational complexity, we first use the K-means algorithm to transform the input Divide clusters. Each cluster Include vectors, recorded as the individual feature matrix within the cluster .
[0102] Step 3.2: For each cluster, calculate the intra-cluster attention feature.
[0103] For each cluster , calculate its cluster attention. First, Perform linear transformation to generate query ,key Sum :
[0104]
[0105] in is a learnable weight matrix. Figure 3 As shown in the intra-cluster attention module in , the intra-cluster attention score is calculated based on the individual feature matrix within the cluster:
[0106]
[0107] Then, a softmax operation is performed on the attention scores to obtain the attention features within the cluster:
[0108]
[0109] Step 3.3: Get the center of each cluster by taking the mean of the vectors within the cluster.
[0110] This step obtains the center of each cluster by taking the mean of the vectors within the cluster:
[0111]
[0112] in It is a cluster The vectors, It is a cluster The size of is the center selection strategy, here it is the mean.
[0113] Step 3.4: Stack all cluster centers into a cluster center matrix and calculate the inter-cluster attention features.
[0114] Stack all cluster centers into a cluster center matrix .right Perform linear transformation to generate query ,key Sum :
[0115]
[0116] in, is a learnable weight matrix.
[0117] The attention mechanism is used to calculate the attention score between clusters:
[0118]
[0119] Then perform a softmax operation on the attention score to obtain the inter-cluster attention feature:
[0120]
[0121] Step 3.5: Concatenate the intra-cluster attention features and inter-cluster attention features after linear transformation, and then obtain the dynamic weight value through the activation function.
[0122] In order to dynamically adjust the contribution of intra-cluster and inter-cluster attention, a dynamic fusion weight mechanism is introduced here. , the attention features within the cluster and inter-cluster attention features After linear transformation, they are spliced together. Dynamic weight values are generated through an activation function sigmoid. :
[0123]
[0124] in 、 and is a learnable weight matrix, is the splicing symbol, which splices the inter-cluster attention features and the intra-cluster attention features. is the sigmoid activation function.
[0125] Step 3.6: Use the dynamic weight value to dynamically weight the intra-cluster attention feature and the inter-cluster attention feature to obtain the weighted attention feature.
[0126] The dynamic weighting formula is as follows:
[0127]
[0128] Step 3.7: After linearly transforming the individual embedding features of each input, the value vector is obtained, and then multiplied by the weighted attention feature to obtain the global feature corresponding to the individual embedding feature.
[0129] For each input , after linear transformation, we get the value vector , and then multiplied by the weighted attention feature to obtain the global feature corresponding to the individual embedding feature:
[0130]
[0131] Step 3.8: Stack the global features corresponding to all the input individual embedding features to obtain the output global features.
[0132] All output Stacked into the final output matrix , thereby completing the global feature extraction.
[0133] Compared with the traditional self-attention mechanism, the group attention mechanism reduces the computational complexity from Reduce to ,in is the average size of each cluster. Take the closest When the integer is , the computational complexity reaches the theoretical optimum, which is The global feature supplementation module not only reduces the computational complexity of the model, but also captures global information between individuals and has high interpretability. Subsequently, the output of the global feature supplementation module is fused with the output of the Transformer encoder and input into the inference module.
[0134] The output of the global feature supplementation module is spliced with the output of the Transformer encoder, and the spliced fusion features are input into the inference module to complete the group segmentation and group behavior recognition tasks.
[0135] The reasoning module performs the following operations:
[0136] Step 4.1: Input the fused features obtained by fusing the global features and the encoding features into the feedforward neural network to predict the subgroup cardinality and obtain the predicted subgroup cardinality.
[0137] In this step, the fusion features are used to predict the subgroup cardinality through a feedforward neural network to obtain the predicted subgroup cardinality.
[0138] Step 4.2: Based on the fused features, calculate the feature similarity between individuals and construct the feature adjacency matrix.
[0139] This step builds a feature adjacency matrix based on the fused features, calculating the feature similarity between individuals. Each individual's features are the concatenation of the coded features and the corresponding individual features from the global features. The specific implementation of constructing the feature adjacency matrix is a relatively mature technology in this field and will not be detailed here.
[0140] Step 4.3: Based on the subgroup cardinality, perform graph clustering on the feature adjacency matrix to obtain the subgroup segmentation results.
[0141] Then, the feature adjacency matrix is subjected to graph clustering, and the result of graph clustering is the subgroup segmentation result. The specific implementation of graph clustering is also a relatively mature technical means in this field, which will not be described in detail here.
[0142] Step 4.4: Input the subgroup segmentation results into the group behavior classifier to obtain the group behavior recognition results.
[0143] In this step, the individual segmentation results will be further sent to the classifier to realize the classification of group behavior.
[0144] In a specific embodiment, when training the second network model, the inference module adopts the following loss function to ensure that the model can accurately capture the complex structure and dynamic changes of group behavior.
[0145] The subgroup cardinality loss function is defined as follows:
[0146]
[0147] in, is the loss function of the subgroup cardinality, using mean square error, It is The predicted subgroup cardinality of samples, It is The true subgroup cardinality of samples.
[0148] The subgroup segmentation loss function is as follows:
[0149]
[0150] is the subgroup segmentation loss, using binary cross entropy loss, is the predicted subgroup membership, ranging from , is the actual subgroup membership relationship, with a value of 0 or 1. is the total number of samples.
[0151] The loss function for group behavior recognition can be defined as follows:
[0152]
[0153] in, It is The sample in The predicted probability of the class, It is The sample in The true distribution of classes, is the number of behavior categories, is the total number of samples.
[0154] The total loss function It can be expressed as the following formula:
[0155]
[0156] in, Represents the weight coefficient, which can control the contribution of the three losses to model training, thereby optimizing the overall performance of the model.
[0157] Step S4: input the collected individual behavior data into the trained second network model to obtain group behavior recognition results.
[0158] After the second network model is trained, for the group to be identified, the behavioral data of each individual is obtained through sensors and input into the trained second network model to obtain accurate prediction results.
[0159] In this application, the technical solution of this application is also verified through experiments. The experiments are based on the UT-Group dataset reconstructed from UT-Data and the self-made WBSensor dataset, and are compared with the current mainstream subgrouping algorithms and group behavior recognition algorithms. The results of the subgrouping task are measured using F1 score, precision, recall rate, mean average precision (mAP), and IOU-AOC indicators, while the group behavior recognition task is measured using F1 score, precision, recall rate, and accuracy indicators. The experimental results are shown in Tables 1 and 2:
[0160] Table 1
[0161]
[0162] Table 2
[0163]
[0164] The experimental results of this proposed method are divided into two parts, evaluating the model's performance in two tasks: group segmentation and group behavior recognition. In the group segmentation task, this proposed method achieved an F1 score of 42.40%, a precision of 39.54%, a recall of 45.70%, a mean average precision of 61.75%, and an IOU-AOC of 22.00% on the UT-Group dataset. On the WBSensor dataset, this method achieved an F1 score of 46.91%, a precision of 42.42%, a recall of 52.45%, a mean average precision of 65.43%, and an IOU-AOC of 25.81%. This method outperformed other methods in all metrics. In the group behavior recognition task, our proposed method achieved an F1 score of 69.51%, a precision of 71.48%, a recall of 67.65%, and an accuracy of 68.79% on the UT-Group dataset. On the WBSensor dataset, our proposed method achieved an F1 score of 75.84%, a precision of 73.45%, a recall of 78.39%, and an accuracy of 72.80%. Our proposed method also outperformed other methods in all metrics for this task. In summary, our proposed method has certain advantages over other algorithms.
[0165] In another embodiment, the present application also proposes a group behavior recognition device based on a multi-task self-supervision framework, including a processor and a memory storing a plurality of computer instructions, wherein the computer instructions implement the steps of the above method when executed by the processor.
[0166] Regarding the specific limitations of the group behavior recognition device based on the multi-task self-supervision framework, please refer to the limitations of the group behavior recognition method based on the multi-task self-supervision framework above, which will not be repeated here. The above-mentioned group behavior recognition device based on the multi-task self-supervision framework can be implemented in whole or in part by software, hardware, and a combination thereof. It can be embedded in or independent of the processor in the computer device in the form of hardware, or it can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the above corresponding operations.
[0167] The memory and processor are electrically connected, directly or indirectly, to enable data transmission or interaction. For example, these elements may be electrically connected to each other via one or more communication buses or signal lines. The memory stores a computer program executable on the processor, and the processor executes the computer program stored in the memory to implement the group behavior recognition method based on the multi-task self-supervisory framework in the embodiment of the present invention.
[0168] The memory may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The memory is used to store a program, and the processor executes the program after receiving an execution instruction.
[0169] The processor may be an integrated circuit chip with data processing capabilities. The processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor.
[0170] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A group behavior recognition method based on a multi-task self-supervised framework, characterized by: The group behavior recognition method based on the multi-task self-supervision framework includes: Obtaining individual behavioral data to generate a first training data set and a second training data set; Constructing a first network model including an individual feature extraction module, a Transformer encoder, and a task head, and training the first network model using a first training dataset. During training, the individual embedding features extracted by the individual feature extraction module are input into the Transformer encoder to extract the encoded features, which are then input into each task head. The loss corresponding to each task head is calculated to complete the training of the first network model. Constructing a second network model including an individual feature extraction module, a Transformer encoder, a global feature supplementation module, and an inference module, freezing the parameters of the Transformer encoder, and training the second network model using the second training dataset. During training, the individual embedding features extracted by the individual feature extraction module are input into the global feature supplementation module and the Transformer encoder. The global features and encoded features obtained are fused and input into the inference module. The loss is calculated to complete the training of the second network model. Input the collected individual behavior data into the trained second network model to obtain the group behavior recognition results; The global feature supplementation module performs the following operations: Cluster the group feature matrix formed by splicing individual embedded features into clusters; For each cluster, calculate the intra-cluster attention feature; The center of each cluster is obtained by taking the mean of the vectors within the cluster; All cluster centers are stacked into a cluster center matrix to calculate the inter-cluster attention features; The intra-cluster attention features and inter-cluster attention features are concatenated after linear transformation, and then the dynamic weight value is obtained through the activation function; Dynamically weight the intra-cluster attention features and inter-cluster attention features using dynamic weight values to obtain weighted attention features; After linearly transforming the individual embedding features of each input, the value vector is obtained, which is then multiplied by the weighted attention feature to obtain the global feature corresponding to the individual embedding feature; The global features corresponding to all the input individual embedding features are stacked to obtain the output global features.
2. The group behavior recognition method based on a multi-task self-supervised framework according to claim 1 is characterized in that: The task heads include a contrast learning task head, a behavior prediction task head and a speed conversion task head; The contrastive learning task head includes a first multi-layer perceptron, a second multi-layer perceptron, and a normalization layer connected in sequence. The encoded features are input into the contrastive learning task head to obtain contrastive learning features. The loss corresponding to the contrastive learning task head is InfoNCE loss. The behavior prediction task head adopts a long short-term memory network, the encoding feature is input into the behavior prediction task head to obtain a behavior prediction value, and the loss corresponding to the behavior prediction task head is the mean square error loss between the behavior prediction value and the true value; The speed transformation task head includes a mean pooling layer, a fully connected layer and a softmax function. The encoded features are input into the behavior speed transformation task head to obtain the probability distribution of the predicted transformation ratio. The loss corresponding to the speed transformation task head is the cross entropy loss between the probability distribution of the predicted transformation ratio and the actual ratio.
3. The group behavior recognition method based on a multi-task self-supervised framework according to claim 2 is characterized in that: The total loss function used in training the first network model is the sum of the losses corresponding to each task head.
4. The group behavior recognition method based on a multi-task self-supervised framework according to claim 1 is characterized in that: The reasoning module performs the following operations: The fused features obtained by fusing the global features and the encoding features are input into the feedforward neural network to predict the subgroup cardinality and obtain the predicted subgroup cardinality; Based on the fusion features, the feature similarity between individuals is calculated and the feature adjacency matrix is constructed; According to the subgroup cardinality, the feature adjacency matrix is subjected to graph clustering operation to obtain the subgroup segmentation result; The subgroup segmentation results are input into the group behavior classifier to obtain the group behavior recognition results.
5. The method for group behavior recognition based on a multi-task self-supervised framework according to claim 4 is characterized in that: The second network model is trained using the following loss functions: subgroup cardinality loss, subgroup segmentation loss, and group behavior recognition loss.
6. A group behavior recognition device based on a multi-task self-supervisory framework, comprising a processor and a memory storing a plurality of computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Group behavior identification method based on multi-modal fusion and implicit interactive relationship learning
CN115719510A
Self-supervised human behavior recognition method based on time-frequency contrast learning
CN117113164A
Group behavior recognition method based on multi-scale feature extraction
CN120108034A