Group behavior identification method and device based on multi-task self-supervision framework

Through the individual behavior understanding module and global feature supplementary module of the multi-task self-supervision framework, the problems of sub-group division and noise filtering in large-scale groups are solved, the accuracy and robustness of group behavior analysis are improved, and the computational complexity is reduced.

CN120354157AActive Publication Date: 2025-07-22ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202510858515.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-22
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively and adaptively divide subgroups in large-scale groups and filter noise individuals, resulting in insufficient accuracy and robustness of group behavior analysis in complex scenarios.

Method used

The multi-task self-supervision framework is adopted, and an individual behavior understanding module is formed through a Transformer encoder and multiple task heads. Individual behavior characteristics are extracted from the multimodal sensor data, and attention is calculated within and between clusters through the global feature supplement module, reducing the computational complexity and improving the accuracy and robustness of group segmentation and behavior recognition.

Benefits of technology

It significantly improves the accuracy and robustness of group segmentation and behavior recognition under self-supervised conditions, reduces the computational complexity, can capture global information between individuals and generate more comprehensive and richer feature representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354157A_ABST
    Figure CN120354157A_ABST
Patent Text Reader

Abstract

The invention discloses a group behavior recognition method and device based on a multi-task self-supervision framework. The method comprises the following steps: firstly, constructing a first network model comprising an individual feature extraction module, a Transform encoder and task heads, training the first network model by adopting a first training data set, calculating loss corresponding to each task head, and completing training of the first network model; and then constructing a second network model comprising an individual feature extraction module, a Transform encoder, a global feature supplement module and a reasoning module, freezing parameters of the Transform encoder, and training the second network model by adopting a second training data set. And inputting the collected individual behavior data into the trained second network model to obtain a group behavior recognition result. According to the method, the calculation complexity is remarkably reduced, and the accuracy and robustness of group segmentation and behavior recognition under the self-supervision condition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of behavior recognition, and particularly relates to a method and device for group behavior recognition based on a multi-task self-supervised framework. Background Art

[0002] Group behavior analysis based on sensor data has become a research hotspot in the current cross-field of artificial intelligence and the Internet of Things. This technology reveals the dynamic characteristics of group behavior by capturing the interaction patterns among multiple individuals, and has important application values in fields such as smart cities, intelligent transportation, and social computing. However, group behavior analysis in actual scenarios still faces significant challenges, mainly because group behavior is not a simple superposition of individual behaviors, but requires integrating individual behavior data and complex interaction relationships, and a modeling framework from local to global needs to be established.

[0003] In recent years, with the popularization of Internet of Things technology, especially the wide application of wireless sensor networks and wearable devices, massive sensor data provides a rich information source for group behavior research. At the same time, the progress of deep learning technology has further promoted the development of this field. Compared with traditional vision-based methods, the sensor data solution has advantages such as low cost, wide applicable scenarios, and strong privacy protection. In addition, modern intelligent terminals are generally integrated with multi-modal sensors, providing reliable technical support for the real-time monitoring and analysis of group behavior, and greatly improving the practicality and scalability of this research.

[0004] Although group behavior analysis based on sensor data shows broad prospects, this field still faces many challenges. First, sensor data usually has a large volume, and many data lack clear labels, making it difficult to directly apply traditional supervised learning methods. Therefore, the self-supervised learning framework has gradually become a research hotspot. It learns general representations from unlabeled data by designing pre-training tasks, and then fine-tunes in downstream tasks to complete the given tasks. This approach effectively alleviates the problem of insufficient labeled data. In addition, in large-scale group scenarios, the group structure is often more complex. This is because large-scale groups usually contain multiple dynamically changing subgroups, and there are also irrelevant interfering individuals at the same time. How to adaptively divide sub-groups and effectively filter out noise individuals to improve the model's understanding ability of complex scenarios is still one of the current research difficulties. Summary of the Invention

[0005] The purpose of this application is to provide a method for group behavior recognition based on a multi-task self-supervised framework to improve the accuracy and robustness of group segmentation and behavior recognition under self-supervised conditions.

[0006] To achieve the above purpose, the technical solution of this application is as follows: A method for group behavior recognition based on a multi-task self-supervised framework, including: Obtain the behavioral data of an individual, and generate a first training dataset and a second training dataset; Construct a first network model including an individual feature extraction module, a Transformer encoder, and task heads. Use the first training dataset to train the first network model. During training, the individual embedding features extracted by the individual feature extraction module are input into the Transformer encoder to extract encoded features, and then input into each task head to calculate the losses corresponding to each task head, completing the training of the first network model; Construct a second network model including an individual feature extraction module, a Transformer encoder, a global feature supplement module, and an inference module. Freeze the parameters of the Transformer encoder. Use the second training dataset to train the second network model. During training, the individual embedding features extracted by the individual feature extraction module are input into the global feature supplement module and the Transformer encoder. The global features and encoded features obtained respectively are fused and then input into the inference module to calculate the loss, completing the training of the second network model; Input the collected individual behavioral data into the trained second network model to obtain the group behavior recognition result.

[0007] Further, the task heads include a contrastive learning task head, a behavior prediction task head, and a speed prediction task head; The contrastive learning task head includes a first multi-layer perceptron, a second multi-layer perceptron, and a normalization layer connected in sequence. The encoded features are input into the contrastive learning task head to obtain contrastive learning features. The loss corresponding to the contrastive learning task head is the InfoNCE loss; The behavior prediction task head uses a long short-term memory network. The encoded features are input into the behavior prediction task head to obtain behavior prediction values. The loss corresponding to the behavior prediction task head is the mean square error loss between the behavior prediction values and the true values; The speed transformation task head includes a mean pooling layer, a fully connected layer, and a softmax function. The encoded features are input into the behavior speed transformation task head to obtain the probability distribution of the predicted transformation ratio. The loss corresponding to the speed transformation task head is the cross-entropy loss between the probability distribution of the predicted transformation ratio and the true ratio.

[0008] Further, when training the first network model, the total loss function used is the sum of the losses corresponding to each task head.

[0009] Further, the global feature supplement module performs the following operations: Cluster the group feature matrix formed by splicing individual embedding features into each cluster; For each cluster, calculate the in-cluster attention features; By taking the mean of the vectors within the cluster, the center of each cluster is obtained; Stack all the cluster centers into a cluster center matrix and calculate the inter-cluster attention features; Concatenate the intra-cluster attention features and the inter-cluster attention features after linear transformation, and then obtain the dynamic weight values through an activation function; Use the dynamic weight values to dynamically weight the intra-cluster attention features and the inter-cluster attention features to obtain the weighted attention features; Perform a linear transformation on each input individual embedding feature to obtain a value vector, and then multiply it by the weighted attention features to obtain the global feature corresponding to the individual embedding feature; Stack the global features corresponding to all the input individual embedding features to obtain the output global features.

[0010] Furthermore, the inference module performs the following operations: Input the fusion features obtained by fusing the global features and the encoded features into a feed-forward neural network to predict the subgroup cardinality, and obtain the predicted subgroup cardinality; Based on the fusion features, calculate the feature similarity between individuals and construct a feature adjacency matrix; According to the subgroup cardinality, perform spectral clustering operations on the feature adjacency matrix to obtain the subgroup segmentation result; Input the subgroup segmentation result into the group behavior classifier to obtain the group behavior recognition result.

[0011] Furthermore, when training the second network model, the loss function used includes subgroup cardinality loss, subgroup segmentation loss, and group behavior recognition loss.

[0012] This application also proposes a group behavior recognition device based on a multi-task self-supervised framework, including a processor and a memory storing a number of computer instructions. When the computer instructions are executed by the processor, the steps of the above method are implemented.

[0013] A group behavior recognition method and device based on a multi-task self-supervised framework proposed in this application extracts rich individual behavior features from multi-modal sensor data through an individual behavior understanding module composed of a Transformer encoder and multiple task heads, improving the accuracy and robustness of group segmentation and behavior recognition under self-supervised conditions. In addition, this application significantly reduces the computational complexity by grouping the input data and calculating attention within and between clusters respectively. The global feature supplement module used can not only reduce the model computational complexity but also capture the global information between individuals, and has high interpretability. Description of the Drawings

[0014] Figure 1Flowchart of the group behavior recognition method based on the multi-task self-supervised framework of the present application.

[0015] Figure 2 Schematic diagram of the first network model structure of the embodiment of the present application.

[0016] Figure 3 Schematic diagram of the second network model structure of the embodiment of the present application.

[0017] Figure 4 Schematic diagram of the global feature supplement module structure of the present application. Detailed implementation manners

[0018] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0019] An embodiment of the present application, as Figure 1 shown, proposes a group behavior recognition method based on a multi-task self-supervised framework, including: Step S1, obtaining the behavior data of individuals, and generating a first training data set and a second training data set.

[0020] For each participating individual, collect the behavior data of the individual. The behavior data of the individual can be collected through the wrist sensor of each individual, or through the smart terminal carried by the individual. The behavior data of the individual can include accelerometer and gyroscope data, or other sensor data.

[0021] Adopt the sliding window technology to segment the collected continuous time series data. Considering the continuity and periodicity characteristics of human behaviors, set the window length to 4 seconds and the overlap rate to 50%. In this way, it can not only ensure that each sample contains sufficient behavior information, but also ensure the continuity between samples. Through this segmentation method, representative individual behavior segments can be extracted from the original data stream.

[0022] However, in the framework of self-supervised learning, simply relying on the sliding window technology to expand the data is far from enough. In order to further improve the performance of the model, data augmentation also needs to be performed on each sample.

[0023] In order to fully exploit the potential of these data, the following common data augmentation strategies for time series data are introduced to further expand and enrich the training data set.

[0024] Jitter: Simulate the sensor error in the real scenario by adding random noise to the sensor data. This method can effectively enhance the robustness of the model to noise. Its formula is shown as follows:

[0025] Among them, is the original data point, is the random noise that follows a normal distribution with a mean of 0 and a variance of .

[0026] Scaling: Scale the sensor data to simulate the behavioral differences between different individuals. Linearly scale the data within each time window, and the formula is:

[0027] Among them, is the scaling factor, and are the upper and lower limits of the scaling factor.

[0028] Amplitude distortion: Enhance the data by introducing non-linear amplitude changes in the sensor data to simulate the intensity fluctuations of the sensor signal. The formula is:

[0029] Among them, is the amplitude distortion function, is the distortion amplitude, is the frequency, is the phase.

[0030] Time distortion: Enhance the data by performing linear or non-linear transformations on the time axis of the sensor data to simulate changes on the time scale. The linear change is to fast forward or slow down the original data. Assume is the original time series, represents time, represents the sensor observation value at time . The formula is:

[0031] Among them, when , the time series is accelerated, that is, the time axis is compressed and the change speed of the sequence is increased. When , the time series is slowed down, that is, the time axis is stretched and the change speed of the sequence is decreased. When , the time series remains unchanged. The linear change of time is used to complete the speed transformation task in the upstream task, and the speed magnification is selected as: .

[0032] The formula for non-linear transformation is:

[0033] Among them, is the time distortion function, is the distortion amplitude, is the frequency, is the phase.

[0034] When actually performing data augmentation, usually one to four data augmentation methods are simultaneously applied to the data of each individual, so as to achieve the purpose of data expansion.

[0035] Then, the preprocessed data set is divided. For example, 80% of it is used as the first training data set, and the other 20% is used as the second training data set.

[0036] Step S2: Construct a first network model including an individual feature extraction module, a Transformer encoder, and a task head. Use the first training data set to train the first network model. During training, the individual embedding features extracted by the individual feature extraction module are input into the Transformer encoder to extract encoded features, and then input into each task head to calculate the losses corresponding to each task head, and complete the training of the first network model.

[0037] As Figure 2 shown, in this embodiment, a first network model including an individual feature extraction module, a Transformer encoder, and a task head is constructed for training the individual feature extraction module and the Transformer encoder.

[0038] Execute the first-stage training, input the first training data set into the individual feature extraction module to obtain individual embedding features. The individual feature extraction module can use the feature extraction modules of existing technologies for feature extraction, such as CNN, LSTM, etc.

[0039] After the individual embedding features are extracted, input them into the Transformer encoder to comprehensively and three-dimensionally extract the deep features of individual activities.

[0040] Traditional self-supervised upstream tasks only use a single upstream task. A single task can only capture specific patterns related to the task in the data, which may lead to insufficiently comprehensive learned feature representations. This limitation will directly affect the generalization ability of the model, resulting in poor performance on downstream tasks. This application presents a multi-task individual behavior understanding module, which consists of a Transformer encoder and a task head. The individual behavior understanding module is the core component of the entire self-supervised framework, aiming to extract rich individual behavior features from multi-modal sensor data.

[0041] Among them, the Transformer encoder is the basic architecture of this module, which is used to capture the temporal dependence relationship and cross-modal interaction characteristics in the sensor data. Given the input sensor data , where is the number of individuals, is the feature dimension. The Transformer encoder encodes the input data through the multi - head self - attention mechanism and the feed - forward neural network.

[0042] The calculation process of the multi - head self - attention mechanism is as follows:

[0043] Among them 、 and represent the query, key, and value matrices respectively, is the dimension of the key vector. The multi - head attention enhances the model's expressive power by projecting the input into multiple sub - spaces and calculating the attention weights in parallel. Its formula is as follows:

[0044] Among them , and are all learnable projection matrices, is the number of attention heads. In this embodiment, the number of attention heads is selected as 2. After the multi - head attention, the feed - forward neural network performs a non - linear transformation on the features of each time step:

[0045] Among them 、 and 、 are learnable weights and biases. Finally, the Transformer encoder stacks two layers of the above - mentioned modules to gradually extract high - level behavioral feature representations and obtain the encoded features.

[0046] After the Transformer encoder, this embodiment designs three task heads, which are respectively used for contrast learning, behavior prediction, and speed prediction. These task heads take the output of the Transformer encoder as the input, share the output features of the Transformer encoder, but guide the model to learn diverse behavioral features through different objective functions. Multi - task learning can significantly improve the model's performance, generalization ability, and robustness through shared feature representations and the synergy between tasks. The following is a detailed description of the three task heads selected and used in this embodiment.

[0047] Contrastive learning task head. The contrastive learning task aims to learn discriminative feature representations by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs. Contrastive learning can be used to reduce the feature distance of activities in the same category while magnifying the distance between different categories, making the model more discriminative. Positive sample pairs are constructed by data augmentation of the input real data, and any data within the same category can form positive sample pairs. Whether it is the data before or after augmentation, as long as they belong to different categories, they can form negative sample pairs. At least ensure that there is at least one positive / negative sample pair among all real data, that the augmented data forms at least one positive sample pair with its corresponding real data, and that the augmented data randomly forms positive / negative sample pairs with each other.

[0048] The contrastive learning task head includes a first multi-layer perceptron, a second multi-layer perceptron, and a normalization layer connected in sequence, which transforms the encoded features output by the Transformer encoder into contrastive learning features , which is used to calculate the objective function. The objective function is based on the InfoNCE loss, and its formula is as follows:

[0049] where, is the contrastive learning feature of the positive sample pair, is the contrastive learning feature of the negative sample pair, is the cosine similarity function, is the temperature parameter. exp The function is the exponential function with the natural constant e as the base. is the number of samples in the current batch, is in the current batch is the total number of negative samples.

[0050] Behavior prediction task head. The behavior prediction task aims to predict sensor data within a certain number of future time steps based on sensor data at the current moment. This task models historical data to capture the temporal dynamics and physical laws in sensor data, thereby realizing the prediction of an individual's future behavior.

[0051] The core goal of the behavior prediction task is to predict sensor data within the next time steps from sensor data at the current moment , where is the historical time step length, is the feature dimension of the sensor data. Given the output features of the Transformer encoder, the behavior prediction task head generates predicted values for future time steps through a decoder module (long short-term memory network LSTM):

[0052] Among them, is the decoder function, is the predicted future sensor data. To optimize the behavior prediction task, the mean squared error (MSE) loss function is adopted, which is defined as follows:

[0053] Among them, is the true value of the th sensor channel at the future time step, is the value of the th sensor channel predicted by the model at the future time step. The MSE loss function guides the model to learn accurate temporal dynamics and physical laws by minimizing the difference between the predicted value and the true value.

[0054] The speed transformation task head. The speed transformation task aims to perform a time-scale transformation (acceleration or slow-down) on the input temporal sensor data and let the model select the correct transformation ratio from a predefined set of ratios. This task captures the dynamic patterns in the sensor data by modeling the time-scale change characteristics of the temporal data, thereby achieving accurate classification of the speed transformation ratio.

[0055] The core objective of the speed transformation task is to perform a time-scale transformation (acceleration or slow-down) on the input temporal sensor data and let the model select the correct transformation ratio from the predefined set of ratios. The speed transformation task extracts feature representations through a Transformer encoder and uses a classification module (including a mean pooling layer, a fully connected layer, and a softmax function) to predict the probability distribution of the transformation ratio. The formula is:

[0056] Among them and are the learnable weights and biases, is the size of the predefined set of ratios (in this task ), is the mean pooling, represents the th transformation ratio in the set of ratios, is the encoded feature output by the Transformer encoder.

[0057] To optimize the speed transformation task, the cross-entropy loss function is adopted, which is defined as follows:

[0058] Among them is the one-hot encoding of the true magnification, and is the probability of the th magnification predicted by the model. The cross-entropy loss function guides the model to learn accurate time-scale transformation characteristics by minimizing the difference between the true magnification and the predicted probability distribution.

[0059] By , and respectively calculate the loss functions of the three tasks, and add the above three losses during the training process as the overall loss function. Through the first-stage training of the above three tasks, the model can mine information in the data from different perspectives through different tasks. By simultaneously optimizing the three related tasks, the performance and generalization ability of the model can be significantly improved. This multi-angle learning can generate more comprehensive and rich feature representations, thereby improving the performance of the model in downstream tasks.

[0060] Step S3: Construct a second network model including an individual feature extraction module, a Transformer encoder, a global feature supplementation module, and an inference module. Freeze the parameters of the Transformer encoder, and use the second training dataset to train the second network model. During the training, the individual embedding features extracted by the individual feature extraction module are input into the global feature supplementation module and the Transformer encoder, and the fused global features and encoded features are input into the inference module to calculate the loss, and complete the training of the second network model.

[0061] The second network model of this embodiment is as Figure 3 shown, including an individual feature extraction module, a Transformer encoder, a global feature supplementation module, and an inference module.

[0062] In the second stage of training, use the remaining 20% of the training set data as the training data to train the second network model. During the training, by freezing the parameters of the Transformer encoder in the individual behavior understanding module obtained from the first-stage training, ensure that the ability to understand individual activities is still available during the feature extraction process. Use another part of the labeled data to train other modules such as the subgroup cardinality prediction FFN and the group behavior classifier, so that the model can complete the two downstream tasks of group segmentation and group behavior recognition.

[0063] First, the labeled input data is passed through the individual feature extraction module to generate corresponding individual behavior embeddings. This process is similar to the first stage and aims to generate the initial features of the individual. Subsequently, these embeddings are input into the Transformer encoder with frozen parameters, which has fully captured the deep semantic information of individual behaviors through multi-task learning in the first stage. By freezing its parameters, further adjustment of the encoder in the second stage is avoided, thus retaining its general feature representation ability learned on the enhanced data. At the same time, the output of the individual feature extraction module is concatenated into a group feature matrix and input into the subsequent lightweight global feature supplementation module.

[0064] The group feature matrix is input into the global feature supplementation module to model the interaction relationships among individuals and fuse with the encoded features output by the Transformer encoder. Traditional global feature extraction methods face the problem of high computational complexity when dealing with large-scale data. To solve this problem, in this application, the input data is grouped and the attention is calculated within and between clusters respectively, significantly reducing the computational complexity.

[0065] As Figure 4 shown, the input of the global feature supplementation module is the group feature matrix, that is, the set of all individual embedding features, containing vectors, and the dimension of each vector is , which can be represented as a matrix . The global feature supplementation module performs the following operations: Step 3.1: Cluster the group feature matrix formed by concatenating individual embedding features into each cluster.

[0066] To reduce the computational complexity, first use the K-means algorithm to divide the input into clusters. Each cluster contains vectors, denoted as the within-cluster individual feature matrix .

[0067] Step 3.2: For each cluster, calculate the within-cluster attention feature.

[0068] For each cluster , calculate its within-cluster attention. First, perform a linear transformation on to generate the query , key and value :

[0069] where is the learnable weight matrix. AsFigure 3 As shown in the intra-cluster attention module in

[0070] Subsequently, a softmax operation is performed on the attention scores to obtain the intra-cluster attention features:

[0071] Step 3.3: By taking the mean of the intra-cluster vectors, the center of each cluster is obtained.

[0072] In this step, the center of each cluster is obtained by taking the mean of the intra-cluster vectors:

[0073] where is the th vector in the cluster is the size of the cluster, is the center selection strategy, which is taking the mean here.

[0074] Step 3.4: Stack all the cluster centers into a cluster center matrix and calculate the inter-cluster attention features.

[0075] Stack all the cluster centers into a cluster center matrix . Perform a linear transformation on to generate a query , a key and a value :

[0076] where is a learnable weight matrix.

[0077] The inter-cluster attention scores are calculated using the attention mechanism:

[0078] Subsequently, a softmax operation is performed on the attention scores to obtain the inter-cluster attention features:

[0079] Step 3.5: Concatenate the intra-cluster attention features and the inter-cluster attention features after linear transformation, and then obtain the dynamic weight values through an activation function.

[0080] To dynamically adjust the contributions of intra-cluster and inter-cluster attention, a dynamic fusion weight mechanism is introduced here. For each input , its intra - cluster attention features and inter - cluster attention features are concatenated after linear transformation. A dynamic weight value is generated through an activation function sigmoid :

[0081] where , and are learnable weight matrices, is the concatenation symbol, which concatenates the inter - cluster attention features and intra - cluster attention features, is the sigmoid activation function.

[0082] Step 3.6: Dynamically weight the intra - cluster attention features and inter - cluster attention features using the dynamic weight value to obtain weighted attention features.

[0083] The dynamic weighting formula is as follows:

[0084] Step 3.7: After linear transformation of each input individual embedding feature, a value vector is obtained, and then it is multiplied by the weighted attention features to obtain the global feature corresponding to the individual embedding feature.

[0085] For each input , after its linear transformation, a value vector is obtained, and then it is multiplied by the weighted attention features to obtain the global feature corresponding to the individual embedding feature:

[0086] Step 3.8: Stack the global features corresponding to all input individual embedding features to obtain the output global feature.

[0087] Stack all the output into the final output matrix , thus completing the extraction of global features.

[0088] Compared with the traditional self - attention mechanism, the grouped attention mechanism reduces the computational complexity from to , where is the average size of each cluster. When takes the integer closest to , the computational complexity reaches the theoretical optimum, which is . The global feature supplementation module can not only reduce the model's computational complexity, but also capture the global information among individuals and has high interpretability. Subsequently, the output of the global feature supplementation module will be fused with the output of the Transformer encoder and input into the inference module.

[0089] Concatenate the output of the global feature supplementation module with the output of the Transformer encoder, and input the concatenated fused features into the inference module to complete the group segmentation and group behavior recognition tasks.

[0090] The inference module performs the following operations: Step 4.1: Input the fused features obtained by fusing the global features and the encoded features into a feed-forward neural network to predict the subgroup cardinality, and obtain the predicted subgroup cardinality.

[0091] In this step, the fused features are used to predict the subgroup cardinality through a feed-forward neural network to obtain the predicted subgroup cardinality.

[0092] Step 4.2: Based on the fused features, calculate the feature similarity between individuals and construct a feature adjacency matrix.

[0093] In this step, based on the fused features, calculate the feature similarity between individuals to construct a feature adjacency matrix, where the feature of each individual is the concatenation of the corresponding individual features in the encoded features and the global features. The specific implementation of constructing the feature adjacency matrix is a relatively mature technical means in this field and will not be elaborated here.

[0094] Step 4.3: According to the subgroup cardinality, perform spectral clustering operations on the feature adjacency matrix to obtain the subgroup segmentation result.

[0095] Subsequently, perform spectral clustering on the feature adjacency matrix, and the result of spectral clustering is the subgroup segmentation result. The specific implementation of spectral clustering is also a relatively mature technical means in this field and will not be elaborated here.

[0096] Step 4.4: Input the subgroup segmentation result into the group behavior classifier to obtain the group behavior recognition result.

[0097] In this step, the individual segmentation result will be further sent to the classifier to achieve the classification of group behaviors.

[0098] In a specific embodiment, when training the second network model, the inference module uses the following loss function to ensure that the model can accurately capture the complex structure and dynamic changes of group behaviors.

[0099] The subgroup cardinality loss function is defined as follows:

[0100] Where is the loss function for the subgroup cardinality. The mean squared error is selected. is the predicted subgroup cardinality of the th sample. is the true subgroup cardinality of the

[0101] The subgroup segmentation loss function is as follows:

[0102] is the subgroup segmentation loss. The binary cross-entropy loss is selected. is the predicted subgroup membership, and its value range is , is the true subgroup membership, and its value is 0 or 1. is the total number of samples.

[0103] The group behavior recognition loss function can be defined as follows:

[0104] Among them, is the predicted probability of the th sample for the th class. is the true distribution of the th sample for the th class. is the number of behavior classes. is the total number of samples.

[0105] The total loss function can be expressed by the following formula:

[0106] Among them, represents the weight coefficient, which can control the contributions of the three losses to the model training, and thus optimize the overall performance of the model.

[0107] Step S4: Input the collected individual behavior data into the trained second network model to obtain the group behavior recognition result.

[0108] After training the second network model, for the group to be recognized, the behavior data of each individual is obtained through sensors and input into the trained second network model, and then an accurate prediction result can be obtained.

[0109] In this application, the technical solution of this application was also experimentally verified. The experiment was based on the UT-Group dataset reconstructed from UT-Data and the self-made WBSensor dataset, and was compared with the current mainstream subgroup division algorithms and group behavior recognition algorithms. The results of the subgroup division task were measured using F1 score, precision, recall, mean average precision mAP, and IOU-AOC metrics, and the group behavior recognition task was measured using F1 score, precision, recall, and accuracy metrics. The experimental results are shown in Tables 1 and 2 as follows: Table 1

[0110] Table 2

[0111] The experimental results of the method of this application are divided into two parts, which respectively evaluate the performance of the model in two tasks: group segmentation and group behavior recognition. In the subgroup division subtask, the F1 value of the method of this application on UT-Group is 42.40%, the precision is 39.54%, the recall is 45.70%, the mean average precision is 61.75%, and the IOU-AOC is 22.00%; the F1 value on the WBSensor dataset is 46.91%, the precision is 42.42%, the recall is 52.45%, the mean average precision is 65.43%, and the IOU-AOC is 25.81%. This method is superior to other methods in all metrics. In the group behavior recognition task, the F1 value of the method of this application on UT-Group is 69.51%, the precision is 71.48%, the recall is 67.65%, and the accuracy is 68.79%; the F1 value on the WBSensor dataset is 75.84%, the precision is 73.45%, the recall is 78.39%, and the accuracy is 72.80%. In all metrics of this task, this method is also superior to other methods. In summary, the method of this application has certain superiority compared with other algorithms.

[0112] In another embodiment, this application also proposes a group behavior recognition device based on a multi-task self-supervised framework, including a processor and a memory storing a number of computer instructions, and the computer instructions, when executed by the processor, implement the steps of the above method.

[0113] For the specific limitations of the group behavior recognition device based on the multi-task self-supervised framework, reference may be made to the limitations of the group behavior recognition method based on the multi-task self-supervised framework in the foregoing text, which will not be elaborated herein. The above-mentioned group behavior recognition device based on the multi-task self-supervised framework can be implemented in whole or in part by software, hardware, and their combination. It can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations above.

[0114] The memory and the processor are electrically connected directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory stores a computer program that can run on the processor, and the processor realizes the group behavior recognition method based on the multi-task self-supervised framework in the embodiments of the present invention by running the computer program stored in the memory.

[0115] Among them, the memory can be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. The memory is used to store the program, and the processor executes the program after receiving the execution instruction.

[0116] The processor may be an integrated circuit chip with data processing capabilities. The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0117] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A method for group behavior recognition based on a multi-task self-supervised framework, characterized in that The group behavior recognition method based on the multi-task self-supervised framework includes: Obtain the behavior data of individuals to generate a first training data set and a second training data set; Construct a first network model including an individual feature extraction module, a Transformer encoder, and task heads. Use the first training data set to train the first network model. During training, the individual embedding features extracted by the individual feature extraction module are input into the Transformer encoder to extract encoded features, and then input into each task head to calculate the losses corresponding to each task head, completing the training of the first network model; Construct a second network model including an individual feature extraction module, a Transformer encoder, a global feature supplement module, and an inference module. Freeze the parameters of the Transformer encoder and use the second training data set to train the second network model. During training, the individual embedding features extracted by the individual feature extraction module are input into the global feature supplement module and the Transformer encoder, and the global features and encoded features obtained respectively are fused and then input into the inference module to calculate the loss, completing the training of the second network model; Input the collected individual behavior data into the trained second network model to obtain the group behavior recognition result.

2. The group behavior recognition method based on the multi-task self-supervised framework according to claim 1, characterized in that The task heads include a contrast learning task head, a behavior prediction task head, and a speed prediction task head; The contrast learning task head includes a first multi-layer perceptron, a second multi-layer perceptron, and a normalization layer connected in sequence. The encoded features are input into the contrast learning task head to obtain contrast learning features. The loss corresponding to the contrast learning task head is the InfoNCE loss; The behavior prediction task head uses a long short-term memory network. The encoded features are input into the behavior prediction task head to obtain behavior prediction values. The loss corresponding to the behavior prediction task head is the mean square error loss between the behavior prediction values and the true values; The speed transformation task head includes a mean pooling layer, a fully connected layer, and a softmax function. The encoded features are input into the behavior speed transformation task head to obtain the probability distribution of the predicted transformation magnification. The loss corresponding to the speed transformation task head is the cross-entropy loss between the probability distribution of the predicted transformation magnification and the true magnification.

3. The group behavior recognition method based on the multi-task self-supervised framework according to claim 2, wherein When training the first network model, the total loss function used is the sum of the losses corresponding to each task head.

4. The group behavior recognition method based on the multi-task self-supervised framework according to claim 1, wherein The global feature supplement module performs the following operations: Cluster the group feature matrix formed by splicing individual embedding features into each cluster; For each cluster, calculate the intra-cluster attention features; Obtain the center of each cluster by taking the mean of the vectors within the cluster; Stack all cluster centers into a cluster center matrix and calculate the inter-cluster attention features; Splice the intra-cluster attention features and the inter-cluster attention features after linear transformation, and then obtain the dynamic weight values through an activation function; Use the dynamic weight values to dynamically weight the intra-cluster attention features and the inter-cluster attention features to obtain the weighted attention features; Perform a linear transformation on each input individual embedding feature to obtain a value vector, and then multiply it by the weighted attention features to obtain the global features corresponding to the individual embedding features; Stack the global features corresponding to the embedded features of all the input individuals to obtain the output global features.

5. The group behavior recognition method based on the multi-task self-supervised framework according to claim 1, characterized in that The inference module performs the following operations: Input the fused features obtained by fusing the global features and the encoded features into a feed-forward neural network to predict the subgroup cardinality, and obtain the predicted subgroup cardinality; Based on the fused features, calculate the feature similarity between individuals and construct a feature adjacency matrix; According to the subgroup cardinality, perform spectral clustering operations on the feature adjacency matrix to obtain the subgroup segmentation result; Input the subgroup segmentation result into the group behavior classifier to obtain the group behavior recognition result.

6. The method for group behavior recognition based on a multi-task self-supervised framework according to claim 5, wherein When training the second network model, the loss function used includes subgroup cardinality loss, subgroup segmentation loss, and group behavior recognition loss.

7. A group behavior recognition device based on a multi-task self-supervised framework, comprising a processor and a memory storing a number of computer instructions, characterized in that, When the computer instructions are executed by the processor, the steps of the method described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • System and method for image mapping and visual attention

    CA2868135A1

  • Group behavior identification method based on multi-modal fusion and implicit interactive relationship learning

    CN115719510A

  • Self-supervised group behavior identification method based on sparse graph causal time sequence coding and identification system thereof

    CN116797972A

  • Self-supervised human behavior recognition method based on time-frequency contrast learning

    CN117113164A

  • Personnel group prediction method and device based on self-supervised graph attention network

    CN117435935A

Cited By

  • Typing and grading system for chronic venous insufficiency

    CN121148732A

  • Behavior recognition method based on single-view unmanned aerial vehicle cluster trajectory analysis

    CN121170644A