Group behavior recognition method and system based on attitude estimation and space-time modeling

By combining a dual-branch occlusion perception head and a posture estimation network with spatiotemporal modeling and crowd semantic analysis, the problem of group behavior recognition in occluded and dense scenes is solved, and efficient and accurate group behavior recognition and real-time warning are achieved.

CN120599693APending Publication Date: 2025-09-05WUHAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510641072.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing monitoring systems have difficulty accurately identifying individual behaviors and group interaction patterns in scenes with occlusion and dense crowds, and their computational efficiency is low, making them prone to tracking loss and trajectory disruption.

Method used

It adopts a dual-branch occlusion perception head and posture estimation network, combines spatiotemporal modeling and crowd semantic analysis, and achieves high-precision group behavior recognition through multi-resolution feature fusion and multi-dimensional feature matching.

Benefits of technology

It improves the accuracy and computational efficiency of behavior recognition in occluded and dense scenes, enhances the ability to understand complex group interactions, and provides efficient and accurate intelligent monitoring support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599693A_ABST
    Figure CN120599693A_ABST
Patent Text Reader

Abstract

The invention discloses a group behavior recognition method and system based on attitude estimation and space-time modeling, and belongs to the field of behavior recognition, and the method comprises the steps: S1, carrying out the target detection of an input video frame, and obtaining a single-person image detection frame and the position information of an abnormal object; s2, based on a single-person image detection frame, generating a high-precision key point heat map through an attitude estimation network; s3, converting the high-precision key point heat map, performing weighted fusion on the converted high-precision key point heat map and re-identification features, and performing feature matching on the re-identification features after weighted fusion to obtain an identity label; s4, inputting the high-precision key point heat map sequence into the space-time modeling network for identification to obtain position information and behavior categories; and S5, based on the identity label, the position information, the behavior category and the abnormal object position information of each person, identifying the group behavior. By identifying and analyzing the group abnormal behavior in the public place, the public safety management efficiency is improved, and the dynamic evaluation of the group behavior is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of behavior recognition, and in particular relates to a group behavior recognition method and system based on posture estimation and spatiotemporal modeling. Background Art

[0002] Traditional surveillance systems only provide video capture and storage capabilities, but are unable to provide real-time analysis and early warning, requiring manual monitoring. Manual monitoring has significant limitations. For one thing, the massive amount of video data far exceeds the processing capacity of security personnel, resulting in insufficient monitoring coverage. Furthermore, long hours can lead to fatigue, resulting in false or missed detections, severely impacting monitoring efficiency.

[0003] Existing methods for group behavior recognition can be mainly divided into two major technical routes: global feature analysis-based and trajectory analysis-based. In the global feature analysis method, researchers usually use 3D convolutional networks or spatiotemporal attention mechanisms to extract scene-level spatiotemporal features. Although such methods can capture the overall movement trends and energy distribution of the crowd, they treat the entire scene as a single analysis unit, making it difficult to accurately distinguish individual behavior characteristics and identify complex group interaction patterns. In the trajectory analysis-based method, the algorithm uses optical flow estimation or social force models to track individual movement trajectories and model group dynamics. Although such methods can provide more detailed motion analysis, their performance is highly dependent on accurate target tracking results. When occlusion or dense crowds occur in real scenes, tracking loss and trajectory breakage are prone to occur, resulting in a significant decrease in behavior recognition performance.

[0004] Therefore, it is necessary to design a group behavior recognition method and system based on posture estimation and spatiotemporal modeling that can integrate multi-dimensional information to address the above problems. Summary of the Invention

[0005] The purpose of the present invention is to provide a group behavior recognition method based on posture estimation and spatiotemporal modeling to address the problems of tracking loss and trajectory interruption that are prone to occur when there is occlusion or dense crowds. Through a dual-branch occlusion perception head and identity tracking that integrates posture estimation and spatiotemporal modeling, it improves the adaptability to occlusion and complex scenes while taking into account computational efficiency, providing a more accurate and efficient solution for group behavior recognition, effectively improving the adaptability of the algorithm in occlusion and dense scenes, and enhancing the ability to understand complex group interactive behaviors, thereby providing more efficient and accurate analysis support for intelligent monitoring systems in the fields of smart cities and public safety.

[0006] According to one aspect of this specification, a group behavior recognition method based on posture estimation and spatiotemporal modeling is provided, comprising:

[0007] S1. Use the target detection model to detect people and abnormal objects in the input video frame to obtain the single person image detection frame and abnormal object location information;

[0008] S2. Input the single person image detection frame into the pose estimation network to generate a high-precision key point heat map including occlusion conditions;

[0009] S3. Use the residual network to extract features from the high-precision key point heat map to obtain re-identification features. At the same time, the high-precision key point heat map is converted into a posture attention map, and weighted fused with the obtained re-identification features. Then, feature matching is performed on the weighted fused re-identification features to obtain the identity of each person.

[0010] S4: The high-precision key point heat map in S2 is stacked according to each person's identity to form a time-series high-precision key point heat map, and input into the spatiotemporal modeling network for behavior recognition to obtain each person's location information and behavior category;

[0011] S5. Based on each person's identity, location information, behavior category, and abnormal object location information, crowd semantic analysis is used to identify group behavior.

[0012] Furthermore, the S1 includes:

[0013] YOLOv5 is used as the target detection model, in which a cross-stage partial connection Darknet53 (CSPDarknet53) network is used to extract multi-scale features of the input video frame;

[0014] Input the multi-scale features of the input video frame into the feature pyramid network module to extract the multi-scale context features of the input video frame;

[0015] Based on the multi-scale context features of the input video frame, a path aggregation network is used for fusion;

[0016] Based on the fused multi-scale features, the single person image detection frame and abnormal object location information are obtained through the detection head and post-processing algorithm.

[0017] Furthermore, the construction of the posture estimation network in S2 includes:

[0018] A pose estimation network is constructed based on the High-Resolution Network (HRNet). The two-dimensional convolution of the high-resolution network is replaced by a dual-branch occlusion perception head to predict high-precision key points. A multi-resolution feature fusion mechanism and an occlusion processing module are established for feature interaction and improving estimation accuracy in cases of occlusion.

[0019] Furthermore, the step S3 performs feature matching on the weighted fused re-identification features, including:

[0020] Based on the weighted fusion of the re-identification features, a threshold judgment is performed. If it is higher than the set threshold, the identity is updated;

[0021] If it is not higher than the set threshold, the position and posture matching of the weighted fused re-identification features is performed, and after obtaining the fusion distance matrix, the threshold judgment is performed again to obtain the identity.

[0022] Furthermore, the construction of the spatiotemporal modeling network in S4 includes:

[0023] A spatiotemporal modeling network is constructed based on a lightweight 3D convolutional network, including convolutional layers, several residual blocks, pooling layers, and fully connected layers to capture spatiotemporal features from stacked high-precision keypoint heat maps.

[0024] Furthermore, the crowd semantic analysis in S5 includes:

[0025] Crowd semantic analysis is performed based on spatial distribution characteristics, interaction intensity indicators and abnormal object association characteristics.

[0026] Furthermore, the occlusion processing module includes parallel feature processing branches and an attention mechanism to enhance the representation capability of key features.

[0027] According to one aspect of this specification, a group behavior recognition system based on posture estimation and spatiotemporal modeling is provided, comprising:

[0028] The target detection module is used to detect people and abnormal objects in the input video frame using the target detection model, and obtain the single person image detection frame and abnormal object location information;

[0029] The pose estimation module is used to input the single-person image detection box into the pose estimation network to generate a high-precision key point heat map including occlusion;

[0030] The identity identification module is used to extract features from the high-precision key point heat map using a residual network to obtain re-identification features. The high-precision key point heat map is converted into a posture attention map, which is weightedly fused with the obtained re-identification features. The weighted fused re-identification features are then matched to obtain each person's identity.

[0031] The spatiotemporal modeling module is used to stack high-precision key point heat maps according to each person's identity to form a high-precision key point heat map sequence, and input it into the spatiotemporal modeling network for behavior recognition to obtain each person's location information and behavior category;

[0032] The group behavior recognition module is used to identify group behaviors using crowd semantic analysis based on each person's identity, location information, behavior category, and abnormal object location information.

[0033] According to one aspect of this specification, an electronic device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the group behavior recognition method based on posture estimation and spatiotemporal modeling when executing the computer program.

[0034] According to one aspect of the present specification, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the group behavior recognition method based on posture estimation and spatiotemporal modeling are implemented.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] 1. To address the problems of key point loss and insufficient target tracking accuracy in occlusion scenarios, the present invention adopts a dual-branch occlusion perception head and a posture estimation network, respectively, to alleviate the problems of key point and target loss in dense crowds through multi-resolution feature fusion, dual-branch prediction and multi-dimensional feature matching.

[0037] 2. To address the problem of insufficient modeling of the association between individual and group behaviors, the present invention adopts a spatiotemporal modeling network and crowd semantic analysis fusion framework. Through spatiotemporal feature extraction and customized crowd semantic rule discrimination, it not only realizes the analysis of complex group interaction behaviors, but also automatically identifies potential security risks, providing real-time warnings for security management. It is suitable for monitoring crowd activities in public places and improving emergency response capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0039] Figure 1 is a flow chart of a method according to an embodiment of the present invention;

[0040] Figure 2 This is a structural diagram of a dual-branch occlusion sensing head according to an embodiment of the present invention;

[0041] Figure 3 This is a structural diagram of the convolutional attention module in the dual-branch occlusion perception head according to an embodiment of the present invention;

[0042] Figure 4This is a diagram of the identity tracking structure with gesture enhancement according to an embodiment of the present invention;

[0043] Figure 5 This is a diagram of the lightweight three-dimensional convolutional network structure of an embodiment of the present invention. DETAILED DESCRIPTION

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0045] like Figure 1 As shown, an embodiment of the present invention provides a group behavior recognition method based on posture estimation and spatiotemporal modeling, including: step 1, using a target detection model to locate and detect people and abnormal objects (such as banners, knives, etc.) in video frames; step 2, extracting a single-person image area based on the detection result frame, and generating a high-precision key point heat map including occlusion through a posture estimation network; step 3, converting the high-precision key point heat map into a posture attention map, weightedly fused with the re-identification feature, and performing cross-frame target tracking through feature matching; step 4, inputting a multi-frame high-precision key point heat map sequence into the spatiotemporal modeling network to predict the behavior category of each person; step 5, comprehensively analyzing the position, identity ID, behavior category and abnormal object information of each person, and determining the group behavior category through preset semantic rules.

[0046] Specifically, the embodiment of the present invention also provides the specific content of step 1: using YOLOv5 as the target detection model, obtaining an input image frame, using YOLOv5 to detect people and abnormal objects therein, and obtaining their detection frames. Specifically, the input image frame is first scaled to a fixed size, and the pixel values ​​are normalized, and then CSPDarknet53 is used as the backbone network to extract multi-scale features of the image. CSPDarknet53 is composed of multiple CSP modules. Each CSP module divides the input feature map into two parts, one part undergoes dense convolution operations, and the other part directly skips the connection, and finally the two parts of features are spliced. Then, the multi-scale features of the input video frame are passed through the feature pyramid network module, and the receptive field is expanded through pooling operations of different scales, and multi-scale contextual features are extracted to enhance the spatial perception ability of the network. The path aggregation network aggregates paths at different scales, enhancing information flow and feature reconstruction, and achieving multi-scale feature fusion. Finally, the detection head uses the fused feature map to predict the location of the detection box and category probability. After post-processing such as non-maximum suppression (NMS) and confidence filtering, the final detection result, namely the detection box and the detected category (people and abnormal objects), is obtained.

[0047] Specifically, an embodiment of the present invention also provides the specific content of step 2: the posture estimation network adopted is the improved HRNet network. Since HRNet can maintain high-resolution feature maps and extract high-level semantic features, it is used as the backbone network, but the two-dimensional convolution ultimately used to predict key points is replaced by a dual-branch occlusion perception head. Through multi-resolution feature fusion and a dual-branch occlusion perception head, a high-precision key point heat map is generated under occlusion conditions, thereby improving the network's ability to extract key points of occluded human bodies.

[0048] Specifically, first the cropped image block Through two 3×3 convolution layers with a stride of 2, the resolution is obtained. Figure 1 Feature map of / 4 , the formula is as follows:

[0049] (1)

[0050] Specifically, it then passes through a backbone network consisting of four stages. Each stage consists of multiple parallel resolution streams. The first stage contains only one high-resolution stream. Starting from the second stage, a lower resolution parallel stream is added. Each resolution stream in each stage is composed of four residual blocks. Each residual block contains two 3×3 convolutional layers, and each convolutional layer is followed by Batch Normalization and ReLU activation functions. Multi-resolution feature fusion is performed after each stage, and features are exchanged between streams of different resolutions. If information is transmitted from a low-resolution stream to a high-resolution stream, bilinear upsampling and 1×1 convolution are used to align the number of feature channels. If information is transmitted from a high-resolution stream to a low-resolution stream, 3×3 convolution downsampling with a step size of 2 is used to achieve the target resolution. Stages feature streams, the input features are represented as , each feature map after fusion It is composed of the information and of all input features.

[0051] (2)

[0052] Among them, the transformation function The specific form of is determined by the resolution relationship, and the formula is as follows:

[0053] (3)

[0054] Specifically, the fusion features in the high-resolution stream in the fourth stage are retained at the end As the input of the dual-branch occlusion perception head, Figure 2 As shown in the figure, the dual-branch occlusion-aware head decomposes the keypoint detection task of a single person into the keypoint detection tasks of the occluder and the occluded person. It consists of two convolutional branches with identical structures, each of which is composed of two convolutional block attention modules and four convolutional layers with a kernel size of 3. Residual connections are included between each pair of convolutional layers to prevent gradient vanishing and promote information fusion between keypoints in the human body. The last module of the occluder branch is the final keypoint prediction head.

[0055] Specifically, if Figure 3 As shown, the convolutional attention module consists of channel attention and spatial attention, and the input feature map First, it will be sent to the channel attention part for processing. In this process, the feature map will undergo maximum pooling and average pooling operations respectively, so as to aggregate features on each channel. The aggregated features will be input into the fully connected layer and the channel attention mask will be generated through the Sigmoid activation function. .

[0056] (4)

[0057] in, is a fully connected layer.

[0058] Specifically, the input feature map With channel attention mask Element-by-element multiplication to obtain the channel enhanced feature map In the convolutional attention module, the spatial attention module aggregates each feature map in the spatial dimension using maximum pooling and average pooling. The aggregated feature map generates a spatial attention mask through splicing, convolution layer and Sigmoid activation function. .

[0059] (5)

[0060] Specifically, the channel enhanced feature map With spatial attention mask Multiply element by element to get the attention-weighted feature map , and then sent to two 3×3 convolutional layers to further extract features. Repeat the above convolutional attention module and two convolutional layers to obtain the final feature map , sent to the key point prediction head to obtain the key point heat map of the occluder .

[0061] (6)

[0062] in, It is the key point convolution prediction head, which consists of a 1×1 convolution layer and an activation function. The number of channels is mapped to the number of key points, the resolution of the input feature map is maintained, a heat map with the same size as the input is generated, and pixel-by-pixel convolution is performed to convert the feature value into the probability of the existence of the key point to obtain the key point heat map.

[0063] Specifically, the other branch is to fuse features and the output features of the previous branch The sum is taken as input, and the key point heat map of the occluded person is obtained through the same structure as the previous branch. .

[0064] Specifically, the embodiment of the present invention also provides the specific content of step 3, such as Figure 4 As shown in Figure 2, after obtaining each person's detection box and key point heat map, the ResNet network is used to extract re-identification features from each person's image block. , and use a convolutional layer to convert the key point heat map into a posture attention map and fused with the re-identification feature weights to obtain the enhanced re-identification feature map .

[0065] (7)

[0066] Specifically, for the current frame, from the tracking pool Get all identity embedding vectors of the previous frame , and then calculate the identity embedding vector of the current frame through cosine similarity Affinity matrix with all vectors in the pool to represent the similarity between two identity embedding vectors.

[0067] (8)

[0068] in is the identity embedding of the i-th person in the current frame, is the identity embedding of the jth person in the previous frame.

[0069] Specifically, if the similarity of some matches in the affinity matrix is ​​higher than the set threshold, the feature match is considered successful and the identity is updated. If the match fails, position and posture matching is performed. Specifically, the overlap of the bounding boxes is first calculated. , as a position constraint, to ensure that the spatial positions of the two people are similar. Then, the normalized posture distance is calculated by normalizing the key point coordinates. As a posture constraint.

[0070] (9)

[0071] Where N is the number of joints in the pose, and are the positions of the kth key point in the current frame and the previous frame respectively.

[0072] Specifically, the position constraints and posture constraints are weighted to obtain the fused distance matrix, and the threshold judgment is performed again to finally obtain the identity and tracking results of each person. The fused distance matrix is ​​as follows:

[0073] (10)

[0074] in, It is a weight parameter between [0,1], which indicates the importance of posture distance in fusion distance; The larger it is, the greater the impact of posture matching on the final distance, and the more emphasis feature matching places on the effect of posture similarity on feature matching. The smaller it is, the higher the weight of position overlap (IOU), and the more emphasis is placed on the effect of position similarity on feature matching.

[0075] Specifically, if the threshold is still not met after the threshold judgment is performed again, it indicates that the current target cannot be reliably associated with any existing identity in the tracking pool. The target is regarded as a newly appeared individual, a new unique identity is assigned to it, and the tracking trajectory is initialized to ensure continuous tracking in subsequent frames.

[0076] Specifically, the embodiment of the present invention also provides the specific content of step 4. The spatiotemporal modeling network uses a lightweight three-dimensional convolutional network to extract the key point heat map of each person from the input video in order to discriminate the target action. Afterwards, 12 frames are grouped as input. These heatmaps are then fed into a multi-stage 3D convolutional neural network for processing. The network's initial stage convolves the input data with a 1×7×7 convolution kernel to extract preliminary spatiotemporal features. The feature maps then pass through multiple residual blocks for further feature extraction. Each residual block consists of three convolutional layers, with the kernel size and number of channels varying across the corresponding residual blocks.

[0077] Specifically, the first stage uses three convolutional layers with sizes of 1×1×1, 1×3×3, and 1×1×1, respectively. In the second stage, the kernel size and number of channels are further increased to 1×1×1, 1×3×4, and 1×1×1, respectively. The subsequent third and fourth stages continue to extract features using larger kernels and more channels, with kernel sizes of 3×1×1, 1×3×3, and 1×1×1, respectively. These four stages are repeated three, four, six, and three times, respectively, to gradually enhance the expressive power of the features.

[0078] Specifically, in each residual block, the convolution operation can be expressed as , where Y is the output feature map, W is the convolution kernel weight, represents the 3D convolution operation, b is the bias term, is the activation function. Through this multi-layered feature extraction mechanism, the network effectively captures rich spatiotemporal features from the stacked keypoint heatmaps. Finally, the feature maps are processed through global average pooling and fully connected layers to output a probability distribution for the action categories. This process enables the system to accurately identify each individual action, resulting in the final action recognition result.

[0079] Specifically, this embodiment of the present invention also provides the details of step 5. Based on the above four steps, the system successfully identifies the location of the abnormal object, the location of each person, their identity, and their behavior category. To further analyze group behavior, a crowd semantic analysis module was designed. This module intelligently identifies group behavior by integrating the behavior category and location information of each person in the image frame, as well as the location of the abnormal object.

[0080] Specifically, the crowd semantic analysis module develops a series of judgment rules based on different behavioral characteristics. For example, gathering behavior is determined by detecting whether the number of people standing or walking in close proximity reaches a preset threshold; dispersal behavior is determined by analyzing whether the distance between historical gatherings has continued to increase; fighting behavior is determined by combining the distance between people and the intensity of the punches; banner-pulling and marching behaviors are identified by identifying the size of the crowd, the position of the banner, and the intensity of the crowd movement; looting and vandalism behaviors are determined by detecting whether anyone is holding dangerous objects (such as knives, sticks, etc.) and the intensity of the punches; running behavior is identified by analyzing the speed and size of the crowd; and sit-ins are determined by determining the distance and number of people between sitting and standing groups. Through this multi-dimensional, multi-rule analysis mechanism, the crowd semantic analysis module can efficiently and accurately identify various group behaviors, providing strong support for scene understanding and decision-making.

[0081] The implementation of each embodiment of the present invention is based on programmed processing by a device with processor functionality. Therefore, in practical engineering, the technical solutions and functions of each embodiment of the present invention are encapsulated into various modules. Based on this reality, and in addition to the aforementioned embodiments, an embodiment of the present invention provides a group behavior recognition system based on posture estimation and spatiotemporal modeling. This system is used to implement the group behavior recognition method based on posture estimation and spatiotemporal modeling described in the aforementioned method embodiment.

[0082] The system includes: a target detection module, which uses a target detection model to perform target detection on people and abnormal objects in the input video frame to obtain a single-person image detection frame and abnormal object location information; a posture estimation module, which uses a posture estimation network to generate a high-precision key point heat map containing occlusion based on the single-person image detection frame; an identity identification module, which converts the high-precision key point heat map containing occlusion into a posture attention map, performs weighted fusion with the re-identification feature, and then performs feature matching on the weighted fused re-identification feature to obtain each person's identity; a spatiotemporal modeling module, which stacks the high-precision key point heat map according to each person's identity to form a high-precision key point heat map sequence, and inputs it into the spatiotemporal modeling network for behavior recognition to obtain each person's location information and behavior category; a group behavior recognition module, which uses crowd semantic analysis to identify group behavior based on each person's identity, location information, behavior category and abnormal object location information.

[0083] The group behavior recognition system based on posture estimation and spatiotemporal modeling provided by the embodiment of the present invention addresses the problems of insufficient modeling of the association between individual and group behaviors, as well as the loss of key points and insufficient target tracking accuracy in occluded scenarios. It adopts several modules to alleviate the problem of key point and target loss in dense crowds through multi-resolution feature fusion, dual-branch prediction and multi-dimensional feature matching. Through spatiotemporal feature extraction and customized semantic rule discrimination, it realizes the analysis of complex group interaction behaviors and automatically identifies potential public security risks, providing real-time warnings for security management. It is suitable for monitoring crowd activities in public places and improving emergency response capabilities.

[0084] Based on the same inventive concept as the aforementioned embodiment, an embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions, thereby implementing a group behavior recognition method based on posture estimation and spatiotemporal modeling as proposed in the aforementioned embodiment.

[0085] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program. When executed by a processor, the program overcomes the problems of key point loss and insufficient target tracking accuracy in occluded scenes, and alleviates the problem of key point and target loss in dense crowds.

[0086] The storage medium can be any non-volatile storage device such as a hard disk, solid-state drive, flash drive, optical disk, etc., which is used to store computer program code and necessary data files. The stored computer program includes: target detection module, posture estimation module, identity identification module, spatiotemporal modeling module and group behavior recognition module.

[0087] Finally, it should be noted that the above specific embodiments are merely representative examples of the present invention. Obviously, the present invention is not limited to the above specific embodiments and is susceptible to numerous variations. Any simple modifications, equivalent variations, and modifications to the above specific embodiments based on the technical essence of the present invention shall be deemed to fall within the scope of protection of the present invention.

Claims

1. A group behavior recognition method based on posture estimation and spatiotemporal modeling, characterized in that: include: S1. Use the target detection model to detect people and abnormal objects in the input video frame to obtain the single person image detection frame and abnormal object location information; S2. Input the single person image detection frame into the pose estimation network to generate a high-precision key point heat map including occlusion conditions; S3. Use the residual network to extract features from the high-precision key point heat map to obtain re-identification features. At the same time, the high-precision key point heat map is converted into a posture attention map, and weighted fused with the obtained re-identification features. Then, feature matching is performed on the weighted fused re-identification features to obtain the identity of each person. S4, stacking the high-precision key point heat map in S2 according to each person's identity to form a high-precision key point heat map sequence, and inputting it into the spatiotemporal modeling network for behavior recognition to obtain each person's location information and behavior category; S5. Based on each person's identity, location information, behavior category, and abnormal object location information, crowd semantic analysis is used to identify group behavior.

2. The group behavior recognition method based on posture estimation and spatiotemporal modeling according to claim 1 is characterized in that: Said S1 comprises: YOLOv5 is used as the target detection model, in which the CSPDarknet53 network is used to extract multi-scale features of the input video frame; Input the multi-scale features of the input video frame into the feature pyramid network module to extract the multi-scale context features of the input video frame; Based on the multi-scale context features of the input video frame, a path aggregation network is used for fusion; Based on the fused multi-scale features, the single person image detection frame and abnormal object location information are obtained through the detection head and post-processing algorithm.

3. The group behavior recognition method based on posture estimation and spatiotemporal modeling according to claim 1 is characterized in that: The construction of the posture estimation network in S2 includes: A pose estimation network is constructed based on a high-resolution network, and the two-dimensional convolution of the high-resolution network is replaced by a dual-branch occlusion perception head to predict high-precision key points; a multi-resolution feature fusion mechanism and an occlusion processing module are established for feature interaction and improving estimation accuracy in cases involving occlusion.

4. The group behavior recognition method based on posture estimation and spatiotemporal modeling according to claim 1 is characterized in that: In S3, feature matching is performed on the weighted fused re-identification features, including: Based on the weighted fusion of the re-identification features, a threshold judgment is performed. If it is higher than the set threshold, the identity is updated; If it is not higher than the set threshold, the position and posture matching of the weighted fused re-identification features is performed, and after obtaining the fusion distance matrix, the threshold judgment is performed again to obtain the identity.

5. The group behavior recognition method based on posture estimation and spatiotemporal modeling according to claim 1 is characterized in that: The construction of the spatiotemporal modeling network in S4 includes: A spatiotemporal modeling network is constructed based on a lightweight 3D convolutional network, including convolutional layers, several residual blocks, pooling layers, and fully connected layers to capture spatiotemporal features from stacked high-precision keypoint heat maps.

6. The group behavior recognition method based on posture estimation and spatiotemporal modeling according to claim 1 is characterized in that: The crowd semantic analysis in S5 includes: Crowd semantic analysis is performed based on spatial distribution characteristics, interaction intensity indicators and abnormal object association characteristics.

7. The group behavior recognition method based on posture estimation and spatiotemporal modeling according to claim 3 is characterized in that: The occlusion processing module includes parallel feature processing branches and an attention mechanism, which is used to enhance the representation capability of key features.

8. A group behavior recognition system based on posture estimation and spatiotemporal modeling, characterized by: include: The target detection module is used to detect people and abnormal objects in the input video frame using the target detection model to obtain the single person image detection frame and abnormal object location information; The pose estimation module is used to input the single-person image detection box into the pose estimation network to generate a high-precision key point heat map including occlusion; The identity identification module is used to extract features from the high-precision key point heat map using a residual network to obtain re-identification features. The high-precision key point heat map is converted into a posture attention map, which is weightedly fused with the obtained re-identification features. The weighted fused re-identification features are then matched to obtain each person's identity. The spatiotemporal modeling module is used to stack high-precision key point heat maps according to each person's identity to form a high-precision key point heat map sequence, and input it into the spatiotemporal modeling network for behavior recognition to obtain each person's location information and behavior category; The group behavior recognition module is used to identify group behaviors using crowd semantic analysis based on each person's identity, location information, behavior category, and abnormal object location information.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the group behavior recognition method based on posture estimation and spatiotemporal modeling are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the group behavior recognition method based on posture estimation and spatiotemporal modeling are implemented.

Citation Information

Cited By

  • Body posture recognition method and system, intelligent terminal and storage medium

    CN121281099A

  • Power transmission line mechanized operation target detection method and device based on deep learning

    CN121415325A

  • Deep Learning-Based Target Detection Method and Equipment for Mechanized Operations on Transmission Lines

    CN121415325B