Open Vocabulary Group Behavior Detection Method for Indoor Security Monitoring Video Scenes

By using the video-text multimodal codec model and the video-individual position information multimodal cross attention model in indoor security monitoring video scenarios, combined with manual and automatic annotation, the problem of identifying the group and group behavior categories to which individuals belong is solved, and open vocabulary detection and improving generalization ability is achieved.

CN119851351BActive Publication Date: 2025-06-17HEFEI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510315321.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-17
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The prior art is difficult to identify the group to which each individual belongs and the behavioral categories of each group in indoor security monitoring video scenarios, and cannot adapt to the needs of open vocabulary detection, and the generalization performance is poor.

Method used

The multimodal cross-attention model based on the video-text multimodal codec model is adopted. The ternary annotation results are obtained through manual fine-grained annotation and automatic global semantic annotation. The model is trained to identify the group to which each individual belongs and the behavioral categories of each group, and the parameter update of the visual encoder and text encoder is constrained by regular terms to reduce knowledge forgetting.

Benefits of technology

It realizes the accurate identification of the group to which each individual belongs and the behavioral categories of each group in the indoor security monitoring video scenario, adapts to the needs of open vocabulary detection, and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851351B_ABST
    Figure CN119851351B_ABST
Patent Text Reader

Abstract

The present invention provides an open-vocabulary group behavior detection method for indoor security monitoring video scenarios, belonging to the field of video action recognition. The steps are as follows: S1: Collect and process indoor scene monitoring videos, obtain valid video segments containing people, and obtain the triple annotation results <video, text, flag> of each valid video segment; S2: For each frame of the video and the corresponding text, use the Swin-B and BERT structures of the CLIP pre-trained model as the image and text encoders respectively; the parameters of Swin-B and BERT are both updated and constrained by regularization terms, and finally the image-text encoder is determined; S3: Construct, train and determine the open-vocabulary group behavior detection model; S4: Input the actual monitoring video into the open-vocabulary group behavior detection model to obtain the behavior category of each group. The present invention can simultaneously identify which group each person in the indoor security monitoring video belongs to, classify the behaviors of each group at the same time, and also meet the open-vocabulary detection requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video action recognition, and particularly to an open-vocabulary group behavior detection method for indoor security monitoring video scenarios. Background Art

[0002] The purpose of group behavior detection for security monitoring video scenarios is to automatically perform group localization on each frame of the monitoring video using computer technology, that is, to identify how many groups there are in the current frame and which group each person belongs to, and at the same time classify the activities carried out by each group. This is of great significance to public security. Since the number of groups and their members in each frame are unknown, this task needs to simultaneously identify group behavior participants, perceive the spatio-temporal correlation relationships of different participants, and identify individual behaviors, which is very challenging.

[0003] Due to these difficulties, most of the existing methods for understanding group activities are limited to a single task of classifying the entire video clip into one of the predefined activity categories, called group behavior recognition. The conventional setting of group behavior recognition usually assumes that each video clip contains only a single group activity, and the activity participants have been pre-determined through manual annotation. However, these assumptions do not hold in many real monitoring scenario videos containing group behaviors, that is, there are often multiple groups performing independent activities in the video, as well as unrelated individuals who do not belong to any group. For example, in the video monitoring footage of a coffee shop, there are usually multiple group behaviors such as checking out, eating, chatting, queuing, and studying.

[0004] Therefore, although group behavior recognition has existed for more than a decade as a representative task of group behavior understanding, its practical application value in real security monitoring video scenarios has great limitations. For example, the patent document with the publication number CN117789094A provides a group behavior detection and recognition method based on deep learning. Although it comprehensively considers multi-level information such as individual motion characteristics, group member interaction characteristics, and group overall characteristics, it still assumes that there is only one group in a frame, and those who do not belong to this group are all abnormal interference items, which does not conform to the actual situation of security monitoring.

[0005] Another example is the group sudden abnormal event detection and localization method provided by the patent document with the publication number CN107506734A, and the violent group behavior detection method based on hierarchical cascading provided by the patent document with the publication number CN105574489A. The core of both is to automatically identify the frame as normal or abnormal, and it is impossible to output which group each member in the monitoring frame belongs to and the specific behavior category of each group.

[0006] In addition, all existing group behavior detection methods for the security monitoring video scenario are closed-set detections, that is, the group behavior categories in the test set or the actual usage scenario are the same as those in the training set, resulting in poor generalization performance of the model. The group behavior detection model trained in one monitoring scenario cannot be applied to a new monitoring scenario. Summary of the Invention

[0007] The technical problem to be solved by the present invention is how to propose a group behavior detection method that can simultaneously identify which group each person in the indoor security monitoring video belongs to, classify the behaviors of each group, and meet the requirements of open vocabulary detection.

[0008] To solve the above technical problems, the present invention provides the following technical solutions: An open vocabulary group behavior detection method for the indoor security monitoring video scenario, comprising the following steps:

[0009] S1: Collect and process the monitoring video data of the indoor scenario, obtain the effective video segments containing people, and obtain the triple annotation results <video, text, flag> of each effective video segment by combining manual fine-grained annotation and automatic global semantic annotation;

[0010] S2: Construct and train an image-text encoder based on the video-text multimodal encoding and decoding model, specifically:

[0011] According to the triple annotation results, for each frame of the video and the corresponding text, use the Swin-B structure of the CLIP pre-trained model as the image encoder and the BERT structure as the text encoder; the parameters of Swin-B and BERT participate in the optimization of the training stage, and at the same time, use a regularization term to constrain the update of the Swin-B and BERT parameters and determine the final image-text encoder;

[0012] S3: Construct and train an open vocabulary group behavior detection model based on the video-individual position information multimodal cross-attention model, and optimize the parameters to determine the open vocabulary group behavior detection model;

[0013] S4: Input the actual monitoring video into the determined open vocabulary group behavior detection model to obtain the behavior categories of each group.

[0014] The present invention utilizes video and text information, extracts features by using the Swin-B structure in the CLIP pre-trained model as the image encoder and the BERT structure as the text encoder, measures the importance of parameters by the degree of gradient change during parameter fine-tuning during training, and designs a regularization term to constrain the parameter update process of the visual encoder and the text encoder. It not only effectively utilizes the parameters of the pre-trained model to adapt to the group behavior detection task, but also avoids the unconstrained update of Swin-B parameters and BERT parameters from causing forgetting of the knowledge learned in the CLIP image-text overall semantic similarity task, reduces knowledge forgetting, and makes it suitable for the open vocabulary scenario.

[0015] Preferably, the specific process of step S1 is as follows:

[0016] S11: Collect monitoring video data of indoor scenes;

[0017] S12: Define various group behavior categories, including fighting, dancing, queuing, checking out, eating, studying or working, taking pictures, chatting, playing ball, walking, running, and outliers;

[0018] S13: Set the length of the video segment and segment the monitoring data video to obtain a number of video segments;

[0019] S14: Use a pedestrian detection and tracking tool to perform human detection and tracking on all video segments, obtain the personnel detection and tracking frames, remove the video segments without people, and obtain the remaining valid video segments;

[0020] S15: Use manual fine-grained annotation and automatic global semantic annotation to annotate the valid video segments to obtain the triple annotation result <video, text, flag>;

[0021] S16: Divide the valid video segments into a training set, a validation set, and a test set.

[0022] Preferably, the specific process of step S15 is as follows:

[0023] S151: Perform manual fine-grained annotation: On the basis of obtaining the personnel detection and tracking frames, annotate the group number and the corresponding group behavior category of each tracked personnel individual, and use the smallest rectangular bounding box jointly formed by the personnel with the same group number as the position information of the group;

[0024] S152: On the basis of the manual fine-grained annotation, perform automatic global semantic annotation: For a certain video to be annotated , if it is judged that it has j types of group behavior categories, then annotate it as j strings. According to j strings and the total number of group behavior categoriesA Automatically generate A - j additional strings, representing the group behavior categories that do not exist in the video segment;

[0025] S153: Use A strings to annotate the video to obtain the triple annotation result <video , text , flag >, where the text represents all defined group behavior categories existing in the video , and the flag is 1 or 0, indicating affirmation and negation of the text respectively, that is, the flag being 1 means that the group behavior category represented by the text exists in the video , and the flag being 0 means that the group behavior category represented by the text does not exist in the video .

[0026] The present invention provides an effective data preparation and annotation strategy, including a combination of manual fine-grained annotation and automatic global semantic annotation. In particular, the automatic global semantic annotation process can quickly generate strings describing the content of video segments. The generated strings not only contain the existing group behavior categories but also clearly indicate the behavior types that are not included, thereby being able to automatically generate a large number of <video, text, flag> triple annotation results, greatly enriching the diversity of training data and supporting model training under open vocabulary settings.

[0027] Preferably, in step S16, when dividing the valid video segments into a training set, a validation set, and a test set, it is restricted that the group behavior categories in the test set do not appear in the training set, but the group behavior categories in the training set and the validation set are the same.

[0028] The present invention restricts the group behavior categories in the test set from appearing in the training set when dividing the valid video segments, but the group behavior categories in the training set and the validation set are the same, which can meet the open vocabulary setting and identify the behavior categories that do not appear in the training set. That is, the present invention can still work effectively even when the group behavior categories in the test set or the actual usage scenario and the behavior categories in the training set have little overlap, improving the generalization ability of group behavior detection in new monitoring scenarios.

[0029] Preferably, the specific process of step S2 is as follows:

[0030] S21: For the video For each frame, the Swin-B structure in the CLIP pre-trained model is used as the image encoder to extract the image features of each frame. The parameters of Swin-B participate in the optimization and fine-tuning during the training phase, and the video is represented as a feature sequence . According to the feature sequence , the video 's video feature representation is obtained;

[0031] S22: For the text , the BERT structure in the CLIP pre-trained model is used as the text encoder to extract the text 's text feature representation , and the parameters of BERT participate in the optimization and fine-tuning during the training phase;

[0032] S23: Construct a video-text multimodal encoding and decoding model, concatenate the video feature representation and the text feature representation along the feature channel dimension to obtain the video-text multimodal fusion feature;

[0033] S24: Obtain the probability value from the video-text multimodal fusion feature, and set it as the probability value when the flag is 1. Then the probability that the flag is 0 is 1 - ;

[0034] S25: Determine the loss value of the video-text multimodal encoding and decoding model according to the probabilities that the flag is 1 and 0;

[0035] S26: Set the regularization constraint loss function as a regular term to constrain the update process of the Swin-B parameters of the visual encoder and the BERT parameters of the text encoder;

[0036] S27: Define the total training loss function , and optimize the video-text multimodal encoding and decoding model on the training set using the gradient descent method to minimize the of the validation set, and obtain the finally determined image encoder and text encoder.

[0037] The present invention uses the regularization constraint loss function Constraining the updates of Swin-B parameters and BERT parameters can avoid the knowledge forgetting problem caused by the unconstrained updates of Swin-B parameters and BERT parameters in completely different scenarios of the targeted group behavior detection task and the contrastive learning task of the overall semantic similarity between pre-trained text and images in the CLIP multimodal model, thus preventing the loss of the ability to represent the concept features of group behaviors that do not exist in the training set. That is, the present invention can better achieve open-vocabulary group behavior detection.

[0038] Preferably, in step S21, according to the feature sequence to obtain the video of the video feature representation The specific process is as follows: The feature sequence is mapped to a feature through a fully connected layer; the feature is passed through a tangent non-linear transformation layer to obtain a feature ; the feature is passed through a linear mapping layer and a Softmax layer to obtain a weight vector , the magnitude of its value represents the weight of the feature of the corresponding frame in the final video feature representation of the entire video; the feature and the weight vector are multiplied to obtain the video feature representation after frame fusion .

[0039] Preferably, the specific process of step S3 is as follows:

[0040] S31: Obtain all frame image data of each video , the temporal position information of all individuals in each frame, the group number to which the individual belongs, and the group behavior category corresponding to the group, and record them, and obtain the <video , text , flag , label > triple annotation result corresponding to the video;

[0041] S32: Design two encoders and decoders composed of a standard four-layer Transformer structure, and use the output of the encoder as the input of the decoder to construct a video-individual position information multimodal cross-attention model;

[0042] S33: Process all frames of the video through the optimized image encoder Swin-B to obtain the feature representations of all frames, and use a fully connected layer to process the temporal position information of all individuals in each frame to obtain an individual position embedding representation sequence,

[0043] S34: Input the feature representations of all frames and the sequence of individual position embedding representations for each frame into the encoder of the video-individual position information multi-modal cross-attention model;

[0044] S35: The decoder outputs the group membership embedding representation vectors and group behavior category embedding representation vectors corresponding to all individuals;

[0045] S36: Use the clustering loss function and the classification loss function as supervision signals to guide the optimization of the open-vocabulary group behavior detection model based on the video-individual position information multi-modal cross-attention model;

[0046] S37: After the optimization is completed, obtain the determined open-vocabulary group behavior detection model.

[0047] By processing continuous-frame video segments, combining the position information of each individual and the features of video frames, and implementing multi-modal cross-attention modeling of video-individual position information through the Transformer structure, the present invention can identify the group to which each individual in the video belongs and classify the behavior of each group, thereby achieving more accurate group behavior detection.

[0048] Preferably, in step S31, the process of obtaining the video The specific process of obtaining the temporal position information of all individuals in each frame is as follows: Use a pedestrian detection and tracking tool to obtain the minimum rectangular bounding box of the area where all individuals are located in each frame, and use to represent the center coordinates and width and height of the minimum rectangular bounding box of the th individual. Represent the temporal position information of the m th individual as , where represents the total number of individuals in this video , and n represents the rd frame of the video n .

[0049] Preferably, the specific process of step S36 is as follows:

[0050] S361: Use the clustering loss function to constrain the group membership embedding representation vectors of all individuals belonging to the same group in the video to be consistent;

[0051] S362: Process the text corresponding to the video through the optimized text encoder BERT to obtain the vector representation based on the text description of the group behavior category;

[0052] Use the classification loss function Constrain the similarity between the group behavior category embedding representation vector corresponding to each individual and the vector representation based on the literal description of the group behavior category to tend to 1;

[0053] S363: Under the constraint of , complete the parameter optimization training of the open-vocabulary group behavior detection model based on the video-individual location information multi-modal cross-attention model by the stochastic gradient descent method.

[0054] Preferably, the specific process of step S4 is as follows:

[0055] S41: Split the actual surveillance video to obtain video segments ;

[0056] S42: Use the pedestrian detection and tracking tool to obtain the temporal position information of all individuals in each frame of the video segment ;

[0057] S43: Use the open-vocabulary group behavior detection model based on the video-individual location information multi-modal cross-attention model to obtain the group membership embedding representation vector and the group behavior category embedding representation vector of all individuals;

[0058] S44: Cluster the group membership embedding representation vectors of all individuals through the clustering algorithm, and individuals with the same group behavior category belong to the same group;

[0059] S45: There are a total of a individuals in the same group. Perform an inner product operation on the group behavior category embedding representation vectors corresponding to all individuals belonging to the same group and the vector representation based on the literal description of the group behavior category, and take the category corresponding to the maximum value of the inner product operation result as the behavior category of each individual;

[0060] Select a The most frequent behavior category among the a individuals as the final group behavior category of the group. If the behavior categories of the a individuals are all different, then select the behavior category with the largest inner product operation result among the

[0061] Compared with the prior art, the advantages of the present invention are as follows: (1) When multiple people appear in the surveillance video image simultaneously, the prior art can only output a group behavior category for the entire image, while the present invention can simultaneously detect which group each person belongs to and identify the behavior category corresponding to the group in one model, and can also identify outliers who do not belong to any group; (2) When performing group behavior detection, a regularization term is designed to constrain the parameter update process of the visual encoder and the text encoder, which can not only effectively utilize the parameters of the pre-trained model to adapt it to the group behavior detection task, but also reduce knowledge forgetting and meet the open vocabulary detection requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is a flowchart of an embodiment of the present invention;

[0063] Figure 2 is a pre-trained model framework of an image-text feature extractor of a video-text multimodal model in an embodiment of the present invention;

[0064] Figure 3 is an open vocabulary group behavior detection model framework of a video-individual position information multimodal cross-attention model in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0066] Embodiment

[0067] As Figure 1 shown, this embodiment provides an open vocabulary group behavior detection method for an indoor security surveillance video scenario, including the following steps:

[0068] S1: Collect and process the surveillance video data of the indoor scenario, obtain the effective video segments containing people, and obtain the triple annotation results <video, text, flag> of each effective video segment by combining manual fine-grained annotation and automatic global semantic annotation. The specific process is as follows:

[0069] S11: Collect real indoor scenes, including surveillance video data of various types such as libraries, restaurants, and indoor stadiums. In this embodiment, the above three types of scenes are taken as examples, and each type contains video data for a period of time captured from multiple different surveillance shooting angles. In this embodiment, 24-hour video data captured from 3 different surveillance shooting angles is selected, totaling 216 hours of video data;

[0070] S12: Define possible group behavior categories. In this embodiment, 12 categories are defined, including fighting, dancing, queuing, checking out, eating, studying or working, taking pictures, chatting, playing ball, walking, running, and loners;

[0071] S13: Set the length of the video segment. In this embodiment, 3 seconds is taken as the length of a video segment, and the original 216-hour video data is segmented to obtain 259,200 video segments. The original surveillance video frame rate is 25 frames per second, that is, each video segment has 75 frames;

[0072] S14: Use existing open-source pedestrian detection and tracking tools to perform human detection and tracking on all video segments to obtain person detection and tracking frames, and remove video segments without people. In this embodiment, a total of 170,000 valid video segments are obtained;

[0073] S15: Label the 170,000 video segments using manual fine-grained annotation and automatic global semantic annotation. Specifically:

[0074] S151: First, perform manual fine-grained annotation: Based on the obtained person detection and tracking frames, label the group number and corresponding group behavior category of each tracked person individual, and use the smallest rectangular bounding box formed by people with the same group number as the position information of the group;

[0075] S152: On the basis of manual fine-grained annotation, to make full use of the semantic information of existing image-text multimodal large models, perform automatic global semantic annotation:

[0076] For a video segment to be labeled , if it is determined that there are j types of group behavior categories, then label it with j strings. According to j strings and the total number of group behavior categories A automatically generate ( A - j ) additional strings to represent the group behavior categories that do not exist in this video segment.

[0077] In this embodiment, the video segment to be labeled has 3 behaviors of queuing, chatting, and loners, that is jIf it is equal to 3, the annotation result is 3 strings, namely "There is a group of people queuing in this video", "There is a group of people chatting in this video" and "There are outliers in this video". According to these 3 annotation results, 9 additional strings can be automatically generated based on the behavior category data, that is, the defined group behavior categories, which are 12 categories in this embodiment: "This video does not contain a fighting group", "This video does not contain a dancing group", "This video does not contain a checkout group", "This video does not contain an eating group", "This video does not contain a learning (working) group", "This video does not contain a photographing group", "This video does not contain a ball-playing group", "This video does not contain a walking group", "This video does not contain a running group". It can be seen that for each video clip 12 string annotations can be quickly obtained;

[0078] S153: Use 12 strings to annotate the video clip to obtain a triple annotation result <video , text , flag >. Among them, the text is automatically converted from . In this embodiment , indicates that all defined group behavior categories exist in the video . The flag is 1 or 0, indicating affirmation and negation of the text respectively. That is, the flag being 1 means that the group behavior category represented by the text exists in the video . The flag being 0 means that the group behavior category represented by the text does not exist in the video . The triple annotation results of this embodiment are shown in the following table:

[0079]

[0080] A fast global semantic annotation provided by this embodiment can quickly generate a string describing the content of a video segment. The generated string not only includes existing group behavior categories but also clearly indicates the behavior types that are not included, thereby enabling the automatic generation of a large number of <video, text, logo> triple annotation results, greatly enriching the diversity of training data, and supporting model training under open vocabulary settings.

[0081] S16: Divide 170,000 video segments into a training set, a validation set, and a test set. To meet the open vocabulary setting, during the division, it is restricted that the group behavior categories in the test set do not appear in the training set, but the group behavior categories in the training set and the validation set are the same, enabling effective operation even when there is basically no overlap between the group behavior categories in the test set or actual usage scenarios and those in the training set, improving the generalization ability of group behavior detection in new monitoring scenarios.

[0082] S2: As Figure 2 shown, construct and train an image - text encoder based on a video - text multimodal codec model. The specific process is as follows:

[0083] S21: For each frame of the video In this embodiment, there are 75 frames in total. Use the Swin - B structure in the commonly used CLIP pre - trained model in the field of image feature extraction as the image encoder to extract the image features of each frame. Here, the parameters of the CLIP pre - trained model Swin - B participate in the optimization and fine - tuning during the training phase, aiming to obtain a visual feature representation that is more conducive to group behavior detection. The dimension of the image feature of each frame is 1 * 512, that is, the video is represented as , and its dimension is 75 * 512;

[0084] To reduce the consumption of feature calculation resources, this feature sequence is mapped through a fully - connected layer to a feature with a dimension of 75 * 256. In this embodiment, the fully - connected layer here consists of 256 neurons;

[0085] The feature passes through a tangent linear transformation layer to obtain a feature with a dimension of 75 * 256; The feature passes through a linear mapping layer and a Softmax layer to obtain a weight vector with a dimension of 75 * 1 and a value range of 0 - 1. The magnitude of its value represents the weight of the corresponding frame feature in the final video feature representation of the entire video;

[0086] Multiply by Multiply to obtain the video feature representation after frame fusion ;

[0087] S22: For the text , use the BERT structure in the CLIP pre-trained model commonly used in the field of text feature extraction as the text encoder to extract the feature representation of the text, with the feature dimension being 1*512. Here, the parameters of BERT participate in the optimization and fine-tuning during the training stage, aiming to obtain a text feature representation more conducive to group behavior detection;

[0088] S23: Construct a video-text multimodal encoding and decoding model, and splice the video feature representation and the text feature representation along the feature channel dimension to obtain a video-text multimodal fusion feature with a dimension of 1*768;

[0089] S24: After being decoded by a fully connected layer and a Sigmoid layer, the fusion feature obtains a probability value , which is set as the probability value of the flag being 1, then the probability of the flag being 0 is 1 - ;

[0090] S25: Obtain the calculation formula for the loss value of the video-text multimodal encoding and decoding model according to the probabilities of the flag being 1 and 0: ;

[0091] S26: Since the group behavior detection task targeted by the present invention and the pre-training image-text overall semantic similarity comparison learning task of the CLIP multimodal model belong to completely different task scenarios, unrestrained updating of the parameters of Swin-B and the parameters of BERT will lead to the problem of knowledge forgetting, that is, forgetting the knowledge learned in the CLIP image-text overall semantic similarity task and losing the ability to represent the group behavior concept features that do not exist in the training set, which is not conducive to open-vocabulary group behavior detection.

[0092] Therefore, the present invention sets a regularization term to constrain the parameter update process of the visual encoder and the text encoder. Specifically: Define the regularization constraint loss function as: , where K represents the total number of parameters in the image encoder Swin-B and the text encoder BERT, k represents the k th parameter; and respectively represent the th parameter after the current fine-tuning update and the original pre-trained parameter. The importance of the -th parameter calculated according to the gradient change to maintain the knowledge memory of the original pre-training task is calculated by the formula: , where represents the number of all <video, text, logo> triples in the training set, represents the -th training sample;

[0093] S27: Define the total training loss function , and optimize the video-text multi-modal encoding and decoding model on the training set using the gradient descent method to minimize the of the validation set, and obtain the final image encoder and text encoder after optimization and fine-tuning for the open-vocabulary group behavior detection task in the indoor security monitoring video scenario.

[0094] S3: As Figure 3 shown, construct and train an open-vocabulary group behavior detection model based on a video-individual location information multi-modal cross-attention model, and optimize the parameters to determine the open-vocabulary group behavior detection model. The specific process is as follows:

[0095] S31: Process 75 consecutive video frames, and obtain the following information for each video frame according to step S1 :

[0096] Information 1: 75-frame image data, denoted as ;

[0097] Information 2: The minimum rectangle bounding box of the area where all individuals are located in each frame, the group number to which the individual belongs, and the group behavior category corresponding to the group, denoted as , where represents the total number of individuals in the video segment , the subscript represents the -th individual, represents the -th individual's center point coordinates and width and height of the minimum rectangle bounding box in the -th frame of the video, that is, the n -th individual's temporal position information in the -th frame. In this embodiment, n , if the -th frame does not contain the individual, then set the value of to . represents the -th individual's group number in the video segment , represents the -th individual in the video segment The behavior categories of the group where it is located;

[0098] In this embodiment, the m representation of the nth individual in the entire video clip is: ;

[0099] Information Three: The 12 groups corresponding to the video clip <Video , text , logo > triple annotation results.

[0100] S32: Design two encoders and decoders composed of standard 4-layer Transformer structures. The output of the encoder is used as the input of the decoder to construct a video-head position information multi-modal cross-attention model;

[0101] S33: The 75-frame image data After being optimized, the image encoder Swin-B calculates and outputs 75-frame feature representations , and the dimension of each frame is 1*256;

[0102] M The individual temporal position information passes through a fully connected layer (consisting of 256 neurons and a RELU non-linear mapping function) to obtain an individual position embedding representation sequence , M For the total number of individuals in the video;

[0103] In this embodiment , ;

[0104] S34: Input the feature representation and the individual position embedding representation sequence into the encoder of the video-individual position information multi-modal cross-attention model. That is, the total input feature of the encoder is ;

[0105] S35: The decoder outputs M the M group membership embedding representation vectors corresponding to M individuals and M the group behavior type embedding representation vectors. That is, the output of the decoder is represented by 2 vectors, denoted as , where ( ) respectively represent the group membership embedding representation vector and the group behavior category representation vector corresponding to the

[0106] S36: Use the clustering loss function and the classification loss function as supervision signals to guide the optimization of the open-vocabulary group behavior detection model based on the video-individual location information multimodal cross-attention model. Specifically:

[0107] S361: Since the number of groups with the meaning of video segments contained in different video segments is different, in order to enable the open-vocabulary group behavior detection model based on the video-individual location information multimodal cross-attention model to be applicable to all video segments to achieve the end-to-end training optimization effect, this embodiment defines that a video segment contains at most 6 groups, that is , define the clustering loss function The calculation formula of is:

[0108]

[0109] Use the clustering loss function To constrain the group membership embedding representation vectors of all individuals belonging to the same group in the video To be consistent;

[0110] S362: Constrain the group behavior category embedding representation vector corresponding to each individual to be closest to the vector representation Based on the literal description of the group behavior category through the classification loss function Specifically:

[0111] First, 12 sentences of text pass through the text encoder BERT in the pre-training step of the image-text feature extractor of the video-text multimodal encoding and decoding model, that is, the text encoder BERT determined after optimization, to obtain 12 vector representations Based on the literal description of the group behavior category , and then for the group behavior category embedding representation of each individual output by the encoder , calculate its similarity with each vector representation Based on the literal description of the group behavior category In this embodiment ), so that when The similarity is close to 1, When the similarity is close to 0, Represents the behavior category of the group where the th individual is located, and define the classification loss The calculation formula of is:

[0112] ;

[0113] S363: In Under the constraints, the parameter optimization training of the open-vocabulary crowd behavior detection model based on the video-individual position information multimodal cross-attention model is completed by the stochastic gradient descent method;

[0114] S37: After the optimization is completed, a determined open-vocabulary crowd behavior detection model is obtained.

[0115] In the embodiment of the present invention, by processing continuous-frame video segments, combining the position information of each individual and the features of video frames, and realizing multimodal cross-attention modeling of video-individual position information through the Transformer structure, it is possible to identify the group to which each individual in the video belongs and classify the behaviors of each group, thereby realizing more accurate crowd behavior detection.

[0116] S4: Input the actual monitored video into the determined open-vocabulary crowd behavior detection model to obtain the behavior categories of each group, specifically:

[0117] S41: In actual application, first segment the indoor security monitoring video with a step size of 75 frames. For the indoor security monitoring video segment with 75 frames ;

[0118] S42: Use the same pedestrian detection and tracking algorithm as in step S1 to obtain the temporal position information of all individuals in each frame of the video segment: , , m represents the th individual in the video segment m , M represents the total number of individuals;

[0119] S43: Use the determined open-vocabulary crowd behavior detection model in step S3 to obtain the M group membership embedding representation vectors of M individuals and the M group behavior category embedding representation vectors of individuals, that is, obtain 2

[0120] vector representations: ;

[0121] S44: Cluster through the KMeans clustering algorithm (the number of cluster centers is set to 6), and individuals with the same group behavior category belong to the same group;

[0121] S45: Finally, perform an inner product operation on the group behavior category embedding representation vectors corresponding to all individuals belonging to the same group and the vector representation based on the text description of the group behavior category, and take the category corresponding to the maximum value of the inner product operation result as the behavior category of each individual;

[0122] There are a total ofa individuals, select a the most frequent behavior category among the a individuals as the final group behavior category of this group. If a the behavior categories of all

[0123] individuals are different, then select the behavior category with the largest inner product operation result among the k individuals as the final group behavior category of this group; For example: If a group has 5 individuals, which are respectively and their corresponding embedded representation vectors of group behavior categories are respectively ; Each embedded representation vector of group behavior category performs an inner product operation with 12 vector representations based on the literal description of group behavior categories, and select the category corresponding to the maximum value of the inner product operation as the behavior category of this individual, denoted as , where represents the category number, is the cosine similarity obtained from the inner product operation; Count the mode in

[0124] as the final group behavior category. If the behavior categories of the 5 individuals are all different, select the category corresponding to the maximum value in as the final group behavior category. The above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: They can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An open vocabulary group behavior detection method for indoor security surveillance video scenes, characterized in that: The following steps are involved: S1: Collect and process surveillance video data of indoor scenes, obtain valid video clips containing people, and combine manual fine-grained annotation and automatic global semantic annotation to obtain the ternary annotation results <video, text, logo> of each valid video clip; S2: Build and train an image-text encoder based on a video-text multimodal encoding and decoding model, specifically: According to the ternary annotation results, for each frame of the video and the corresponding text, the Swin-B structure of the CLIP pre-trained model is used as the image encoder and the BERT structure is used as the text encoder; the parameters of Swin-B and BERT are optimized in the training phase, and the regularization term is used to constrain the update of Swin-B and BERT parameters and determine the final image-text encoder; S3: Construct and train an open vocabulary group behavior detection model based on the video-individual position information multimodal cross-attention model, and perform parameter optimization to determine the open vocabulary group behavior detection model; S4: Input the actual surveillance video into the determined open vocabulary group behavior detection model to obtain the behavior category of each group; The specific process of step S4 is: S41: Segment the actual surveillance video to obtain video clips ; S42: Obtain video clips using pedestrian detection and tracking tools The temporal position information of all individuals in each frame; S43: Use the open vocabulary group behavior detection model based on the video-individual position information multimodal cross-attention model to obtain the group belonging embedding representation vectors and group behavior category embedding representation vectors of all individuals; S44: Cluster the group belonging embedding representation vectors of all individuals through a clustering algorithm, and individuals with the same group behavior category belong to the same group; S45: Perform an inner product operation on the group behavior category embedding representation vector corresponding to all individuals belonging to the same group and the vector representation based on the text description of the group behavior category, and take the category corresponding to the maximum value of the inner product operation result as the behavior category of each individual.

2. The open vocabulary group behavior detection method for indoor security monitoring video scenes according to claim 1 is characterized in that: The specific process of step S1 is: S11: Collect surveillance video data of indoor scenes; S12: Define various types of group behavior, including fighting, dancing, queuing, checking out, eating, studying or working, taking pictures, chatting, playing ball, walking, running, and loner; S13: setting the length of the video segment and dividing the monitoring data video to obtain a plurality of video segments; S14: using a pedestrian detection and tracking tool to perform human body detection and tracking on all video clips, obtaining a person detection and tracking frame, removing video clips that do not contain people, and obtaining the remaining valid video clips; S15: annotate the effective video clips using manual fine-grained annotation and automatic global semantic annotation to obtain a triplet annotation result <video, text, logo>; S16: Divide the valid video clips into a training set, a validation set, and a test set.

3. The open vocabulary group behavior detection method for indoor security monitoring video scenes according to claim 2 is characterized in that: The specific process of step S15 is: S151: Manual fine-grained labeling: Based on the obtained person detection and tracking frame, label each tracked person with the group number and the corresponding group behavior category, and the minimum rectangular bounding box formed by the persons with the same group number is used as the location information of the group; S152: Based on manual fine-grained annotation, automatic global semantic annotation is performed: for a video to be annotated , if it is determined that j group behavior category, it is marked as j Strings, according to j Total number of strings and group behavior categories A Automatically generated ( A - j ) additional strings, indicating the group behavior categories that are not present in the video clip; S153: Utilization A String to video Labeling, get triple labeling results <video ,text , logo >, where the text Video All defined group behavior categories exist in 1 or 0, respectively, indicating the text Affirmation and negation, i.e. signs 1 means video There is text in The group behavior category represented by 0 means video Text does not exist in Represents the category of group behavior.

4. The open vocabulary group behavior detection method for indoor security monitoring video scenes according to claim 2 is characterized in that: In step S16, when the valid video clips are divided into a training set, a validation set and a test set, the group behavior categories in the test set are restricted to not appear in the training set, but the group behavior categories in the training set and the validation set are consistent.

5. The open vocabulary group behavior detection method for indoor security monitoring video scenes according to claim 4 is characterized in that: The specific process of step S2 is: S21: For Video For each frame of the video, the Swin-B structure in the CLIP pre-trained model is used as the image encoder to extract the image features of each frame. The parameters of Swin-B are optimized and fine-tuned in the training phase. Represented as a feature sequence , according to the feature sequence Get the video Video feature representation ; S22: For text , using the BERT structure in the CLIP pre-trained model as a text encoder to extract text Text feature representation , BERT’s parameters are optimized and fine-tuned during the training phase; S23: Construct a video-text multimodal encoding and decoding model to represent video features and text feature representation Splicing along the feature channel dimension to obtain video-text multimodal fusion features; S24: Obtaining probability values ​​based on video-text multimodal fusion features , set as a flag The probability value is 1, then the sign The probability of being 0 is 1- ; S25: According to the signs Determine the loss value of the video-text multimodal encoding and decoding model for the probability of 1 and 0 ; S26: Setting the regularized constraint loss function , as a regular term to constrain the update process of the Swin-B parameters of the visual encoder and the BERT parameters of the text encoder; S27: Define the total training loss function , and use the gradient descent method to optimize the video-text multimodal encoding and decoding model on the training set so that the validation set Minimum, finalized image encoder and text encoder are obtained.

6. The open vocabulary group behavior detection method for indoor security monitoring video scenes according to claim 5, characterized in that: In step S21, according to the feature sequence Get the video Video feature representation The specific process is: feature sequence Mapped into features through a fully connected layer ;feature After a tangent nonlinear transformation layer, the features are obtained ;feature After a linear mapping layer and a Softmax layer, the weight vector is obtained , the value of which indicates the weight of the feature of the corresponding frame in the final video feature representation of the entire video; feature and the weight vector Multiply to get the video feature representation after frame fusion .

7. The open vocabulary group behavior detection method for indoor security monitoring video scenes according to claim 6, characterized in that: The specific process of step S3 is: S31: Get every video All frame image data, the temporal position information of all individuals in each frame, the group number of the individual, and the group behavior category corresponding to the group are recorded, and the video is obtained. Corresponding <Video ,text , logo >Triple labeling results; S32: Design two encoders and decoders consisting of a standard four-layer Transformer structure. The output of the encoder is used as the input of the decoder to build a video-individual position information multimodal cross-attention model; S33: Processing video with the optimized image encoder Swin-B All frames of the image are processed to obtain the feature representation of all frames, and a fully connected layer is used to process the temporal position information of all individuals in each frame to obtain the individual position embedding representation sequence; S34: The feature representations of all frames and the embedding representations of all individual positions in each frame are input into the encoder of the video-individual position information multimodal cross attention model; S35: The decoder outputs the group belonging embedding representation vector and group behavior category embedding representation vector corresponding to all individuals; S36: Using clustering loss function and classification loss function as supervision signals to guide the optimization of open vocabulary group behavior detection model based on video-individual position information multimodal cross-attention model; S37: After the optimization is completed, a determined open vocabulary group behavior detection model is obtained.

8. The open vocabulary group behavior detection method for indoor security monitoring video scenes according to claim 7, characterized in that: In step S31, obtain the video The specific process of obtaining the temporal position information of all individuals in each frame is as follows: using the pedestrian detection and tracking tool to obtain the minimum rectangular bounding box of the area where all individuals are located in each frame, Indicates The center point coordinates and width and height of the minimum rectangular bounding box of each individual m The temporal position information of each individual is expressed as ,in Indicates that the video The total number of individuals in n Video No. n frame.

9. The open vocabulary group behavior detection method for indoor security monitoring video scenes according to claim 8, characterized in that: The specific process of step S36 is: S361: Through clustering loss function Constraint Video The group belonging embedding representation vectors of all individuals belonging to the same group are consistent; S362: Processing Video with Optimized Text Encoder BERT Corresponding text , obtain the vector representation based on the text description of the group behavior category; Through the classification loss function Constrain the similarity between the embedding representation vector of the group behavior category corresponding to each individual and the vector representation based on the text description of the group behavior category to be close to 1; S363: Under the constraint of , the parameter optimization training of the open vocabulary group behavior detection model based on the video-individual position information multimodal cross-attention model is completed by the stochastic gradient descent method.

10. The open vocabulary group behavior detection method for indoor security monitoring video scenes according to claim 9, characterized in that: In step S4, there are a total of a Individuals, select a The most common behavior category among the individuals is taken as the final group behavior category of the group. a The behavior categories of each individual are different, so choose a The behavior category with the largest inner product operation result among the individuals is taken as the final group behavior category of the group.

Citation Information

Patent Citations

  • Layered stack based violent group behavior detection method

    CN105574489A

  • Method for detecting and locating emergent abnormal event of group

    CN107506734A

  • Group behavior detection and identification method and system based on deep learning

    CN117789094A

  • Volleyball group behavior identification method based on multi-modal information fusion

    CN111401174A

  • Action detection method and device, storage medium and electronic equipment

    CN118711255A