Weakly supervised group behavior recognition method based on motion diagram guided by visual conceptual knowledge
By extracting the visual concept knowledge and statistical information of individual actions, calculating the action diagram of the video, and combining semantic representation for enhancement and fusion, the problem of difficult connection between individual action visual information and semantic concepts in group behavior recognition tasks under weak supervision conditions is solved, and efficient and accurate group behavior recognition is achieved.
Patent Information
- Application Number
- CN202510080237.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-03
AI Technical Summary
Under weak supervision conditions, it is difficult for group behavior recognition tasks to effectively utilize the link between the visual information of individual actions and their semantic concepts, resulting in poor recognition performance.
By constructing a neural network, the visual concept knowledge of individual actions and statistical information related to group behavior are extracted, the action diagram of the video is calculated, and the action diagram is enhanced and fusion is enhanced and fusion based on semantic representation. Finally, the classification head of the fully connected layer structure is used to classify group behavior.
It realizes efficient and accurate identification of group behavior under weak supervision conditions, and improves recognition performance, especially in scenarios where individual labeling is difficult to obtain.
Smart Images

Figure CN120088852A_ABST
Abstract
Description
Technical Field
[0001] Based on deep learning technology, the present invention studies a method for identifying group behaviors under weakly supervised conditions. First, a neural network is used to learn and extract visual concept knowledge of individual actions, and statistical information related to group behaviors is obtained through statistical methods; then, a neural network is used to learn and extract visual features of the video, and the action map of the video is calculated using the visual concept knowledge; then, the statistical information related to group behaviors is used to strengthen the connection between the action map and specific group behaviors, and the fusion between action maps is carried out by combining the semantic representation of individual action categories; finally, a classification head with a fully connected layer structure is used for group behavior classification to achieve group behavior recognition under weakly supervised conditions. The present invention belongs to the field of computer vision, and specifically relates to technologies such as deep learning and behavior recognition. Background Art
[0002] With the continuous development of information technology, the task of identifying group behaviors in video data has received extensive attention from many research scholars. This task aims to analyze a given video sequence and identify the behavior categories jointly expressed by multiple related individuals therein. The task of identifying group behaviors has important practical application values. In the field of sports event analysis, using group behavior recognition technology, training and competition videos can be analyzed, athlete data can be statistically analyzed, providing assistance for coaches to guide training and deploy strategies, or classifying and sorting different sports scenes to provide users with efficient and accurate highlight retrieval services. In the field of surveillance videos, by analyzing group behaviors, abnormal situations of personnel can be automatically perceived, and real-time early warnings of abnormal events and dangerous behaviors can be realized, providing technical support for ensuring the personal and property safety of the masses and maintaining social security and stability.
[0003] In the task of identifying group behaviors, fully supervised algorithms align the visual representations of individuals with action concepts by using the annotations of individuals in the scene, and then infer group behaviors. These annotations usually include the individual bounding boxes and their action labels in the test stage. Although they show good performance, in practical applications, annotating the labels of a large number of individuals is time-consuming and expensive.
[0004] Therefore, weakly supervised group behavior recognition methods have been proposed, aiming to handle scenarios where individual annotations are difficult to obtain. They usually use object detectors to detect the positions of individuals, or use attention mechanisms to implicitly capture key regions to make up for the lack of individual annotation information. However, these methods lack a clear connection between the visual information of individual actions and their semantic concepts, and this connection has been proven to be beneficial for identifying group behaviors by fully supervised methods.
[0005] The present invention proposes a weakly supervised group behavior recognition method based on an action graph guided by visual concept knowledge. First, a neural network is constructed to learn and extract the visual concept knowledge of individual actions, and statistical information related to group behaviors is obtained through statistical methods; then, a neural network is used to learn and extract the visual features of the video, and the action graph of the video is calculated using the visual concept knowledge; afterwards, the statistical information related to group behaviors is used to enhance the connection between the action graph and specific group behaviors, and the fusion between action graphs is performed in combination with the semantic representation of individual action categories; finally, a classification head with a fully connected layer structure is used for group behavior classification to achieve group behavior recognition under weakly supervised conditions. Summary of the Invention
[0006] Different from the existing weakly supervised group behavior recognition methods, the present invention proposes a weakly supervised group behavior recognition method based on an action graph guided by visual concept knowledge. First, a neural network is constructed to learn and extract the visual concept knowledge of individual actions, and statistical information related to group behaviors is obtained through statistical methods. The visual concept knowledge of individual actions is represented by the mean of the individual action features extracted by the neural network, expressing the general visual representation of individual actions; the statistical information related to group behaviors is obtained by statistically analyzing the spatial distribution of individual actions under different group behaviors, describing the spatial distribution patterns of individual actions corresponding to each group behavior. Then, a neural network is used to learn and extract the visual features of the input video, and the action graph of the video is calculated using the visual concept knowledge. The visual concept knowledge and the visual features of the input video are matched in the frequency domain to efficiently capture the regions related to individual actions. Afterwards, the statistical information is used to enhance the connection between the action graph and specific group behaviors, and the fusion between action graphs is performed in combination with the semantic representation of individual action categories. The semantic representation of individual action categories is obtained by the Word Embedding method, explicitly indicating the corresponding relationship between the action graph and the semantic concept of individual actions. Finally, a classification head with a fully connected layer structure is used for group behavior classification to achieve group behavior recognition under weakly supervised conditions. The main process of this method is as shown in the appendix Figure 1 and can be divided into the following four steps: obtaining visual concept knowledge and statistical information, calculating the action graph of the input video, enhancing and fusing the action graph, and classifying group behaviors.
[0007] (1) Obtaining visual concept knowledge and statistical information
[0008] The present invention constructs a neural network and extracts the visual representation of individual actions, and defines the mean of the features of each individual action as the feature prototype of this type of individual action to vectorially represent visual concept knowledge. The present invention statistically analyzes the spatial distribution of each individual action under different group behaviors, summarizes and marks it on the corresponding plan view to obtain a behavior-action relationship diagram, so as to vectorially represent the statistical information related to group behaviors. The visual concept knowledge of individual actions and the statistical information related to group behaviors are both obtained from the training set data and directly transferred to the test environment to adapt to the weakly supervised conditions where individual annotations of test samples are not required.
[0009] (2) Calculate the action diagram of the input video
[0010] The visual features of the input video are extracted by a neural network, and the backbone of this network can be initialized with the network parameters for obtaining visual concept knowledge in step (1). The visual features of the input video are matched with the visual concept knowledge in the frequency domain to obtain the action diagram of the input video. This diagram implies the occurrence probability of each individual action in the input video in space. Different from existing methods, the present invention calculates the action diagram to capture the key areas where actions occur, and can explicitly align the semantic concepts of individual actions with the visual features of the input video.
[0011] (3) Enhancement and fusion of the action diagram
[0012] The present invention performs a dot product of vectors between the action diagram and the statistical information to provide the spatial distribution information of individual behaviors corresponding to specific group behaviors. And the word embeddings of the semantic representations of individual action categories are spliced onto the enhanced action diagram to explicitly indicate the corresponding relationship between the action diagram and the semantic concepts of individual actions, assisting in the fusion between action diagrams to obtain the group features of the input video.
[0013] (4) Group behavior classification
[0014] The present invention constructs a classification head with a fully connected layer structure to classify the group features of the input video, so as to realize the recognition of group behaviors under weakly supervised conditions.
[0015] Compared with the prior art, the present invention has the following obvious advantages and beneficial effects:
[0016] First of all, the visual concept knowledge of individual actions and the statistical information related to group behaviors proposed by the present invention are both obtained from the training set data and can be directly transferred to the test environment as general information about individual actions and group behaviors to adapt to test scenarios where it is difficult to obtain individual annotations, and alleviate the performance disadvantage caused by the lack of individual annotations under weakly supervised conditions. Secondly, different from existing methods, the action diagram proposed by the present invention can explicitly align the semantic concepts of individual actions with the visual features of the input video to make up for the defect that the captured area is separated from the action semantics in weakly supervised scenarios.
[0017] Experiments have confirmed that under the framework of a deep neural network based on the I3D (Inflated 3D ConvNet) backbone, the present invention can achieve 92.7% MCA (Multi-class Classification Accuracy) on the Volleyball Dataset. In addition, when only 25% and 50% of the training samples are used, the present invention can still obtain 88.3% and 91.0% MCA. Therefore, applying this method to the weakly supervised group behavior recognition task is feasible for efficiently and accurately recognizing group behaviors and has important application value. Description of the Drawings
[0018] Figure 1 Flowchart of the weakly supervised group behavior recognition method based on visual concept knowledge-guided action graphs
[0019] Figure 2 Schematic diagram of obtaining visual concept knowledge and calculating the action graph of the input video
[0020] Figure 3 Schematic diagram of action graph fusion Detailed Implementation Manner
[0021] According to the above description, the following is a specific implementation process, but the scope protected by this patent is not limited to this implementation process.
[0022] Step 1: Obtain visual concept knowledge and statistical information
[0023] Step 1.1: Obtain visual concept knowledge
[0024] The algorithm proposed by the present invention is implemented based on the mainstream open-source deep learning framework PyTorch. The flowchart of this step is shown in the attached Figure 2 first row.
[0025] Using I3D (inflated 3D ConvNet) as the backbone network, input the training set video samples to obtain visual features with the shape of h×w×C. Here, h, w, and C are the height, width, and the number of feature channels respectively. Then, use the RoiAlign operation to crop out the visual representation of each individual from the visual features of the samples, with the shape of p×p×C. Here, p is the height (equal to the width) of the representation block. Use a FC layer (Fully Connected Layer) to encode the visual representation of each individual into 1×D, and then input the encoded representation into another FC layer to obtain the predicted value of the action category of each individual. Here, D is the feature length. At the same time, use the Average Pooling operation to downsample the visual features of the sample video to 1×1×C, and then input it into a FC layer to obtain the predicted value of the group behavior category of the sample. Use formula (1) to supervise the classification of individual actions and group behaviors:
[0026]
[0027] Among them, represents the Cross-Entropy Loss, are the prediction results of group behaviors and individual actions, g and a n are the true value labels of group behaviors and individual actions, N is the number of individuals in the input video, and λ pre is the weight scalar used to balance the two classification losses. In the present invention, h, w, p, C, and D are taken as 90, 160, 7, 832, and 256 respectively, and λ pre is taken as 1. Use the Adaptive Moment Estimation (Adam) algorithm to optimize the above loss function, with the initial learning rate of 1e-4 and train for 50 rounds.
[0028] After the network training is completed, take the average value of the visual representations of each type of individual according to the corresponding action category to obtain K a feature prototypes of individual actions, with the shape of p×p×C, as the vectorized representation of the visual concept knowledge of individual actions. Here, K a is the number of action categories.
[0029] Step 1.2: Obtain statistical information
[0030] For a group behavior scenario with regular behavior areas and individual position distributions, such as a ball sports scenario, the present invention first performs a perspective normalization operation on the training set samples to exclude the interference of camera perspective changes on statistical information. In this step, the M-LSD (Mobile Line Segment Detection) algorithm is used to detect the landmark lines in the motion area, such as the field lines of a stadium. Then, the intersection points between the landmark lines are calculated, and a three-dimensional affine matrix from the camera perspective of the sample to the top view perspective is calculated based on the coordinates of the intersection points. According to this matrix, a projection transformation from the camera perspective to the top view perspective is performed on the sample, and the resolution of the sample after transformation is the same as the original resolution. In the present invention, the intersection points in the top view perspective are defined as the landmark points of the standard behavior areas in the top-down view, such as the boundary corner points of the stadium and the intersection points of the field lines.
[0031] After perspective normalization, for all individuals in the training set samples, in this step, the bottom center point and its neighborhood of their bounding boxes are marked into K g ×K a sub-images, where K g represents the number of group behavior categories. This marking process depends on the individual action category corresponding to each individual and the group behavior category to which it belongs. In the present invention, the neighborhood width is set to 7, and the sub-image resolution is set to 90×160. Then, all the sub-images are stitched together to obtain a behavior-action relationship graph of the form K g ×K a ×h×w as a vectorized representation of the statistical information related to group behavior.
[0032] Step 2: Calculate the action graph of the input video
[0033] The process flow diagram of this step is shown in the attached Figure 2 second row. For the input video, the I3D backbone network is used to extract its visual features, with the shape of h×w×C. The parameters of this backbone network can be initialized with the backbone network parameters in Step 1.1.
[0034] In this step, the visual features of the input video are matched with the feature prototypes of K a individual actions in the frequency domain to obtain the action graph of the input video. First, under a given individual action category, zero-padding is performed on the visual features of the input video and the feature prototype of this individual action, and their height and width of the resolution are extended spatially to (h+p-1) and (w+p-1). Then, two-dimensional fast Fourier transforms are performed on both of them, and the complex conjugate of the transformation result of the feature prototype is taken. The obtained results are subjected to vector dot multiplication, and then two-dimensional inverse fast Fourier transform is performed to obtain C feature maps. The feature maps are averaged to obtain an action sub-image corresponding to this individual action with a height of (h+p-1) and a width of (w+p-1). The zero-padded parts of the sub-image are cropped to make its resolution h×w. Stitch Ka Sub - graphs corresponding to a group of individual actions are obtained to form an action graph of the input video, with the shape of K a ×h×w.
[0035] Step 3: Enhancement and fusion of the action graph
[0036] Step 3.1: Action graph enhancement
[0037] In this step, the action graph is first broadcast to K g ×K a ×h×w to correspond to K g categories of group behaviors. Then, it is multiplied by the behavior - action relationship graph in vector dot - product to supplement the spatial distribution information of individual behaviors corresponding to specific group behaviors. After that, according to the corresponding group behavior categories, the results of the dot - product are split into K g groups of enhanced action graphs.
[0038] Step 3.2: Obtaining semantic representations
[0039] In this step, first, K a individual action labels are encoded into one - hot vectors. Then, an FC layer is used to embed them into a d - dimensional latent space to obtain semantic features of individual action categories with the shape of K a ×d. This semantic representation is used to explicitly indicate the correspondence between the action graph and the semantic concept of individual actions. In the present invention, the value of the feature length d is 128.
[0040] Step 3.3: Action graph fusion at the individual action level
[0041] The flow diagram of this step is shown in the appendix Figure 3 . For any given group behavior category, in this step, the corresponding action graph is first flattened into K a vectors of h×w, and an FC layer is used to adjust the length of each vector to D. Then, the K a vectors are concatenated with the semantic representations of the corresponding individual actions, and another FC layer is used to adjust the length of each vector to D. Finally, the K a vectors are stacked and flattened to obtain a vector with a length of K a ×D, and an FC layer is used to fuse them to obtain the group feature corresponding to the given group behavior category of the input video, with the shape of 1×D.
[0042] Step 3.4: Action graph fusion at the group behavior level
[0043] For K g categories of group behaviors, repeat Step 3.3 to obtain K ggroup features corresponding to the given group behavior categories. Stacking K g group features to obtain the group features of the input video, with a shape of K g ×D.
[0044] Step 4: Group behavior classification
[0045] In this step, an FC layer is used as the classification head to classify the group features of the input video, obtaining the predicted scores of K g group behaviors. In addition, the visual features of the input video obtained in Step 2 are downsampled to 1×1×C using average pooling operation, and then input into an FC layer for group behavior classification, obtaining another K g predicted scores. The final classification result is the average of the two sets of predicted scores. The classification of group behaviors is supervised using Equation (2):
[0046]
[0047] where, represents the prediction obtained from the visual features of the input video, is the prediction obtained from the group features of the input video, and g is the ground truth label of the group behavior. λ main is the weight scalar used to balance the two classification losses. In the present invention, λ main takes 3. The above loss function is optimized using the Adaptive Moment Estimation (Adam) algorithm, with an initial learning rate of 5e-4 and trained for 50 epochs.
Claims
1. A weakly supervised group behavior recognition method based on action graphs guided by visual concept knowledge, characterized by: Firstly, a neural network is constructed to learn and extract the visual concept knowledge of individual actions, and statistical information related to group behavior is obtained through statistical methods. Then, the neural network is used to learn and extract the visual features of the video, and the visual concept knowledge is used to calculate the action graph of the video. After that, the statistical information related to group behavior is used to enhance the connection between the action graph and the specific group behavior, and the action graphs are fused in combination with the semantic representation of individual action categories. Finally, the classification head with a fully connected layer structure is used to classify group behavior, so as to realize group behavior recognition under weak supervision conditions.
2. The method according to claim 1, characterized in that: Step 1: Acquire visual concept knowledge and statistical information Step 1.1: Acquire visual concept knowledge Input the training set video samples, build the backbone network to extract the visual features of the sample video, and its shape is h×w×C; where h, w, and C are the height, width, and number of feature channels respectively; then use the RoiAlign operation to crop the visual representation of each individual from the visual features of the sample, and its shape is p×p×C; where p is the height of the representation block (equal to the width); use an FC layer to encode the visual representation of each individual into 1×D, and then input the encoded representation into another FC layer to obtain the predicted value of the action category of each individual; where D is the feature length; at the same time, use the average pooling operation to downsample the visual features of the sample video to 1×1×C, and then input an FC layer to obtain the predicted value of the group behavior category of the sample; use formula (1) to supervise the classification of individual actions and group behaviors: in, represents the cross-entropy loss (Cross-Entropy Loss), and is the predicted result of group behavior and individual action, g and a n is the true value label of group behavior and individual action, N is the number of individuals in the input video, and λ pre is the weight scalar used to balance the two classification losses; After the network training is completed, the average value of each individual visual representation is taken according to the corresponding action category to obtain K a The characteristic prototype of each individual action is in the form of p×p×C, which is used as the vectorized representation of the visual concept knowledge of the individual action; a is the number of action categories; Step 1.2: Get statistics For group behavior scenes with regular behavior areas and individual position distribution, the training set samples are first normalized to eliminate the interference of camera perspective changes on statistical information; the M-LSD algorithm is used to detect the landmarks of the motion area and calculate the intersections between the landmarks; then the three-dimensional affine matrix from the camera perspective of the sample to the top view perspective is calculated based on the coordinates of the intersection points; based on the matrix, the sample is projected from the camera perspective to the top view perspective, and the resolution of the transformed sample is consistent with the original resolution; After view normalization, for all individuals in the training set, the bottom center point and neighborhood of their bounding boxes are marked to K g ×K a In the subgraph, K g represents the number of group behavior categories; the labeling process depends on the individual action category corresponding to each individual and the group behavior category to which it belongs; then, all subgraphs are spliced to obtain a K g ×K a ×h×w behavior-action relationship graph as a vectorized representation of statistical information related to group behavior.
3. The method according to claim 1, characterized in that: Step 2: Compute the action graph of the input video Input video, build backbone network to extract its visual features, the shape is h×w×C; combine the visual features of input video with K a The feature prototypes of individual actions are matched in the frequency domain to obtain the action graph of the input video; first, under a given individual action category, the visual features of the input video and the feature prototype of the individual action are padded with zeros, and the height and width of their resolutions are spatially extended to (h+p-1) and (w+p-1); then, a two-dimensional fast Fourier transform is performed on both, and the complex conjugate of the transformation result of the feature prototype is taken; the result is vector dot multiplication, and then a two-dimensional inverse fast Fourier transform is performed to obtain C feature maps; the feature maps are averaged to obtain an action sub-map with a height of (h+p-1) and a width of (w+p-1) corresponding to the individual action; the corresponding zero-padded part of the sub-map is cropped to make its resolution h×w; K a The subgraphs corresponding to the individual actions are grouped to obtain the action graph of the input video, which is in the form of K a ×h×w.
4. The detection method according to claim 1, characterized in that: Step 3: Action graph enhancement and fusion Step 3.1: Action Graph Enhancement First, broadcast the action graph to K g ×K a ×h×w, corresponding to K g group behavior categories; then perform vector dot multiplication with the behavior-action relationship graph to supplement the spatial distribution information of individual behaviors corresponding to specific group behaviors; then, split the dot multiplication result into K g Group enhanced action graph; Step 3.2: Obtaining semantic representation First, K a The individual action labels are encoded as one-hot vectors; then, they are embedded into a d-dimensional latent space using an FC layer to obtain a K-dimensional latent space. a ×d semantic features of individual action categories; This semantic representation is used to explicitly indicate the correspondence between action graphs and individual action semantic concepts; Step 3.3: Individual action-level action graph fusion For any given group behavior category, first flatten its corresponding action graph into K a h×w vectors, and use an FC layer to adjust the length of each vector to D; then, K a vectors are concatenated with the semantic representation of the corresponding individual action, and another FC layer is used to adjust the length of each vector to D; finally, K a The vectors are stacked and flattened to obtain a vector of length K a ×D vector, and use an FC layer to fuse them to obtain the group feature of the input video corresponding to the given group behavior category, whose shape is 1×D; Step 3.4: Group behavior-level action graph fusion For K g group behavior categories, repeat step 3.3 to get K g group characteristics corresponding to a given group behavior category; Stacking K g The group features of the input video are obtained, and their shape is K g ×D.
5. The method according to claim 1, characterized in that: Step 4: Classification of group behavior Use an FC layer as the classification head to classify the group features of the input video and get K g In addition, the visual features of the input video obtained in step 2 are downsampled to 1×1×C using an average pooling operation, and then input into an FC layer for group behavior classification to obtain another K g prediction scores; the final classification result is the average of the two groups of prediction scores; the classification of group behavior is supervised using formula (2): in, represents the predictions derived from the visual features of the input video, is the prediction derived from the group characteristics of the input video, g is the ground-truth label of the group behavior; main is a weight scalar used to balance the two classification losses.