Fine-grained video behavior recognition method based on attribute guidance
Through a fine-grained video behavior recognition network based on attribute guidance, using atomic attribute generation and semantic embedding methods, the problem of difficult traditional technology to identify fine-grained behaviors is solved, and the accurate identification and understanding of the actions of human bodies and mechanical equipment is achieved.
Patent Information
- Application Number
- CN202510262627.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to effectively identify fine-grained behaviors of complex actions, especially in human activities and mechanical equipment actions, and traditional methods are difficult to capture subtle differences in actions.
A fine-grained video behavior recognition network based on attribute guidance is proposed. Through atomic attribute generation and attribute guidance tip fine-tuning, semantic embeddings are generated and constrained in semantic space to improve the accuracy of action recognition.
The accurate identification of fine-grained movements of human bodies and mechanical equipment is achieved, and the understanding of complex movements is improved. The experimental results show superiority in the robotic arm behavior recognition dataset.
Smart Images

Figure CN120088708A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and image processing, and relates to the use of a deep convolutional neural network for fine-grained behavior recognition of videos. Specifically, it relates to a fine-grained video behavior recognition method based on attribute guidance. Background Art
[0002] Video behavior recognition is of great significance for a wide range of visual analysis applications, such as intelligent monitoring, social scene understanding, and sports video analysis. With the great success of deep learning in various video understanding tasks, many spatio-temporal feature learning behavior recognition methods (such as STM, TSM, and I3D) and long-term modeling methods (such as TCN, TDRN, and SlowFast) have been proposed for video behavior recognition. To improve the network's ability to perceive motion patterns, TEA (Yan Li, Bin Ji, Xintian Shi, Jianguo Zhang, Bin Kang, and Limin Wang. TEA: temporal excitation and aggregation for action recognition. In IEEE CVPR, June, pages 906–915, 2020) uses the temporal differences at the feature level to activate motion-sensitive channels, forming a hierarchical residual structure for behavior recognition. ViViT (Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lucic, and Cordelia Schmid. ViViT: A Video Vision Transformer. In IEEE, ICCV, October, pages 6816--6826, 2021) captures long-range dependencies with an extensible self-attention mechanism, thus outperforming convolutional neural networks.
[0003] Since there are only subtle differences between many human activities in the professional field, behavior recognition has recently evolved to a more fine-grained level. Traditional behavior recognition mainly focuses on identifying visual appearance cues of objects and backgrounds, such as gymnastics and basketball, while fine-grained behavior recognition differentiates action categories at the sub-category level. For example, the action pairs in existing fine-grained datasets based on high-definition sports videos are more sensitive to temporal correlations than objects and backgrounds. The FineGym dataset consists of various subtle aspects, such as postures, movement ranges, and dramatic limb deformations. That is to say, fine-grained behavior recognition is significantly more challenging than traditional behavior recognition.
[0004] Recently, vision-language models have gradually emerged, demonstrating how language can facilitate communication between different models. For example, after pre-training, CLIP (Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger and Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision. In ICML, July, volume 139, pages 8748-8763, 2021) enables natural language to be used to refer to learned visual concepts. X-CLIP (Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang and Haibin Ling. Expanding Language-Image Pretrained Models for General Video Recognition. In ECCV, October, Part IV, volume 13664, pages 1–18, 2022) adds trainable temporal layers to its image encoder and creates video-specific text prompts. In addition, we also noticed that the action category descriptions of these methods are usually extremely abstract and concise, which makes it difficult for the text encoder to capture complex activity semantics. Therefore, our goal is to use atomic attributes, specifically sub-actions decoupled in time, to match the video semantics and thus improve the understanding level of complex actions. Summary of the Invention
[0005] The present invention aims to overcome the deficiencies of traditional methods and technologies, and proposes a fine-grained video action recognition network based on attribute guidance. This network consists of two main modules: atomic attribute generation and attribute-guided prompt fine-tuning. Specifically, given a dual-combination prompt composed of a class label and atomic attributes (obtained through atomic attribute generation), we provide these inputs to a frozen text encoder to generate text embedding representations. Then, we use the video as a conditional input to fine-tune these text embedding representations to generate semantic responses for the class and attributes, thereby supporting action recognition.
[0006] The technical solution of the present invention is as follows:
[0007] A fine-grained video behavior recognition method based on attribute guidance, the method comprising the following steps:
[0008] First step, prepare initial training data: The training data includes video segments obtained by splitting the video and text descriptions of action categories;
[0009] Second step, generation of atomic action attributes: Query the text descriptions of action categories to a large language model (LLMs) to expand the category labels, and then generate corresponding atomic action attributes;
[0010] Third step, generate video tokens, category prompts, and attribute prompts: The video segments obtain video tokens through a video encoder; the action categories and atomic action attributes obtain category prompts and attribute prompts respectively through a text encoder;
[0011] Fourth step, obtain category embeddings and attribute embeddings through two video-conditioned encoders: Use a video-conditioned category encoder to obtain category embeddings, and obtain attribute embeddings through a video-conditioned attribute encoder. These two embeddings represent important relationship transformations for action recognition;
[0012] Fifth step, construct a Joint Semantic Space Constraint (JSC) to align the two video-conditioned embeddings obtained in the fourth step in the latent space;
[0013] Sixth step, construct a loss function and train the network;
[0014] Loss function Measures the gap between the action category labels predicted by the input video in the training data and the true category labels;
[0015]
[0016] Among them, is a cross-entropy loss, used to measure the degree of prediction deviation, and then guide the model training, so that the action prediction result of the model can be closer to the real situation. is the Joint Semantic Space Constraint loss, used to distinguish different actions and normal or abnormal states within each action according to the characteristics of the dataset, and make the positive and negative sample pair embeddings align or separate as required through relevant constraints. λ is a hyperparameter that weighs the importance of the loss function, used to adjust the relative weights of the classification objective loss and the Joint Semantic Space Constraint loss in the overall training framework;
[0017] Seventh step, perform fine-grained video behavior recognition through the trained network.
[0018] The beneficial effects of the present invention are:
[0019] The present invention is a framework for fine-grained video action recognition based on an attribute-guided deep neural network, and finally obtains the classification results of fine-grained actions, with the following characteristics:
[0020] 1. The system is easy to construct and can recognize fine-grained actions of human bodies and mechanical equipment;
[0021] 2. We propose to use atomic attributes as the main clue to better understand complex human actions and robotic arm actions;
[0022] 3. We perform multi-step processing on the features of videos and texts to generate semantic embeddings and further constrain the prediction results in the semantic space;
[0023] 4. The experimental results have advanced superiority in two mainstream robotic arm action recognition datasets. Description of the Drawings
[0024] Figure 1 It is a specific implementation network framework diagram. Specific Embodiments
[0025] The following will make a detailed description of the attribute-guided fine-grained video action recognition method of the present invention in combination with embodiments and drawings.
[0026] An attribute-guided fine-grained video action recognition method, as Figure 1 shown, includes the following steps:
[0027] (1) Prepare initial training data
[0028] The initial training data includes video segments obtained by splitting videos and text descriptions of actions in the segments.
[0029] Sample the processed video segments at a fixed frame rate and then input them into the network through data augmentation. During the data augmentation process, the following two methods are mainly used:
[0030] 1. Geometric transformation methods: These methods enhance the dataset by changing the geometric shape of the image. The specific operations include mirror flipping, random rotation, cropping, irregular deformation, and scaling, etc. These operations can generate diverse image samples without changing the image content.
[0031] 2. Color transformation category: Different from geometric transformation, color transformation category methods directly change the pixel values of images, thus affecting the visual content of images. Common operations include adding noise, applying blur effects, adjusting color balance, randomly erasing or filling partial regions, etc. These methods can effectively simulate different lighting conditions and color distortion situations, further enhancing the diversity of the dataset. The initial text description of the behavior in the video in the training data is a behavior description that can be disassembled into atomic action attributes, called action categories.
[0032] (2) Generation of atomic action attributes
[0033] For each action category, design a corresponding prompt template to query a large language model (specifically ChatGPT-4) to expand the category label, generate a behavior description of atomic action attributes, and then obtain a text description of the behavior in the video containing local context information at different time intervals.
[0034] (3) Generation of video tokens, category prompts, and attribute prompts
[0035] Use the video encoder f(·||θ v ) to encode the video clip v in the training data to generate the video token X v , and at the same time put its corresponding action category C and atomic action attribute A into the text encoder for encoding to obtain the category prompt X c and the attribute prompt X a :
[0036]
[0037] Among them, θ v represents the parameters of the video encoder, represents the parameters of the text encoder.
[0038] (4) Obtain category embeddings and attribute embeddings through two video conditional encoders
[0039] First, perform spatial saliency merging on the video tokens, and apply the density peak clustering algorithm based on K-nearest neighbors DPC-KNN to identify discriminative video tokens and merge non-critical tokens (such as the background). For a video token X v , the spatial token in its t-th frame is denoted as Calculate the local density ρ for each token i and its distance metric δ i :
[0040]
[0041] Among them, represents One of the K nearest spatial markers, where K represents the number of neighboring spatial markers, and ρ j represents the local density of other spatial markers;
[0042] ρ i ×δ i Markers with high values are determined as cluster centers, and the remaining markers are assigned to the nearest center. Finally, the spatial markers of the t-th frame Aggregated spatial markers
[0043] For the video marker X v After merging the spatial saliency, we get Using the video-conditioned attribute encoder and combining with the attribute prompt for video-conditioned attribute encoding, for Performing multi-scale temporal downsampling, we get where 2 L is the downsampling factor, and then we get the video-conditioned attribute embedding
[0044]
[0045] where, MHCA attr (·) represents the multi-head cross-attention operation, and X a represents the attribute prompt obtained in the third step.
[0046] Using the same method, for the result obtained after merging the spatial saliency of the video marker by using the video-conditioned class encoder and combining with the class prompt Performing video-conditioned class encoding, we get the video-conditioned class embedding After that, the final class prediction p cls is obtained by applying average pooling and an action head with a multi-layer perceptron (MLP) structure:
[0047]
[0048] The cross-entropy objective function can be used to optimize the action prediction p cls :
[0049]
[0050] where B is the number of training samples, y is the true action label, represents the class prediction of the b-th training sample.
[0051] (5) Construct the joint semantic space constraint
[0052] Given the characteristics of this dataset, each action category contains instances of abnormal execution, which are slightly different from those of normal execution. Therefore, we propose a joint semantic space constraint method, aiming to effectively distinguish different action categories and further identify the normal and abnormal states in each action category. First, a category prototype library is constructed Among them, represents the normal prototype of the N cls th category, represents the abnormal prototype of the N cls th category; the category prototype library M stores and updates the normal and abnormal prototypes based on the classification results obtained from the formula of action prediction p cls .
[0053] Next, according to the action prediction classification result p cls , select the attributes corresponding to the identified category from , and retrieve a pair of category prototypes. For example, if p cls is classified as the i-th category (a normal category), the normal and abnormal prototypes of this category will be retrieved from the category prototype library M and the attributes Apply a constraint to ensure that the similarity score between the attribute and the normal prototype exceeds that between the attribute and the abnormal prototype by α, as follows:
[0054]
[0055] Among them, B represents the number of training samples, K(l) represents the set of all samples belonging to the same category as the l-th sample, K ∈ K(l) represents a certain sample in this set, represents the category to which the k sample is classified, S(·) represents the cosine similarity, a i represents the central embedding of all attributes of the category where the k sample is located, represents the average pooling of all attributes of the category where the k sample is located, is the normal prototype of the category where the k sample is located, is the abnormal prototype of the category where the k sample is located. The goal of is to encourage the embeddings of the positive pairs to align with each other, while promoting the embeddings of the negative pairs
[0056] (6) Construct the loss function and train the network
[0057] The framework jointly trains the classification objective constraint and the joint semantic space constraint :
[0058]
[0059] Among them, λ is a hyperparameter used to weigh the importance of the loss function.
[0060] In the inference stage, the central embedding a of the attribute i and the class prototype The highest similarity score between them is used as the final recognition result p:
[0061]
[0062] Among them, S(·) represents the cosine similarity, N cls is the total number of classes, and argmax is the operation of selecting the class where the highest score is located.
[0063] Table 1 shows the accuracy comparison of the behavior recognition results of the present invention
[0064]
[0065]
[0066] The behavior recognition results of the present invention and the comparison with other methods are shown in Table 1. The present invention uses the commonly used Averaged per-class accuracy (Mean%) and Averaged per-video accuracy (Top-1%) as the evaluation criteria for the accuracy of behavior recognition. The larger the Mean and Top-1, the higher the accuracy of behavior recognition. The present invention uses two datasets (RobAVA-S, RobAVA-R) for robotic arm action recognition to verify the effectiveness of the proposed attribute-guided fine-grained video behavior recognition method. Specifically, RobAVA-S is a dataset of robotic arm behavior videos collected from a simulated environment with fine annotations, containing 17,000 videos, divided into 3 coarse-grained action categories and 8 fine-grained action categories. RobAVA-R is a dataset of robotic arm behavior videos collected in a real-world scenario, containing 22,573 videos, divided into 9 coarse-grained action categories and 100 fine-grained action categories. In terms of experimental settings, the input videos are processed using a primary sparse sampling and enhancement strategy, with a resolution of 224×224. The method of the present invention is trained on 2 NVIDIA-A40 graphics cards, using the AdamW training optimizer. In the basic training stage, the model is iteratively trained 30 times on the RobAVA-S and RobAVA-R datasets, and the learning rate is set to 8×10 -6 , and the weight decay is set to 0.001.
[0067] The comparison methods include I3D Non-Local-R101 (Xiaolong Wang, Ross B. Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In IEEE, CVPR, June, pages 7794–7803, 2018), TSM (Ji Lin, Chuang Gan, and Song Han. TSM: temporal shift module for efficient video understanding. In IEEE ICCV, October, pages 7082–7092, 2019), TEA (Yan Li, Bin Ji, Xintian Shi, Jianguo Zhang, Bin Kang, and Limin Wang. TEA: temporal excitation and aggregation for action recognition. In IEEE CVPR, June, pages 906–915, 2020), TAM (Zhaoyang Liu, Limin Wang, Wayne Wu, Chen Qian, and Tong Lu. TAM: temporal adaptive module for video recognition. In IEEE / CVF, ICCV, October, pages 13688–13698. IEEE, 2021), ViViT-L (nurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lucic, and Cordelia Schmid. Vivit: A video vision transformer. In IEEE / CVF, ICCV, October, pages 6816–6826, 2021), TimeSformer-L (Gedas Bertasius, Heng Wang, and Lorenzo Torresani.Is space-timeattention all you need for video understanding?In ICML,July,Virtual Event,pages 813–824,2021),VideoSwin-L(Ze Liu,Jia Ning,Yue Cao,Yixuan Wei,ZhengZhang,Stephen Lin,and Han Hu.Video swin transformer.In IEEE / CVF,CVPR,June,pages 3192–3201,2022),MViT-H(hen Yan,XuehanXiong,Anurag Arnab,Zhichao Lu,MiZhang,Chen Sun,and Cordelia Schmid.Multiview transformers for videorecognition.In IEEE / CVF,CVPR,June,pages 3323–3333,2022),UniFormerViT-B / 16(unchang Li,Yali Wang,Peng Gao,Guanglu Song,Yu Liu,Hongsheng Li,and YuQiao.Uniformer:Unified transformer for efficient spatial-temporalrepresentation learning.In ICLR,April.OpenReview.net,2022),UniFormerV2 ViT-L / 14(Kunchang Li,Yali Wang,Yinan He,Yizhuo Li,Yi Wang,Limin Wang,and YuQiao.UniFormerV2:spatiotemporal learning by arming image vits with videouniformer.In IEEE / CVF,ICCV,October,pages 5373–5382,2023),ActionCLIPViT-B / 16(Mengmeng Wang,Jiazheng Xing,and Yong Liu.Action-clip:A new paradigm forvideo action recognition.CoRR,abs / 2109.08472, 2021), X-CLIP ViT-L / 14 (olin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, and Haibin Ling. Expanding language-image pretrained models for general video recognition. In ECCV, October, pages 1–18, 2022), Wu et.al. ViT-L / 14 (Wenhao Wu, Zhun Sun, and Wanli Ouyang. Revisiting classifier: Transferring vision-language models for video recognition. In AAAI, February, pages 2847–2855, 2023).
Claims
1. A fine-grained video behavior recognition method based on attribute guidance, characterized in that: The method comprises the following steps: The first step is to prepare the initial training data: the training data includes the video clips after video segmentation and the text description of the action category; The second step is to generate atomic action attributes: query the text description of the action category to the large language model to expand the category label and generate the corresponding atomic action attributes; Step 3: Generate video tags, category hints, and attribute hints: The video clips are passed through a video encoder to obtain video tags; The action category and atomic action attributes are respectively given category hints and attribute hints by the text encoder; Step 4: Get category embedding and attribute embedding through two video conditional encoders: get category embedding using video conditional category encoder, and get attribute embedding using video conditional attribute encoder; these two embeddings represent the important relation transformations for action recognition; In the fifth step, we construct a joint semantic space constraint to align the two video-conditional embeddings obtained in the fourth step in the latent space. Step 6: Construct the loss function and train the network; Loss Function Measuring the gap between the action category labels predicted by the input video and the true category labels in the training data; in, is a cross entropy loss, which is used to measure the degree of prediction deviation and guide model training so that the action prediction results of the model can be closer to the actual situation; It is a joint semantic space constraint loss, which is used to distinguish different actions and normal or abnormal states within each action according to the characteristics of the data set, and align or separate the embeddings of positive and negative sample pairs as required through relevant constraints; λ is a hyperparameter that weighs the importance of the loss function and is used to adjust the relative weights of the classification target loss and the joint semantic space constraint loss in the overall training framework; The seventh step is to perform fine-grained video behavior recognition through the trained network.
2. The attribute-guided fine-grained video behavior recognition method according to claim 1 is characterized in that: In the first step, the processed video clips are sampled at a fixed frame rate and then input into the network through data enhancement.
3. The attribute-guided fine-grained video behavior recognition method according to claim 1 is characterized in that: In the second step, for each action category, a corresponding prompt template is designed to query the large language model to expand the category label, generate a behavioral description of the atomic action attributes, and then obtain a text description of the behavior in the video containing local context information at different time intervals.
4. The attribute-guided fine-grained video behavior recognition method according to claim 1, characterized in that: In the third step, the video encoder f(·||θ v ) Encode the video clip v in the training data and generate the video tag X v , and put the corresponding action category C and atomic action attribute A into the text encoder Encode and get category prompt X c and attribute hint X a : Among them, θ v Represents the parameters of the video encoder, Represents the parameters of the text encoder.
5. The attribute-guided fine-grained video behavior recognition method according to claim 1, characterized in that: In the fourth step: firstly, spatial saliency merging of video tags is performed, and the density peak clustering algorithm DPC-KNN based on K nearest neighbor is applied to identify the distinguishing video tags and merge the non-critical tags. For a video tag X v , and the spatial mark in the tth frame is recorded as For each mark Calculate the local density ρ i and its distance index δ i : in, express One of the K neighboring space markers, K represents the number of neighboring space markers, ρ j represents the local density of other spatial markers; ρ i ×δ i The marker with the highest value is determined as the cluster center, and the remaining markers are assigned to the nearest center; finally, the spatial marker of the tth frame is obtained. Spatial markers after aggregation Mark the video with X v After combining the spatial saliency, we get The video conditional attribute encoder is used to encode the video conditional attributes in combination with the attribute prompt. Perform multi-scale temporal downsampling to obtain 2 of them L is the downsampling factor, and then the video conditional attribute embedding is obtained Among them, MHCA attr (·) represents multi-head cross attention operation, X a Indicates the attribute hint obtained in the third step; The same method is used to use the video conditional category encoder and combine the category prompts to merge the spatial saliency of the video tag. Perform video conditional category encoding to obtain video conditional category embedding After that, the final category prediction p cls It is through Applying average pooling and an action head with a multi-layer perceptron structure yields: The cross entropy objective function can be used to optimize the action prediction p cls : Where B is the number of training samples, y is the real action label, Represents the category prediction of the b-th training sample.
6. The attribute-guided fine-grained video behavior recognition method according to claim 1, characterized in that: Step 5: First build a category prototype library in, Indicates the Nth cls The normal prototype of the category, Indicates the Nth cls The category prototype library M will predict p based on the action cls The classification results obtained by the formula are used to store and update normal and abnormal prototypes; Next, according to the action prediction classification result p cls ,from Select the attribute corresponding to the identified category from , and retrieve a pair of category prototypes; apply a constraint to ensure that the similarity score between the attribute and the normal prototype exceeds its similarity score with the abnormal prototype by α, as follows: Where B represents the number of training samples, K(L) represents the set of all samples belonging to the same category as the lth sample, k∈K(l) represents a sample in the set, represents the category into which the k samples are classified, S(·) represents the cosine similarity, and a i represents the central embedding of all attributes of the category where the k samples belong, It means that all attributes of the category where the k samples belong are averaged pooled. is the normal prototype of the category where the k sample belongs, is the abnormal prototype of the category to which the k sample belongs; The goal is to encourage The embeddings of The embeddings are far away from each other.
7. The attribute-guided fine-grained video behavior recognition method according to claim 1, characterized in that: The sixth step: In the inference phase, the center of the attribute is embedded in a i With class prototype The highest similarity score between them is taken as the final recognition result p: Among them, S(·) represents cosine similarity, N cls is the total number of categories, and argmax is the operation of selecting the category with the highest score.