Fine-grained video action recognition method based on multi-granularity framework

By using a multi-grained framework and video-text matching technology in fine-grained video action recognition, fine-grained descriptions and fuse video features, the problem of insufficient accuracy in fine-grained video action recognition is solved, and higher recognition accuracy and video content understanding are achieved.

CN119964226AInactive Publication Date: 2025-05-09BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411713132.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In fine-grained video action recognition, how to improve the accuracy of fine-grained video action recognition, especially when dealing with similar appearance and uneven actions.

Method used

A fine-grained video action recognition method based on a multi-grained framework is designed, and fine-grained descriptions are generated through video-text matching, video features are constructed based on coarse and fine-grained text descriptions, and features are fused through cross-attention modules to achieve more accurate video-text matching.

Benefits of technology

It improves the accuracy of fine-grained video action recognition, can more effectively capture atomic actions and global semantic information in the video, and enhances the understanding and recognition ability of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964226A_ABST
    Figure CN119964226A_ABST
Patent Text Reader

Abstract

The invention discloses a fine-grained video action recognition method based on a multi-granularity framework, and belongs to the technical field of computer vision. The implementation method comprises the following steps of: 1, acquiring a fine-grained text by using an LLM (Language Language Model); 2, performing filtering measurement on the fine-grained text to obtain an optimized fine-grained text; 3, acquiring coarse and fine granularity video features; 4, obtaining mixed granularity video features; specifically, fine-grained video features and coarse-grained video features are subjected to feature fusion by adopting FFN feedforward propagation to obtain mixed video features; 5, identifying video actions; compared with the prior art, in fine-grained video action recognition, the accuracy of fine-grained video action recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a fine-grained video action recognition method based on a multi-granularity framework, and belongs to the technical field of computer vision. Background Art

[0002] Fine-grained video action recognition has attracted widespread attention in many practical application fields in recent years, especially in monitoring and video understanding, human-computer interaction, etc. In the field of intelligent monitoring, fine-grained video action recognition can analyze surveillance videos in real time and identify abnormal behaviors or potential threats, thereby improving security and reducing labor costs. In terms of human-computer interaction, through gesture recognition and motion capture, computers can better understand user intentions and achieve a more natural interactive experience.

[0003] However, fine-grained recognition faces challenges, as it needs to process actions with similar appearances and make detailed distinctions, requiring high accuracy when capturing keyframes. For example, the ambiguous actions "baking cookies" and "making pizza" are both actions performed in the kitchen, they have very similar visual appearances, and they have very similar atomic actions such as "kneading dough". In addition, there are some uneven actions, that is, multiple atomic actions are unevenly distributed in the video and some atomic actions may not be directly related to the global video semantics, for example, "swinging legs" contains the action of "standing".

[0004] Therefore, how to improve the accuracy of fine-grained video action recognition becomes an urgent problem to be solved. Summary of the invention

[0005] The purpose of the present invention is to solve the technical problem of improving the accuracy of fine-grained video action recognition and to propose a fine-grained video action recognition method based on a multi-granularity framework.

[0006] The working principle of the present invention is as follows: First, a fine-grained description generation model based on video-text matching is designed to capture the atomic actions in the video by generating fine-grained descriptions, thereby enhancing the understanding of the video content. The model generates detailed descriptions by combining a pre-trained large language model, which can effectively reflect the specific action features in the video. Next, a filtering metric is proposed to select descriptions corresponding to the atomic actions present in the video and description. Then, coarse and fine-grained text descriptions are used to construct coarse and fine-grained video features respectively, and the two are combined to obtain mixed granularity video features. This can achieve more accurate video-text matching.

[0007] The objective of the present invention is achieved through the following technical solutions:

[0008] A fine-grained video action recognition method based on a multi-granularity framework of the present invention is applied to a video action recognition scenario and comprises the following steps:

[0009] Step 1: Use the LLM large language model to obtain fine-grained text;

[0010] Step 1.1: Build an action label set;

[0011] Step 1.2: Use the user input text to decompose the action question text to form a question label P;

[0012] Step 1.3: Input the action label C and question label P into the LLM model to obtain the fine-grained text S n ,n∈(1,N), where N represents the coarse-grained text T being split into N fine-grained texts;

[0013] Step 2: Filter and measure the fine-grained text using the method shown in formula (1) to obtain optimized fine-grained text;

[0014] Step 2.1: Use the text encoder to encode the coarse-grained and fine-grained texts respectively to obtain the coarse-grained text feature t c And fine-grained text features s c .

[0015] Step 2.2: Calculate the similarity scores σ between the fine-grained text features and the coarse-grained text features respectively c and the difference score δ between fine-grained text features c .

[0016] Step 2.3: Use the similarity score σ obtained in step 2.2 c and the difference score δ c Calculate the filtering metrics of fine-grained text;

[0017]

[0018] Among them, σ c Expressed as the similarity score between fine-grained text features and coarse-grained text features; δ c Expressed as a difference score; N represents the coarse-grained text T is split into N fine-grained texts; TPP c Filtering metrics represented as fine-grained text;

[0019] Step 2.4: Use the filtering metric obtained in 2.3 to filter the fine-grained text, that is, select TPP c The text with a larger score is regarded as fine-grained text.

[0020] Step 3: Obtain coarse and fine-grained video features;

[0021] Step 3.1: Acquisition of video frames;

[0022] Step 3.1.1: Perform frame extraction operation on each input video file to obtain L frame images as input images.

[0023] Step 3.1.2: Perform conventional processing on the input image.

[0024] Step 3.2: Perform spatiotemporal encoding on the input image to obtain the video feature v; respectively encode the coarse-grained text T and the optimized fine-grained text S n ,n∈(1,N) to encode the text and obtain the coarse-grained text feature t c and optimizing fine-grained text features c ;

[0025] Step 3.3: Concatenate the coarse-grained text features with the fine-grained text features to obtain enhanced coarse-grained text features;

[0026] Step 3.4: Obtain coarse-grained video features;

[0027] Step 3.4.1: Perform cross-attention calculation on the enhanced coarse-grained text features and video features to obtain the coarse-grained video weight A1;

[0028] Step 3.4.2: Use the coarse-grained video weight A1 to weight the video feature to obtain the coarse-grained video feature O1;

[0029] Step 3.5: Obtain fine-grained video features;

[0030] Step 3.5.1: Use optimized fine-grained text features and video features to perform cross-attention calculations; obtain fine-grained video weight A2;

[0031] Step 3.5.2: Use the fine-grained video weight A2 to weight the video feature to obtain the fine-grained video feature O2;

[0032] Step 4: Obtain mixed granularity video features; specifically, the fine-grained video features and the coarse-grained video features are fused using FFN feedforward propagation to obtain the mixed video features O;

[0033] Step 5: Video action recognition;

[0034] Step 5.1: Match the mixed-granularity video features with the coarse-grained text features to obtain the similarity matching matrix M c ;

[0035] Step 5.2: Select the category with the highest similarity probability in the similarity matching matrix as the prediction result of the video action

[0036] Beneficial effects:

[0037] Compared with the prior art, the present invention has the following effects:

[0038] The present invention designs a sub-text generator based on a large language model. It does not need to construct a data set and conduct training separately, but directly uses the rich prior knowledge of the large language model to split the global semantic information to obtain rich sub-texts.

[0039] The present invention designs a text prompt perplexity measurement index, which can directly compare the quality of sub-texts generated by different prompts of a large language model, and select high-quality sub-texts as extended sub-texts of global semantics.

[0040] The network architecture designed by the present invention first obtains enhanced global semantic information containing sub-text information through the cross-attention module, and then uses the cross-attention module to combine the enhanced global semantic information with video features to obtain coarse-grained video features. In addition, the cross-attention module combines the rich sub-text with the video features to obtain fine-grained video features. The two embeddings are combined together as the final global video features. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 Network architecture diagram;

[0042] Figure 2 Schematic diagram of subtext generator based on large model;

[0043] Figure 3 Schematic diagram of the impact of different numbers N and TTP on recognition accuracy. DETAILED DESCRIPTION

[0044] In order to better illustrate the purpose and advantages of the present invention, the invention is further described below in conjunction with the accompanying drawings and examples. It should be noted that the implementation of the present invention is not limited to the following embodiments, and any form of modification or change made to the present invention will fall within the protection scope of the present invention.

[0045] During the training process, the cosine annealing optimization strategy is used, EPOCH is set to 30, Batch_Size is set to 256, the initial learning rate is set to 5e-5, the input image size is set to 224*224, 8 frames are selected as input for each video, and 8 frames are obtained from the video using uniform sampling. Random shearing and rotation operations are performed on the input images, the dropout ratio is 0, and the optimizer selects Adamw. The network architecture is as follows Figure 1 shown.

[0046] Example

[0047] like Figure 1 As shown, the fine-grained video action recognition method based on a multi-granularity framework of this embodiment is specifically implemented in the following steps:

[0048] Step 1: Use the LLM large language model to obtain fine-grained text;

[0049] Step 1.1: Build an action label set;

[0050] Step 1.2: Use the user input text to decompose the action question text to form a question label P;

[0051] Step 1.3: Input the action label C and question label P into the LLM model to obtain the fine-grained text S n ,n∈(1,N), where N represents the coarse-grained text T being split into N fine-grained texts;

[0052] In the embodiment, the LLM model is a large language model, the action label is the label of "kick football"; the input question label P is: Please help me split "kick football" into four sub-actions; the output fine-grained text is: 1. Adjust position, 2. Swing legs, 3. Extend legs forward and contact with the ball, 4. Persevere to the end and keep balance. Figure 2 Shown are fine-grained text descriptions of the “diving” action generated using two different large language models.

[0053] Step 2: Filter and measure the fine-grained text using the method shown in formula (1) to obtain optimized fine-grained text;

[0054] Step 2.1: Use the text encoder to encode the coarse-grained and fine-grained texts to obtain the coarse-grained text feature t c And fine-grained text features s c .

[0055] Step 2.2: Calculate the similarity scores σ between the fine-grained text features and the coarse-grained text features respectively c and the difference score δ between fine-grained text features c .

[0056] Step 2.3: Use the similarity score σ obtained in step 2.2 c and the difference score δ c Calculate the filtering metrics of fine-grained text;

[0057]

[0058] Among them, σ c Expressed as the similarity score between fine-grained text features and coarse-grained text features; δ c Expressed as a difference score; N represents the coarse-grained text T is split into N fine-grained texts; TPP c Filtering metrics represented as fine-grained text;

[0059] Step 2.4: Use the filtering metric obtained in 2.3 to filter the fine-grained text, that is, select TPP c The text with a larger score is regarded as fine-grained text.

[0060] In the embodiment, Figure 3 As shown, the number N and TPP are fully explored c Impact on the final recognition accuracy, specifically, the impact of different numbers N (left) and TTP (right) on the recognition accuracy.

[0061] Step 3: Obtain coarse and fine-grained video features;

[0062] Step 3.1: Acquisition of video frames;

[0063] Step 3.1.1: Perform frame extraction operation on each input video file to obtain L frame images as input images.

[0064] Step 3.1.2: Perform conventional processing on the input image.

[0065] In the embodiment, conventional processing such as random cropping and rotation operations are performed on the input image.

[0066] Step 3.2: Perform spatiotemporal encoding on the input image to obtain the video feature v; respectively encode the coarse-grained text T and the optimized fine-grained text S n ,n∈(1,N) to encode the text and obtain the coarse-grained text feature t c and optimizing fine-grained text features c ;

[0067] Step 3.3: Concatenate the coarse-grained text features with the fine-grained text features to obtain enhanced coarse-grained text features;

[0068] Step 3.4: Obtain coarse-grained video features;

[0069] Step 3.4.1: Perform cross-attention calculation on the enhanced coarse-grained text features and video features to obtain the coarse-grained video weight A1;

[0070] Step 3.4.2: Use the coarse-grained video weight A1 to weight the video feature to obtain the coarse-grained video feature O1;

[0071] Step 3.5: Obtain fine-grained video features;

[0072] Step 3.5.1: Use optimized fine-grained text features and video features to perform cross-attention calculations; obtain fine-grained video weight A2;

[0073] Step 3.5.2: Use the fine-grained video weight A2 to weight the video feature to obtain the fine-grained video feature O2;

[0074] In the embodiment, we select L=8 as the number of frames extracted from each video, and perform random cropping and rotation operations on each image as video preprocessing. A direct connection method is used when concatenating coarse-grained text features with fine-grained text features.

[0075] Step 4: Obtain mixed granularity video features; specifically, the fine-grained video features and the coarse-grained video features are fused using FFN feedforward propagation to obtain the mixed video features O;

[0076] In the embodiment, each FFN uses a double-layer linear layer, and the video features after FFN are fused by addition to obtain the mixed video feature O.

[0077] Step 5: Video action recognition;

[0078] Step 5.1: Match the mixed-granularity video features with the coarse-grained text features to obtain the similarity matching matrix M c ;

[0079] Step 5.2: Select the category with the highest similarity probability in the similarity matching matrix as the prediction result of the video action

[0080] In the embodiment, by calculating the similarity matrix M of the mixed granularity video feature and the coarse granularity text feature c As the prediction matrix, and for each action select M c The index with the largest weight in the corresponding row is used as the prediction result of the current action

[0081] To further demonstrate the superiority of the present invention, the ablation experiment is first used to prove the effectiveness of the enhanced coarse-grained video features and fine-grained video features proposed in the present invention as shown in Table 1, proving that the enhanced coarse-grained video features and fine-grained video features proposed in the present invention will improve the Top-1 accuracy and Top-5 accuracy.

[0082] Table 1. Coarse and fine-grained video feature ablation experiment table

[0083]

[0084] Table 2 proves that the present invention is not limited to a certain architecture, but can be migrated to other network architectures, proving that the present invention has strong generalization. To verify the effectiveness of the present invention, the present invention is compared with the existing action recognition algorithm under supervised learning as shown in Table 3. In addition, the optimal effect is achieved in zero-sample and small-sample supervised tasks as shown in Tables 4 and 5.

[0085] Table 2 Generalization experiment table

[0086]

[0087] Table 3. Supervised comparison results

[0088]

[0089] Table 4 Zero sample comparison results

[0090]

[0091] Table 5 Small sample comparison results

[0092]

[0093] The specific description above further illustrates the purpose, technical solutions and beneficial effects of the invention in detail. It should be understood that the above is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A fine-grained video action recognition method based on a multi-granularity framework, characterized by: The following steps are included: Step 1: Use the LLM large language model to obtain fine-grained text; Step 1.1: Build an action label set; Step 1.2: Use the user input text to decompose the action question text to form a question label P; Step 1.3: Input the action label C and question label P into the LLM model to obtain the fine-grained text S n ,n∈(1,N), where N represents the coarse-grained text T being split into N fine-grained texts; Step 2: Filter and measure the fine-grained text using the method shown in formula (1) to obtain optimized fine-grained text; Step 2.1: Use the text encoder to encode the coarse-grained and fine-grained texts respectively to obtain the coarse-grained text feature t c And fine-grained text features s c ; Step 2.2: Calculate the similarity scores σ between the fine-grained text features and the coarse-grained text features respectively c and the difference score δ between fine-grained text features c ; Step 2.3: Use the similarity score σ obtained in step 2.2 c and the difference score δ c Calculate the filtering metrics of fine-grained text; Among them, σ c Expressed as the similarity score between fine-grained text features and coarse-grained text features; δ c Expressed as a difference score; N represents the coarse-grained text T is split into N fine-grained texts; TPP c Filtering metrics represented as fine-grained text; Step 2.4: Use the filtering metric obtained in 2.3 to filter the fine-grained text, that is, select TPP c The ones with large scores are regarded as fine-grained texts; Step 3: Obtain coarse and fine-grained video features; Step 3.1: Acquisition of video frames; Step 3.2: Perform spatiotemporal encoding on the input image to obtain the video feature v; respectively encode the coarse-grained text T and the optimized fine-grained text S n ,n∈(1,N) to encode the text and obtain the coarse-grained text feature t c and optimizing fine-grained text features c ; Step 3.3: Concatenate the coarse-grained text features with the fine-grained text features to obtain enhanced coarse-grained text features; Step 3.4: Obtain coarse-grained video features; Step 3.5: Obtain fine-grained video features; Step 4: Obtain mixed granularity video features; specifically, the fine-grained video features and the coarse-grained video features are fused using FFN feedforward propagation to obtain the mixed video features O; Step 5: Video action recognition; Step 5.1: Match the mixed-granularity video features with the coarse-grained text features to obtain the similarity matching matrix M c ; Step 5.2: Select the category with the highest similarity probability in the similarity matching matrix as the prediction result of the video action 2. The fine-grained video action recognition method based on a multi-granularity framework as claimed in claim 1, characterized in that: Step 3.1 is implemented as follows: Step 3.1.1: Perform frame extraction operation on each input video file to obtain L frame images as input images; Step 3.1.2: Perform conventional processing on the input image.

3. The fine-grained video action recognition method based on a multi-granularity framework as claimed in claim 1, characterized in that: Step 3.4 is implemented as follows: Step 3.4.1: Perform cross-attention calculation on the enhanced coarse-grained text features and video features to obtain the coarse-grained video weight A1; Step 3.4.2: Use the coarse-grained video weight A1 to weight the video features to obtain the coarse-grained video features O1.

4. The fine-grained video action recognition method based on a multi-granularity framework as claimed in claim 1, characterized in that: Step 3.5 is implemented as follows: Step 3.5.1: Use optimized fine-grained text features and video features to perform cross-attention calculations; obtain fine-grained video weight A2; Step 3.5.2: Use the fine-grained video weight A2 to weight the video features to obtain the fine-grained video features O2.

Citation Information

Patent Citations

  • Multi-granularity human body action classification method based on graph convolutional network

    CN115116139A

  • Time sequence action detection method and device based on coarse-to-fine granularity information capture

    CN116229315A