A Learnable Low-Rank Bilinear Behavior Perception Method

By adding video space-time modeling and designing a bilinear fusion mechanism based on the image large model, the problem of insufficient video action recognition performance and efficiency in the existing technology is solved, and the effect of efficiently identifying video behavior in a multimodal framework is achieved.

CN120071445BActive Publication Date: 2025-07-01ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510547601.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-01
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing video action recognition methods have shortcomings in performance and efficiency, especially in a multimodal framework. The method of freezing CLIP parameters and training only additional parameters leads to a degradation in supervision performance and the inability to effectively integrate text and visual modalities.

Method used

A method of learning low-rank bilinear behavior perception is proposed. By adding video spatiotemporal modeling based on the image large model, a video behavior recognition model is constructed, including video encoder, multi-scale aggregator, text encoder and multi-task decoder, and a spatiotemporal adapter and multi-head processing mechanism are designed to perform bilinear fusion and feature interaction.

Benefits of technology

It realizes rapid and efficient migration to video behavior recognition tasks based on the image model. By effectively integrating global and local spatiotemporal features, it improves sensitivity to action details and computational efficiency, and achieves advanced performance on multiple action recognition data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071445B_ABST
    Figure CN120071445B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image or video recognition, and particularly to a learnable low-rank bilinear behavior perception method, including: (1) establishing a framework based on an image large model and adding video spatio-temporal modeling for migration to video tasks; (2) constructing a video behavior recognition model within the framework, including: a video encoder, a multi-scale aggregator, a text encoder, and a multi-task decoder; (3) constructing a training mechanism in which the main branch of the large model is frozen and only newly added parameters are learned, and using a server to train the video behavior recognition model. By optimizing the objective function until the network converges, local optimal network parameters are obtained, and a trained video behavior recognition model is obtained; (4) inputting the video sequence to be recognized into the trained video behavior recognition model to recognize human behaviors. The beneficial effects of the present invention are as follows: it can accurately recognize human behaviors in videos and achieves advanced performance in multiple action recognition datasets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image or video recognition, and in particular to a learnable low-rank bilinear behavior perception method. Background Art

[0002] Video action recognition is a key task in video understanding, which involves identifying specific behaviors or actions in a video. As is well known, directly training large models for video action recognition is highly resource-intensive due to large amounts of video data, high computational requirements, and extended training durations. To overcome these difficulties, the research community has developed an Image-to-Video (I2V) adaptation paradigm that converts pre-trained multimodal vision-language models (VLMs) into video processing frameworks, improving performance and efficiency. The most intuitive approach to I2V adaptation is to directly add temporal modeling to the image encoder of CLIP and then fine-tune the entire network. However, full fine-tuning requires high computational costs and may reduce the original generalization ability of CLIP.

[0003] With the emergence of Parameter-Efficient Fine-Tuning (PEFT), researchers have started exploring methods of freezing the original CLIP parameters and only training a minimal set of additional parameters. Early methods mainly focused on unimodal frameworks. By leveraging the visual branch of CLIP, adding learnable parameters such as adapters, and attaching a linear classification layer, these methods have achieved impressive results in supervised settings. However, excluding the text branch in these methods may lead to the loss of the generalization ability of CLIP, which is one of the fundamental advantages of CLIP.

[0004] When applying PEFT techniques to multimodal frameworks, the present invention observes a significant performance gap compared to the performance of unimodal frameworks. Using ST-Adapter as a representative of the unimodal framework and introducing the text branch of CLIP to convert ST-Adapter into a multimodal framework, as expected, the present invention observes a significant decline in supervised performance when freezing the CLIP parameters while learning its adapter. There are three reasons for this performance gap: (1) Existing visual adapters tend to focus on global temporal augmentation while neglecting more subtle temporal differences, which limits their ability to capture the subtle temporal dynamics crucial for video action recognition. (2) Unimodal adapters are insufficient because they cannot effectively integrate the text and visual modalities, resulting in a misalignment between the two. (3) When the text input is limited, the original contrastive learning supervision becomes insufficient, which leads to weaker guidance during the learning process when training on relatively small datasets. Summary of the Invention

[0005] To overcome the above deficiencies, the present invention aims to provide a learnable low-rank bilinear behavior perception method, which can achieve fast and efficient migration to video behavior recognition tasks by adding only a small number of learnable parameters on the basis of an image model.

[0006] The present invention achieves the above object through the following scheme: A learnable low-rank bilinear behavior perception method, comprising the following steps:

[0007] (1) Establish a framework that migrates from an image large model with video spatio-temporal modeling added to a video task;

[0008] (2) Construct a video behavior recognition model within the framework, including:

[0009] (2.1) Establish a video encoder for spatio-temporal modeling and extracting video-level features;

[0010] (2.2) Establish a multi-scale aggregator for aggregating multi-level features and enhancing features in the video encoder;

[0011] (2.3) Establish a text encoder for capturing semantic information and extracting text features;

[0012] (2.4) Establish a multi-task decoder that integrates features from the multi-scale aggregator and the text encoder and performs multi-modal learning;

[0013] (3) Construct a training mechanism in which the main branch of the large model is frozen and only newly added parameters are learned. Use the server to train the video behavior recognition model. By optimizing the objective function until the network converges, obtain the local optimal network parameters to get the trained video behavior recognition model;

[0014] (4) Input the video sequence to be recognized into the trained video behavior recognition model to recognize human behaviors.

[0015] Preferably, the specific steps of step (1) include:

[0016] (1.1) Establish a video feature extraction network based on an image large model. First, extract frames from the given video, usually uniformly sample, extract even frames, and input them into the video feature extraction network at the same time. Perform single-frame feature extraction and spatio-temporal interaction fusion between different frames at the same time. Finally, perform average pooling in the temporal dimension to output video features;

[0017] (1.2) Establish a text feature extraction network based on an image large model. Input the given text into the text feature extraction network to perform semantic feature extraction of lexical units and semantic interaction fusion between different lexical units at the same time. Finally, output text features;

[0018] (1.3) Input the video features and text features into the decoder to obtain and output the classification result, which is the category corresponding to the behavior.

[0019] Preferably, the specific implementation steps of step (2.1) include:

[0020] (2.1.1) Each image large model Transformer layer is composed of L repeated blocks, following the PEFT paradigm, and a spatio-temporal adapter is added to each block;

[0021] (2.1.2) For the spatio-temporal adapter, given the feature , where T is the number of sampled frames, L is the number of tokens, C1 is the number of feature channels, and R is the set of real number matrices with the shape of (T, L, C1);

[0022] (2.1.3) Introduce a learnable bilinear fusion mechanism , which is expressed as: , where is a learnable weight matrix with the shape of (C1, C1), is a learnable projection matrix with the shape of (C1, C), and C is the number of projected feature channels;

[0023] (2.1.4) Introduce multi-head processing , where the symbol […] represents the concatenation operation, H is the number of heads, is the subspace representation of the given feature divided for the h-th head, is a learnable weight matrix with the shape of (C1 / H, C1 / H) for the h-th head, is a learnable projection matrix with the shape of (C1 / H, C / H) for the h-th head, and C / H is the number of projected feature channels;

[0024] (2.1.5) Introduce low-rank latent space decomposition to perform learnable bilinear fusion of the given features A and B in the shared low-rank latent space, and finally project from the latent space to the desired output dimension. The process is expressed as , where are all learnable projection matrices with the shape of (C1 / H, r) for the h-th head, r is the dimension of the projected latent space and , is a learnable projection matrix with the shape of (r², C / H) for the h-th head;

[0025] (2.1.6) Perform deeper feature interaction on the given original input feature Z, the feature obtained by global time enhancement of the spatio-temporal adapter, the feature obtained by local time difference modeling of the spatio-temporal adapter, the spatio-temporal adapter bilinear fusion module , and finally The feature fusion process of (TED-Adapter++) is described as follows:

[0026] , where is the class token, is the video feature token of the t-th frame.

[0027] Preferably, the specific implementation steps of step (2.2) include:

[0028] (2.2.1) For the class tokens output by each layer of the video encoder, connect them to form the key K and the value V, and use the class token output by the last layer of the video encoder as the query Q. The aggregation process is expressed as the following cross-attention CA operation: , where is the aggregated class token;

[0029] (2.2.2) Then add to for final enhancement: , where is the enhanced class token, having the fine-grained information provided by the shallow video encoder layer and the high-level semantics provided by the deep video encoder layer.

[0030] Preferably, the structure of the text encoder in step (2.3) is: each image large model Transformer layer consists of L repeated blocks, following the PEFT paradigm, and a text adapter is added to each block.

[0031] Preferably, in step (2.4), the multi-task decoder is equipped with multiple different learning tasks, and each task corresponds to a separate head, including a visual classification head, a multi-modal contrast learning head, a cross-modal classification head, and a cross-modal masked language modeling head.

[0032] Preferably, the objective function of the multi-task decoder is , where cs represents the cross-modal classification head, CMC represents the multi-modal contrast learning head, CMLM represents the cross-modal masked language modeling head, and VC represents the visual classification head.

[0033] Preferably, in step (3), the main branch of the large model freezes the training mechanism of only learning new parameters, freezes all the parameters of the original image large model, and only learns and trains the parameters of the spatio-temporal adapter in the video encoder, the multi-scale aggregator, the text adapter in the text encoder, and the visual classification head and the cross-modal masked language modeling head in the multi-task decoder.

[0034] Preferably, the specific implementation steps of step (4) include:

[0035] (4.1) First, extract frames from the given video. Extract T frames from each video in a fixed and uniform sampling manner, stack them in the batch dimension B, and finally form a set of inputs with dimensions BT, C, H, W, where C, H, and W are the number of channels, image height, and width respectively;

[0036] (4.2) Use the trained video action recognition model to extract video features from the video;

[0037] (4.3) Use the trained video action recognition model to extract text features from the given text;

[0038] (4.4) Input the video features and text features obtained in steps (4.2) and (4.3) into the decoder, and output the category with the highest selection score as the action category corresponding to the entire video segment.

[0039] Preferably, the even number of frames is 8, 16, or 32 frames.

[0040] The beneficial effects of the present invention are as follows: The present invention effectively integrates global temporal features, local temporal features, and spatial features into an identification model, and maintains the efficiency of the identification model by controlling the number of learnable parameters. By designing a spatio-temporal adapter, a learnable bilinear fusion strategy is allowed to explore the high-order interaction between temporal and spatial features to adapt to action-related regions, capture all-round action category information, and improve the sensitivity to action details. Multiprocessing is used for parallel learning across subspaces, combined with low-rank latent space decomposition to minimize redundant calculations while retaining basic information, thereby optimizing the computational efficiency. In addition, the present invention designs a multi-scale feature aggregation module for cross-level integration of features, effectively combining low-level details with high-level semantics to achieve more comprehensive video understanding. Using this method can accurately identify human behaviors in videos, achieving advanced performance in multiple action recognition datasets, verifying its effectiveness and robustness in dealing with action-related details. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a schematic flowchart of the steps of the method of the present invention;

[0042] Figure 2 is a schematic diagram of the overall network framework of the method of the present invention;

[0043] Figure 3 is a schematic diagram of an integrated block diagram for effectively integrating global temporal features, local temporal features, and spatial features of the present invention; where (a) is a block diagram of the spatio-temporal adapter; (b) is a block diagram of the multi-head low-rank bilinear fusion in the spatio-temporal adapter;

[0044] Figure 4 It is a visualization diagram of the original space, global time, local time, and spatio-temporal fusion attention score of an embodiment of the present invention;

[0045] Figure 5 It is a visualization diagram of the attention score comparing an embodiment of the present invention with the original M²-CLIP. Detailed implementation manners

[0046] The present invention will be further described below in conjunction with specific implementation examples, but the protection scope of the present invention is not limited thereto:

[0047] Embodiment: As Figure 1 , Figure 2 shown, a learnable low-rank bilinear behavior perception method includes the following steps:

[0048] (1) Establish a framework that is based on an image large model and incorporates video spatio-temporal modeling and migrates to video tasks;

[0049] (2) Construct a video behavior recognition model within the framework, including:

[0050] (2.1) Establish a video encoder for spatio-temporal modeling and extracting video-level features;

[0051] (2.2) Establish a multi-scale aggregator for aggregating multi-level features and enhancing features in the video encoder;

[0052] (2.3) Establish a text encoder for capturing semantic information and extracting text features;

[0053] (2.4) Establish a multi-task decoder that integrates features from the multi-scale aggregator and the text encoder and performs multi-modal learning;

[0054] (3) Construct a training mechanism in which the main branch of the large model is frozen and only newly added parameters are learned. Use the server to train the video behavior recognition model. By optimizing the objective function until the network converges, obtain local optimal network parameters, and obtain a trained video behavior recognition model;

[0055] (4) Input the video sequence to be recognized into the trained video behavior recognition model to recognize human behaviors.

[0056] For this method, on the Kinetics-400 validation set, the Top-1 accuracy reaches 87.8%, and on the Something-something v2 validation set, the Top-1 accuracy reaches 73.0%, showing very good recognition effects.

[0057] The following will introduce each step in more detail:

[0058] (1) Establish a framework that is based on an image large model, incorporates video spatio-temporal modeling, and is migrated to video tasks. By migrating the single-frame image large model, while retaining the image model, a spatio-temporal modeling module is added at an appropriate position to obtain a video model that outperforms full training or full fine-tuning and is efficiently migrated from the image model.

[0059] The specific implementation process is as follows:

[0060] (1.1) Establish a video feature extraction network based on the image large model CLIP. First, extract frames from the given video, usually by uniform sampling, extract 8 frames, stack them in the batch dimension, and finally form a set of inputs with dimensions 8B, 3, H, W, where 3, H, W correspond to the number of channels, image height, and width respectively; at the same time, input them into the video feature extraction network to perform single-frame feature extraction and spatio-temporal interaction fusion between different frames, and finally perform average pooling in the temporal dimension to output video features;

[0061] (1.2) Establish a text feature extraction network based on the image large model CLIP. Input the given text into the text feature extraction network to perform semantic feature extraction of lexical units and semantic interaction fusion between different lexical units, and finally output text features;

[0062] (1.3) Input the video features and text features into the decoder to obtain and output the classification result, and the classification result is the category corresponding to the behavior.

[0063] (2) Build a video behavior recognition model within the framework, including:

[0064] (2.1) Establish a video encoder for spatio-temporal modeling and extracting video-level features.

[0065] As Figure 3 shown in (a), perform spatio-temporal modeling of the video on the image features of the core unit within the Transformer layer. Further, as Figure 3 shown in (b), introduce the multi-head and low-rank mechanism to implement the multi-head low-rank bilinear fusion strategy in the spatio-temporal adapter to explore the high-order interaction between time and space features. The specific steps are as follows:

[0066] (2.1.1) Each image large model Transformer layer consists of L repeated blocks, following the PEFT paradigm, and a spatio-temporal adapter is added to each block;

[0067] (2.1.2) For the spatio-temporal adapter, for a given feature , where T is the number of sampled frames, L is the number of tokens, C1 is the number of feature channels, and R is the set of real number matrices with shape (T, L, C1);

[0068] Introduce a learnable bilinear fusion mechanism in (2.1.3) , which is expressed as: , where is a learnable weight matrix with shape (C1, C1), is a learnable projection matrix with shape (C1, C), and C is the number of feature channels after projection;

[0069] For the bilinear fusion mechanism in (2.1.4), to reduce the computational complexity and enrich the feature interaction, introduce multi-head processing, which is specifically expressed as: , where the symbol […] represents the concatenation operation, H is the number of heads, is the subspace representation of the given feature partitioned for the h-th head, is a learnable weight matrix with shape (C1 / H, C1 / H) for the h-th head, is a learnable projection matrix with shape (C1 / H, C / H) for the h-th head, and C / H is the number of feature channels after projection;

[0070] Furthermore, to reduce the computational complexity, for the bilinear fusion mechanism, introduce low-rank latent space decomposition, so that the given features A and B perform bilinear fusion in the shared low-rank latent space, and finally project from the latent space to the desired output dimension. The process is expressed as , where are all learnable projection matrices with shape (C1 / H, r) for the h-th head, r is the dimension of the projected latent space and , is a learnable projection matrix with shape (r², C / H) for the h-th head;

[0071] Finally, to perform deeper feature interaction, given the original input feature Z, the feature obtained by global time enhancement of the spatio-temporal adapter, the feature obtained by local time difference modeling of the spatio-temporal adapter, the bilinear fusion module of the spatio-temporal adapter , the feature fusion process of the final

[0072] , where is the class token, is the video feature token of the t-th frame.

[0073] (2.2) A multi-scale aggregator is established for aggregating multi-level features and enhancing features in video encoders. For each frame t, its representation is enhanced by aggregating multi-level features from the encoder. This method strengthens the per-frame representation by integrating multi-level features, where early layers contribute fine-grained information and deeper layers capture higher-level semantics. Figure 2 As shown, the connection is performed at the end of the video encoder to perform multi-scale aggregation and enhancement. The specific steps are as follows:

[0074] (2.2.1) Class tokens for each layer output of the video encoder , concatenate them to form the key K and value V, and convert the class token output by the last layer of the video encoder into As query Q, the aggregation process is formulated as the following cross-attention (CA) operation: ,in is the class token for aggregation;

[0075] (2.2.2) Then, Add to On top, make the final enhancement: ,in As an enhanced class token, it possesses the fine-grained information provided by the shallow video encoder layers and the high-level semantics provided by the deep video encoder layers.

[0076] (2.3) Build a text encoder for capturing semantic information and extracting text features; excluding the text branch may lead to the loss of CLIP’s generalization ability, which is one of CLIP’s fundamental advantages. Figure 2 As shown in the figure, the text adapter is integrated into the text encoder to strengthen the semantic alignment between visual and text modalities. Its structure is as follows: each image-based model Transformer layer consists of L repeated blocks, following the PEFT paradigm, and a text adapter is added to each block.

[0077] (2.4) Establish a multi-task decoder that integrates features from the multi-scale aggregator and text encoder and performs multimodal learning. Figure 2 As shown in the figure, the multi-task decoder is equipped with multiple different learning tasks, each task corresponds to a separate head, including visual classification head, multimodal contrastive learning head, cross-modal classification head, and cross-modal mask language modeling head, which aims to use multi-task constraints to improve the joint representation ability of the multimodal framework, which helps to mine a richer set of supervisory signals and guide the model to better align visual and textual modalities while capturing various aspects of semantic information, which not only alleviates the performance differences of supervised learning, but also retains the significant generalization ability of CLIP. Its objective function is:

[0078] , where cs, CMC, CMLM, and VC are all different heads in the decoder. cs represents the cross-modal classification head, CMC represents the multi-modal contrastive learning head, CMLM represents the cross-modal masked language modeling head, and VC represents the visual classification head.

[0079] In this embodiment, the server is used to execute the training of the video behavior recognition model. After the video features and text features pass through the decoder, the behavior categories in each video are output, with a size of B, N, that is, each video corresponds to a label vector of size N, where N is the number of categories in the dataset.

[0080] Figure 4 The highlighted light spots in different image sequences represent the original, global, local, and fusion attention scores of this method, indicating how it focuses on the key action areas.

[0081] (3) Construct a training mechanism in which the main branch of the large model is frozen and only the newly added parameters are learned. Use the server to train the video behavior recognition model. By optimizing the objective function until the network converges, obtain the local optimal network parameters to get the trained video behavior recognition model. Freeze all the parameters of the original image large model, and only learn and train the parameters of the spatio-temporal adapter in the video encoder, the multi-scale aggregator, the text adapter in the text encoder, the visual classification head in the multi-task decoder, and the cross-modal masked language modeling head. Train in an end-to-end manner and optimize the objective function: , until the network converges to obtain the local optimal parameters. The visualization of the attention scores compared with the original M²-CLIP is as Figure 5 shown. In two different video action segments, the highlighted parts show how the improved M²-CLIP focuses on different parts of the video content more accurately.

[0082] (4) Input the video sequence to be recognized into the trained video behavior recognition model to recognize human behaviors. The specific steps are as follows:

[0083] (4.1) First, extract frames from the given video. Extract 8 frames from each video in a fixed uniform sampling manner and stack them in the batch dimension B. Finally, form a set of inputs with a dimension of 8B, 3, H, W, where 3, H, and W are the number of channels, the height, and the width of the image, respectively;

[0084] (4.2) Use the trained video behavior recognition model to extract video features from the video;

[0085] (4.3) Use the trained video behavior recognition model to extract text features from the given text;

[0086] (4.4) Input the video features and text features obtained in steps (4.2) and (4.3) into the decoder, and output the category with the highest selection score as the behavior category corresponding to the entire video segment finally.

[0087] The method of the present invention can, under the drive of an image large model, through the introduction of a spatio-temporal adapter and a multi-scale aggregator, relying on a bilinear fusion mechanism, fully fuse global and local spatio-temporal features. At the same time, by introducing multi-head processing and a low-rank space, the feature interaction efficiency is further improved, and the spatio-temporal features are more effectively fused together. Finally, it is possible to achieve a leading behavior recognition effect with a small number of learnable parameters.

[0088] The above are the specific embodiments of the present invention and the technical principles applied. If changes are made according to the concept of the present invention and the functions and effects generated do not exceed the spirit covered by the specification and the drawings, they should still fall within the protection scope of the present invention.

Claims

1. A learnable low-rank bilinear behavior perception method, characterized by The following steps are involved: (1) Establish a framework that uses a large image model to add video spatiotemporal modeling and migrate it to video tasks; (2) Construct a video behavior recognition model within the framework, including: (2.1) Establish a video encoder for spatiotemporal modeling and extracting video-level features; the specific implementation steps include: (2.1.1) Each image-based model Transformer layer consists of L repeated blocks, following the PEFT paradigm, and each block adds a spatiotemporal adapter; (2.1.2) For the spatiotemporal adapter, given the feature , where T is the number of sampling frames, L is the number of tokens, C1 is the number of feature channels, and R is a set of real matrices with shape (T, L, C1); (2.1.3) Introducing a learnable bilinear fusion mechanism , which is expressed as: ,in, is a learnable weight matrix of shape (C1, C1), is a learnable projection matrix of shape (C1, C), where C is the number of feature channels after projection; (2.1.4) Introducing multi-head processing , where the symbol [...] represents the connection operation, H is the number of heads, is the subspace representation of a given feature partitioned for the h-th head, is the learnable weight matrix of the h-th head with shape (C1 / H, C1 / H), is the learnable projection matrix of the h-th head shape (C1 / H, C / H), where C / H is the number of feature channels after projection; (2.1.5) Introducing low-rank latent space decomposition, given features A and B can perform learnable bilinear fusion in a shared low-rank latent space, and finally projected from the latent space to the desired output dimension. The process is expressed as ,in, are all learnable projection matrices with the hth head shape (C1 / H, r), r is the dimension of the latent space after projection and , is the learnable projection matrix of the h-th head with shape (r², C / H); (2.1.6) Given the original input feature Z, the spatiotemporal adapter global time enhancement feature , Features obtained by local time difference modeling of spatiotemporal adapter , Spatiotemporal adapter bilinear fusion module Conduct deeper feature interactions and finally The feature fusion process is described as follows: ,in, is a class token, is the feature token of the video of the tth frame; (2.2) Establish a multi-scale aggregator for aggregating multi-level features and enhanced features in video encoders; (2.3) Build a text encoder to capture semantic information and extract text features; (2.4) Building a multi-task decoder that integrates features from the multi-scale aggregator and text encoder and performs multimodal learning; (3) Construct a training mechanism that freezes the main branch of a large model and only learns new parameters. Use the server to train the video behavior recognition model. By optimizing the objective function until the network converges, the local optimal network parameters are obtained to obtain a trained video behavior recognition model. (4) Input the video sequence to be identified into the trained video behavior recognition model to identify human behavior.

2. A learnable low-rank bilinear behavior perception method according to claim 1, characterized in that: The specific steps of step (1) include: (1.1) Establish a video feature extraction network based on a large image model. First, extract frames from a given video, usually uniformly sample even frames, and input them into the video feature extraction network to extract single-frame features and perform spatiotemporal interaction fusion between different frames. Finally, perform average pooling in the temporal dimension and output video features. (1.2) Establishing a text feature extraction network based on a large image model, inputting a given text into the text feature extraction network, extracting semantic features of vocabulary units and simultaneously performing semantic interaction fusion between different vocabulary units, and finally outputting text features; (1.3) The video features and text features are input into the decoder to obtain and output the classification result, which is the category of the corresponding behavior.

3. The learnable low-rank bilinear behavior perception method according to claim 1, characterized in that: The specific implementation steps of step (2.2) include: (2.2.1) Class tokens for each layer output of the video encoder , concatenate them to form the key K and value V, and convert the class token output by the last layer of the video encoder into As query Q, the aggregation process is formulated as the following cross-attention CA operation: ,in is the class token for aggregation; (2.2.2) Then Add to On top, make the final enhancement: ,in As an enhanced class token, it possesses the fine-grained information provided by the shallow video encoder layers and the high-level semantics provided by the deep video encoder layers.

4. The learnable low-rank bilinear behavior perception method according to claim 1, characterized in that: The structure of the text encoder in step (2.3) is as follows: each image-based model Transformer layer consists of L repeated blocks, following the PEFT paradigm, and a text adapter is added to each block.

5. The learnable low-rank bilinear behavior perception method according to claim 1, characterized in that: In the step (2.4), the multi-task decoder is equipped with multiple different learning tasks, each task corresponds to a separate head, including a visual classification head, a multimodal contrastive learning head, a cross-modal classification head, and a cross-modal mask language modeling head.

6. A learnable low-rank bilinear behavior perception method according to claim 5, characterized in that: The objective function of the multi-task decoder is , where cs represents the cross-modal classification head, CMC represents the multimodal contrastive learning head, CMLM represents the cross-modal masked language modeling head, and VC represents the visual classification head.

7. A learnable low-rank bilinear behavior perception method according to any one of claims 1 to 6, characterized in that: The main branch of the large model in step (3) freezes the training mechanism of only newly added parameter learning, freezes all parameters of the original image large model, and only learns and trains the parameters of the spatiotemporal adapter, multi-scale aggregator in the video encoder, the text adapter in the text encoder, and the visual classification head and cross-modal mask language modeling head in the multi-task decoder.

8. The learnable low-rank bilinear behavior perception method according to claim 1, characterized in that: The specific implementation steps of step (4) include: (4.1) First, extract frames from a given video. T frames are extracted from each video in a fixed uniform sampling manner and stacked in batch dimension B. Finally, a set of inputs is formed with dimensions BT, C, H, W, where C, H, and W are channels, image height, and width respectively. (4.2) Use the trained video behavior recognition model to extract video features from the video; (4.3) Use the trained video behavior recognition model to extract text features from the given text; (4.4) Input the video features and text features obtained in steps (4.2) and (4.3) into the decoder, and output the category with the highest score as the behavior category corresponding to the entire video.

9. A learnable low-rank bilinear behavior perception method according to claim 2, characterized in that The even-numbered frames are 8, 16 or 32 frames.

Citation Information

Patent Citations

  • Clue language recognition method and system based on low-rank bilinear fusion

    CN116206600A

  • Local learnable query enhanced video recognition method

    CN119625836A