Learable low-rank bilinear behavior perception method

By building a video behavior recognition model based on the image large model, introducing spatiotemporal modeling and multi-scale aggregators, the problem of the gap in video action recognition performance and generalization capabilities in the existing technology is solved, and efficient video action recognition and calculation efficiency are achieved.

CN120071445AActive Publication Date: 2025-05-30ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202510547601.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

There is a gap in performance in existing video action recognition methods, especially in a multimodal framework, when CLIP parameters are frozen and only additional parameters are trained, the supervision performance is significantly reduced, and the text and visual modes cannot be effectively integrated, resulting in a decrease in generalization ability.

Method used

A low-rank bilinear behavior perception method is proposed. By introducing space-time modeling based on image large models, a video behavior recognition model is constructed, including video encoder, multi-scale aggregator, text encoder and multi-task decoder, and a training mechanism for freezing the main branch of the big model to add new parameters to learn.

Benefits of technology

Effectively integrating global time features, local time features and spatial features, improving the ability to recognize video actions, enhancing the computing efficiency of the model, and achieving advanced performance in multiple action recognition data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071445A_ABST
    Figure CN120071445A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image or video recognition, in particular to a learnable low-rank bilinear behavior perception method, which comprises the following steps of: (1) establishing a framework for adding video space-time modeling and migrating to a video task on the basis of an image large model; (2) constructing a video behavior recognition model in the framework, wherein the video behavior recognition model comprises a video encoder, a multi-scale aggregator, a text encoder and a multi-task decoder; (3) constructing a training mechanism of large model main branch freezing and only new parameter learning, training the video behavior recognition model by using a server, and obtaining local optimal network parameters by optimizing an objective function until network convergence to obtain a trained video behavior recognition model; and (4) inputting a video sequence to be identified into the trained video behavior identification model to identify human behaviors. The method has the advantages that the human behaviors in the video can be recognized with high precision, and advanced performance is achieved in a plurality of action recognition data sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image or video recognition, and in particular to a learnable low-rank bilinear behavior perception method. Background Art

[0002] Video action recognition is a key task in video understanding, which involves identifying specific behaviors or actions in videos. As is well known, directly training large models for video action recognition is highly resource-intensive due to large amounts of video data, high computational requirements, and extended training durations. To overcome these difficulties, the research community has developed an image-to-video (I2V) adaptation paradigm that converts pre-trained multimodal vision-language models (VLMs) into video processing frameworks, improving performance and efficiency. The most intuitive approach to I2V adaptation is to directly add temporal modeling to the image encoder of CLIP and then fine-tune the entire network. However, full fine-tuning requires high computational costs and may reduce the original generalization ability of CLIP.

[0003] With the emergence of parameter-efficient fine-tuning (PEFT), researchers have begun to explore methods of freezing the original CLIP parameters and only training a minimal set of additional parameters. Early methods mainly focused on unimodal frameworks. By leveraging the visual branch of CLIP, adding learnable parameters such as adapters, and attaching a linear classification layer, these methods have achieved impressive results in supervised settings. However, excluding the text branch in these methods may lead to the loss of the generalization ability of CLIP, which is one of the fundamental advantages of CLIP.

[0004] When applying PEFT techniques to multimodal frameworks, the present invention observes a significant performance gap compared to the performance of unimodal frameworks. Using ST-Adapter as a representative of the unimodal framework and introducing the text branch of CLIP to convert ST-Adapter into a multimodal framework, as expected, the present invention observes a significant decline in supervised performance when freezing the CLIP parameters while learning its adapter. There are three reasons for this performance gap: (1) Existing visual adapters tend to focus on global temporal augmentation while ignoring finer temporal differences, which limits their ability to capture the subtle temporal dynamics crucial for video action recognition. (2) Unimodal adapters are insufficient because they cannot effectively integrate the text and visual modalities, resulting in a misalignment between the two. (3) When the text input is limited, the original contrastive learning supervision becomes insufficient, which leads to weaker guidance during the learning process when training on relatively small datasets. Summary of the Invention

[0005] To overcome the above deficiencies, the present invention aims to provide a learnable low-rank bilinear behavior perception method, which can achieve fast and efficient migration to video behavior recognition tasks by adding only a small number of learnable parameters on the basis of an image model.

[0006] The present invention achieves the above object through the following scheme: A learnable low-rank bilinear behavior perception method, comprising the following steps: (1) Establish a framework that migrates from an image large model by adding video spatio-temporal modeling to a video task; (2) Construct a video behavior recognition model within the framework, including: (2.1) Establish a video encoder for spatio-temporal modeling and extracting video-level features; (2.2) Establish a multi-scale aggregator for aggregating multi-level features and enhancing features in the video encoder; (2.3) Establish a text encoder for capturing semantic information and extracting text features; (2.4) Establish a multi-task decoder for integrating features from the multi-scale aggregator and the text encoder and performing multi-modal learning; (3) Construct a training mechanism in which the main branch of the large model is frozen and only newly added parameters are learned. Use the server to train the video behavior recognition model. By optimizing the objective function until the network converges, obtain local optimal network parameters to get a trained video behavior recognition model; (4) Input the video sequence to be recognized into the trained video behavior recognition model to recognize human behaviors.

[0007] Preferably, the specific steps of step (1) include: (1.1) Establish a video feature extraction network based on an image large model. First, extract frames from the given video, usually uniformly sample, extract even frames, and input them into the video feature extraction network at the same time to perform single-frame feature extraction and spatio-temporal interaction fusion between different frames. Finally, perform average pooling in the temporal dimension to output video features; (1.2) Establish a text feature extraction network based on an image large model. Input the given text into the text feature extraction network to perform semantic feature extraction of lexical units and semantic interaction fusion between different lexical units, and finally output text features; (1.3) Input the video features and text features into the decoder to obtain and output the classification result, and the classification result is the category of the corresponding behavior.

[0008] Preferably, the specific implementation steps of step (2.1) include: (2.1.1) Each image large model Transformer layer consists of L repeated blocks, following the PEFT paradigm, and a spatio-temporal adapter is added to each block; (2.1.2) For the spatio-temporal adapter, given the feature , where T is the number of sampled frames, L is the number of tokens, C1 is the number of feature channels, and R is the set of real number matrices of shape (T, L, C1); (2.1.3) Introduce a learnable bilinear fusion mechanism , which is expressed as: , where is a learnable weight matrix of shape (C1, C1), is a learnable projection matrix of shape (C1, C), and C is the number of projected feature channels; (2.1.4) Introduce multi-head processing , where the symbol […] represents the concatenation operation, H is the number of heads, is the subspace representation of the given feature partitioned for the h-th head, is a learnable weight matrix of shape (C1 / H, C1 / H) for the h-th head, is a learnable projection matrix of shape (C1 / H, C / H) for the h-th head, and C / H is the number of projected feature channels; (2.1.5) Introduce low-rank latent space decomposition to perform learnable bilinear fusion of the given features A and B in a shared low-rank latent space, and finally project from the latent space to the desired output dimension. The process is expressed as , where are all learnable projection matrices of shape (C1 / H, r) for the h-th head, r is the dimension of the projected latent space and , is a learnable projection matrix of shape (r², C / H) for the h-th head; (2.1.6) Perform deeper feature interaction on the given original input feature Z, the feature obtained by global time enhancement of the spatio-temporal adapter, the feature obtained by local time difference modeling of the spatio-temporal adapter, the spatio-temporal adapter bilinear fusion module , and finally The feature fusion process of (TED-Adapter++) is expressed as follows: , where is the class token, is the video feature token of the t-th frame.

[0009] Preferably, the specific implementation steps of step (2.2) include: (2.2.1) For the class tokens output by each layer of the video encoder , concatenate them to form the key K and value V, and use the class tokens output by the last layer of the video encoder as the query Q. The aggregation process is expressed as the following cross-attention CA operation: , where is the aggregated class token; (2.2.2) Then add to for final enhancement: , where is the enhanced class token, which has the fine-grained information provided by the shallow video encoder layer and the high-level semantics provided by the deep video encoder layer.

[0010] Preferably, the structure of the text encoder in step (2.3) is: each image large model Transformer layer consists of L repeated blocks, following the PEFT paradigm, and a text adapter is added to each block.

[0011] Preferably, in step (2.4), the multi-task decoder is equipped with multiple different learning tasks, and each task corresponds to a separate head, including a visual classification head, a multi-modal contrast learning head, a cross-modal classification head, and a cross-modal masked language modeling head.

[0012] Preferably, the objective function of the multi-task decoder is , where cs represents the cross-modal classification head, CMC represents the multi-modal contrast learning head, CMLM represents the cross-modal masked language modeling head, and VC represents the visual classification head.

[0013] Preferably, in step (3), the training mechanism for the large model main branch freezes only the training of newly added parameters. All parameters of the original image large model are frozen, and only the parameters of the spatio-temporal adapter in the video encoder, the multi-scale aggregator, the text adapter in the text encoder, and the visual classification head and cross-modal masked language modeling head in the multi-task decoder are trained.

[0014] Preferably, the specific implementation steps of step (4) include: (4.1) First, extract frames from the given video. Extract T frames from each video in a fixed uniform sampling manner, stack them in the batch dimension B, and finally form a set of inputs with dimensions BT, C, H, W, where C, H, and W are the number of channels, image height, and width respectively; (4.2) Use the trained video action recognition model to extract video features from the video; (4.3) Use the trained video action recognition model to extract text features from the given text; (4.4) Input the video features and text features obtained in steps (4.2) and (4.3) into the decoder, and output the category with the highest selection score as the behavior category corresponding to the entire video segment finally.

[0015] Preferably, the even frames are 8, 16 or 32 frames.

[0016] The beneficial effects of the present invention are as follows: The present invention effectively integrates global temporal features, local temporal features, and spatial features into an identification model, and maintains the efficiency of the identification model by controlling the number of learnable parameters. By designing a spatio-temporal adapter, a learnable bilinear fusion strategy is allowed to explore the high-order interaction between temporal and spatial features to adapt to action-related regions, capture all-round action category information, and improve the sensitivity to action details. Multi-head processing is used for parallel learning across subspaces, and combined with low-rank latent space decomposition to minimize redundant calculations while retaining basic information, thereby optimizing the computational efficiency. In addition, the present invention designs a multi-scale feature aggregation module for cross-level integration of features, effectively combining low-level details with high-level semantics to achieve more comprehensive video understanding. Using this method can accurately identify human behaviors in videos, achieving advanced performance in multiple action recognition datasets, verifying its effectiveness and robustness in dealing with action-related details. Description of the Drawings

[0017] Figure 1 is a schematic diagram of the step flow of the method of the present invention; Figure 2 is a schematic diagram of the overall network framework of the method of the present invention; Figure 3 is a schematic diagram of an integrated block diagram for effectively integrating global temporal features, local temporal features, and spatial features of the present invention; wherein (a) is a block diagram of the spatio-temporal adapter; (b) is a block diagram of the multi-head low-rank bilinear fusion in the spatio-temporal adapter; Figure 4 is a visualization diagram of the original space, global time, local time, and spatio-temporal fusion attention scores of an embodiment of the present invention; Figure 5 is a visualization diagram of the attention scores of an embodiment of the present invention compared with the original M²-CLIP. Detailed Embodiments

[0018] The present invention will be further described below in conjunction with specific implementation examples, but the protection scope of the present invention is not limited thereto: Embodiment: As Figure 1 , Figure 2 shown, a learnable low-rank bilinear behavior perception method includes the following steps: (1)Build a framework that is based on an image large model, incorporates video spatio-temporal modeling, and migrates to video tasks; (2)Construct a video action recognition model within the framework, including: (2.1)Build a video encoder for spatio-temporal modeling and extracting video-level features; (2.2)Build a multi-scale aggregator for aggregating multi-level features and enhancing features in the video encoder; (2.3)Build a text encoder for capturing semantic information and extracting text features; (2.4)Build a multi-task decoder that integrates features from the multi-scale aggregator and the text encoder and performs multi-modal learning; (3)Construct a training mechanism in which the main branch of the large model is frozen and only new parameters are learned. Use the server to train the video action recognition model. By optimizing the objective function until the network converges, obtain the local optimal network parameters and get the trained video action recognition model; (4)Input the video sequence to be recognized into the trained video action recognition model to recognize human actions.

[0019] In this method, on the Kinetics-400 validation set, the Top-1 accuracy reaches 87.8%, and on the Something-something v2 validation set, the Top-1 accuracy reaches 73.0%, showing very good recognition effects.

[0020] The following is a more detailed description of each step: (1)Build a framework that is based on an image large model, incorporates video spatio-temporal modeling, and migrates to video tasks. By migrating the single-frame image large model, while retaining the image model, add a spatio-temporal modeling module at an appropriate position to obtain a video model with performance superior to full training or full fine-tuning and efficiently migrated from the image model.

[0021] The specific implementation process is as follows: (1.1)Build a video feature extraction network based on the image large model CLIP. First, extract frames from the given video, usually uniformly sample 8 frames, stack them on the batch dimension, and finally form a set of inputs with dimensions 8B, 3, H, W, where 3, H, and W correspond to the number of channels, image height, and width respectively. At the same time, input them into the video feature extraction network to perform single-frame feature extraction and spatio-temporal interaction fusion between different frames, and finally perform average pooling in the temporal dimension to output video features; (1.2) Establish a text feature extraction network based on the image large model CLIP. Input the given text into the text feature extraction network to extract the semantic features of lexical units and simultaneously perform semantic interaction and fusion between different lexical units, and finally output the text features; (1.3) Input the video features and text features into the decoder to obtain and output the classification result, and the classification result is the category of the corresponding behavior.

[0022] (2) Build a video behavior recognition model within the framework, including: (2.1) Establish a video encoder for spatio-temporal modeling and extracting video-level features.

[0023] As Figure 3 shown in (a), perform spatio-temporal modeling of the video by the image features of the core units within the Transformer layer. Further, as Figure 3 shown in (b), introduce the multi-head and low-rank mechanism to implement the multi-head low-rank bilinear fusion strategy in the spatio-temporal adapter to explore the high-order interaction between temporal and spatial features. The specific steps are as follows: (2.1.1) Each image large model Transformer layer consists of L repeated blocks, following the PEFT paradigm, and a spatio-temporal adapter is added to each block; (2.1.2) For the spatio-temporal adapter, for a given feature , where T is the number of sampled frames, L is the number of tokens, C1 is the number of feature channels, and R is the set of real number matrices with the shape of (T, L, C1); (2.1.3) Introduce a learnable bilinear fusion mechanism , which is expressed as: , where is a learnable weight matrix with the shape of (C1, C1), is a learnable projection matrix with the shape of (C1, C), and C is the number of projected feature channels; (2.1.4) For the bilinear fusion mechanism, to reduce the computational complexity and enrich the feature interaction, introduce multi-head processing, which is specifically expressed as: , where the symbol […] represents the concatenation operation, H is the number of heads, is the subspace representation of the given feature partitioned for the h-th head, is a learnable weight matrix with the shape of (C1 / H, C1 / H) for the h-th head, is a learnable projection matrix with the shape of (C1 / H, C / H) for the h-th head, and C / H is the number of projected feature channels; (2.1.5) Further, to reduce the computational complexity, for the bilinear fusion mechanism, low-rank latent space decomposition is introduced, enabling the bilinear fusion of given features A and B within a shared low-rank latent space, and finally projecting from the latent space to the desired output dimension. The process is expressed as , where are all learnable projection matrices of shape (C1 / H, r) for the h-th head, r is the dimension of the projected latent space and , is a learnable projection matrix of shape (r², C / H) for the h-th head; (2.1.6) Finally, for deeper feature interaction, given the original input feature Z, the feature obtained by global temporal enhancement of the spatio-temporal adapter, the feature obtained by local temporal difference modeling of the spatio-temporal adapter, the spatio-temporal adapter bilinear fusion module , finally (TED-Adapter++)'s feature fusion process is expressed as follows: , where is the class token, is the video feature token for the t-th frame.

[0024] (2.2) Build a multi-scale aggregator for aggregating multi-level features and enhanced features in the video encoder. For each frame t, enhance its representation by aggregating multi-level features from the encoder. This method strengthens the representation of each frame by integrating multi-level features, where the early layers contribute fine-grained information and the deeper layers capture higher-level semantics. As Figure 2 shown, multi-scale aggregation and enhancement are performed at the end of the video encoder. The specific steps are as follows: (2.2.1) For the class tokens output by each layer of the video encoder, connect them to form the key K and value V, and use the class token output by the last layer of the video encoder as the query Q. The aggregation process is expressed as the following cross-attention (CA) operation: , where is the aggregated class token; (2.2.2) Then, add to for the final enhancement: , where is the enhanced class token, possessing the fine-grained information provided by the shallow video encoder layer and the high-level semantics provided by the deep video encoder layer.

[0025] (2.3) Build a text encoder for capturing semantic information and extracting text features; excluding the text branch may lead to the loss of CLIP’s generalization ability, which is one of CLIP’s fundamental advantages. Figure 2 As shown in the figure, the text adapter is integrated into the text encoder to strengthen the semantic alignment between visual and text modalities. Its structure is as follows: each image-based model Transformer layer consists of L repeated blocks, following the PEFT paradigm, and a text adapter is added to each block.

[0026] (2.4) Establish a multi-task decoder that integrates features from the multi-scale aggregator and text encoder and performs multimodal learning. Figure 2 As shown in the figure, the multi-task decoder is equipped with multiple different learning tasks, each task corresponding to a separate head, including visual classification head, multimodal contrastive learning head, cross-modal classification head, and cross-modal mask language modeling head, which aims to use multi-task constraints to improve the joint representation ability of the multimodal framework, which helps to mine a richer set of supervisory signals and guide the model to better align visual and textual modalities while capturing various aspects of semantic information, which not only alleviates the performance differences of supervised learning, but also retains the significant generalization ability of CLIP. Its objective function is: , where cs, CMC, CMLM, and VC are all different heads in the decoder, cs represents the cross-modal classification head, CMC represents the multi-modal contrastive learning head, CMLM represents the cross-modal masked language modeling head, and VC represents the visual classification head.

[0027] This embodiment uses a server to perform training of a video behavior recognition model. After the video features and text features pass through a decoder, the behavior category in each video is output, with a size of B, N. That is, each video corresponds to a label vector of size N, where N is the number of categories corresponding to the data set.

[0028] Figure 4 The highlighted spots in different image sequences represent the original, global, local, and fused attention scores of our method, showing how it focuses on key action areas.

[0029] (3) Construct a training mechanism that freezes the main branch of the large model and only learns new parameters. Use the server to train the video behavior recognition model. By optimizing the objective function until the network converges, obtain the local optimal network parameters, and obtain a trained video behavior recognition model. Freeze all parameters of the original image large model, and only learn and train the parameters of the spatiotemporal adapter, multi-scale aggregator in the video encoder, text adapter in the text encoder, and visual classification head and cross-modal mask language modeling head in the multi-task decoder. Use end-to-end training to optimize the objective function: , until the network converges and obtains the locally optimal parameters. The visualization of the attention scores compared with the original M²-CLIP is as Figure 5 shown. In two different video action segments, the highlighted parts show how the improved M²-CLIP pays more accurate attention to different parts of the video content.

[0030] (4) Input the video sequence to be recognized into the trained video behavior recognition model to identify human behaviors. The specific steps are as follows: (4.1) First, extract frames from the given video. Extract 8 frames from each video in a fixed uniform sampling manner, stack them on the batch dimension B, and finally form a set of inputs with dimensions 8B, 3, H, W, where 3, H, and W are the number of channels, image height, and width, respectively; (4.2) Use the trained video behavior recognition model to extract video features from the video; (4.3) Use the trained video behavior recognition model to extract text features from the given text; (4.4) Input the video features and text features obtained in steps (4.2) and (4.3) into the decoder, and output the category with the highest selection score as the behavior category corresponding to the entire video segment.

[0031] The method of the present invention can, under the drive of the image large model, through the introduction of spatio-temporal adapters and multi-scale aggregators, rely on the bilinear fusion mechanism to fully fuse global and local spatio-temporal features. At the same time, by introducing multi-head processing and low-rank spaces, the feature interaction efficiency is further improved, and the spatio-temporal features are more effectively fused together. Finally, it can achieve leading behavior recognition effects with a small number of learnable parameters.

[0032] The above are the specific embodiments of the present invention and the technical principles applied. If changes are made according to the concept of the present invention and the functions and effects generated do not exceed the spirit covered by the description and drawings, they should still fall within the protection scope of the present invention.

Claims

1. A learnable low-rank bilinear behavior perception method, characterized by The following steps are involved: (1) Establish a framework that uses a large image model to add video spatiotemporal modeling and migrate it to video tasks; (2) Construct a video behavior recognition model within the framework, including: (2.1) Building a video encoder for spatiotemporal modeling and extracting video-level features; (2.2) Establish a multi-scale aggregator for aggregating multi-level features and enhanced features in video encoders; (2.3) Build a text encoder to capture semantic information and extract text features; (2.4) Building a multi-task decoder that integrates features from the multi-scale aggregator and text encoder and performs multimodal learning; (3) Construct a training mechanism that freezes the main branch of a large model and only learns new parameters. Use the server to train the video behavior recognition model. By optimizing the objective function until the network converges, the local optimal network parameters are obtained to obtain a trained video behavior recognition model. (4) Input the video sequence to be identified into the trained video behavior recognition model to identify human behavior.

2. A learnable low-rank bilinear behavior perception method according to claim 1, characterized in that: The specific steps of step (1) include: (1.1) Establish a video feature extraction network based on a large image model. First, extract frames from a given video, usually uniformly sample even frames, and input them into the video feature extraction network to extract single-frame features and perform spatiotemporal interaction fusion between different frames. Finally, perform average pooling in the temporal dimension and output video features. (1.2) Establishing a text feature extraction network based on a large image model, inputting a given text into the text feature extraction network, extracting semantic features of vocabulary units and simultaneously performing semantic interaction fusion between different vocabulary units, and finally outputting text features; (1.3) The video features and text features are input into the decoder to obtain and output the classification result, which is the category of the corresponding behavior.

3. The learnable low-rank bilinear behavior perception method according to claim 1, characterized in that: The specific implementation steps of step (2.1) include: (2.1.1) Each image-based model Transformer layer consists of L repeated blocks, following the PEFT paradigm, and each block adds a spatiotemporal adapter; (2.1.2) For the spatiotemporal adapter, given the feature , where T is the number of sampling frames, L is the number of tokens, C1 is the number of feature channels, and R is a set of real matrices with shape (T, L, C1); (2.1.3) Introducing a learnable bilinear fusion mechanism , which is expressed as: ,in, is a learnable weight matrix of shape (C1, C1), is a learnable projection matrix of shape (C1, C), where C is the number of feature channels after projection; (2.1.4) Introducing multi-head processing , where the symbol [...] represents the connection operation, H is the number of heads, is the subspace representation of a given feature partitioned for the h-th head, is the learnable weight matrix of the h-th head with shape (C1 / H, C1 / H), is the learnable projection matrix of the h-th head shape (C1 / H, C / H), where C / H is the number of feature channels after projection; (2.1.5) Introducing low-rank latent space decomposition, given features A and B can perform learnable bilinear fusion in a shared low-rank latent space, and finally projected from the latent space to the desired output dimension. The process is expressed as ,in, are all learnable projection matrices with the hth head shape (C1 / H, r), r is the dimension of the latent space after projection and , is the learnable projection matrix of the h-th head with shape (r², C / H); (2.1.6) Given the original input feature Z, the spatiotemporal adapter global time enhancement feature , Features obtained by local time difference modeling of spatiotemporal adapter , Spatiotemporal adapter bilinear fusion module Conduct deeper feature interactions and finally The feature fusion process is described as follows: ,in, is a class token, is the feature token of the video in the tth frame.

4. The learnable low-rank bilinear behavior perception method according to claim 1, characterized in that: The specific implementation steps of step (2.2) include: (2.2.1) Class tokens for each layer output of the video encoder , concatenate them to form the key K and value V, and convert the class token output by the last layer of the video encoder into As query Q, the aggregation process is formulated as the following cross-attention CA operation: ,in is the class token for aggregation; (2.2.2) Then Add to On top, make the final enhancement: ,in As an enhanced class token, it possesses the fine-grained information provided by the shallow video encoder layers and the high-level semantics provided by the deep video encoder layers.

5. The learnable low-rank bilinear behavior perception method according to claim 1, characterized in that: The structure of the text encoder in step (2.3) is as follows: each image-based model Transformer layer consists of L repeated blocks, following the PEFT paradigm, and a text adapter is added to each block.

6. The learnable low-rank bilinear behavior perception method according to claim 1, characterized in that: In the step (2.4), the multi-task decoder is equipped with multiple different learning tasks, each task corresponds to a separate head, including a visual classification head, a multimodal contrastive learning head, a cross-modal classification head, and a cross-modal mask language modeling head.

7. The learnable low-rank bilinear behavior perception method according to claim 6, characterized in that: The objective function of the multi-task decoder is , where cs represents the cross-modal classification head, CMC represents the multimodal contrastive learning head, CMLM represents the cross-modal masked language modeling head, and VC represents the visual classification head.

8. A learnable low-rank bilinear behavior perception method according to any one of claims 1 to 7, characterized in that: The main branch of the large model in step (3) freezes the training mechanism of only newly added parameter learning, freezes all parameters of the original image large model, and only learns and trains the parameters of the spatiotemporal adapter, multi-scale aggregator in the video encoder, the text adapter in the text encoder, and the visual classification head and cross-modal mask language modeling head in the multi-task decoder.

9. The learnable low-rank bilinear behavior perception method according to claim 1, characterized in that: The specific implementation steps of step (4) include: (4.1) First, extract frames from a given video. T frames are extracted from each video in a fixed uniform sampling manner and stacked in batch dimension B. Finally, a set of inputs is formed with dimensions BT, C, H, W, where C, H, and W are channels, image height, and width respectively. (4.2) Use the trained video behavior recognition model to extract video features from the video; (4.3) Use the trained video behavior recognition model to extract text features from the given text; (4.4) Input the video features and text features obtained in steps (4.2) and (4.3) into the decoder, and output the category with the highest score as the behavior category corresponding to the entire video.

10. A learnable low-rank bilinear behavior perception method according to claim 2, characterized in that The even-numbered frames are 8, 16 or 32 frames.

Citation Information

Patent Citations

  • Clue language recognition method and system based on low-rank bilinear fusion

    CN116206600A

  • Fine-grained action recognition method based on cross-modal knowledge alignment

    CN118196888A

  • Multi-modal feature learning efficiency optimization method based on low-rank factorization

    CN118568658A

  • Image large model driven video behavior identification space-time parameter efficient fine tuning method

    CN118865498A

  • Local learnable query enhanced video recognition method

    CN119625836A

Cited By

  • Transform-based high-power face video super-resolution processing method

    CN120765460A

  • High-magnification face video super-resolution processing method based on transformer

    CN120765460B