A behavior recognition method based on multi-modal large model knowledge distillation
By combining a cross-modal distillation framework with the teacher model VTR-B/16 and a hierarchical video-text interaction module, the problems of high model complexity and insufficient feature representation in video domain behavior recognition are solved, achieving efficient, accurate and robust behavior recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI INST OF MICROSYSTEM & INFORMATION TECH CHINESE ACAD OF SCI
- Filing Date
- 2026-02-25
- Publication Date
- 2026-06-05
Smart Images

Figure CN122156854A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of human-computer interaction, specifically relating to a behavior recognition method based on multimodal large model knowledge distillation. Background Technology
[0002] In today's rapidly developing technological landscape, human-computer interaction has become a crucial interdisciplinary field in computer science and human communication. Thanks to the widespread application of smart devices and advancements in virtual reality technology, motion recognition technology, as one of the core technologies of human-computer interaction, has attracted extensive attention from researchers. Motion recognition technology enables computers to understand users' intentions and needs by analyzing their body movements, thereby providing a more natural, intuitive, and efficient interaction method.
[0003] Recently, visual language pre-trained models such as CLIP [1] CLIP has attracted increasing attention due to its impressive multimodal general visual knowledge acquired through pre-training on large-scale network datasets, and its related models possess superior zero-shot and few-shot inference capabilities. Recent work has demonstrated the effectiveness of CLIP in transferring to video action recognition models. [2,3,4] Compared to methods pre-trained on datasets like ImageNet, this approach has significant advantages. However, transfer learning relies heavily on large model CLIPs. [2] CLIP itself is computationally very large, which greatly increases the hardware requirements of its downstream task models.
[0004] Knowledge distillation (KD), as one of the core techniques for model compression and transfer learning, has significantly promoted the development of lightweight models. Teacher models, such as the VTR-B / 16 model, exemplify this advancement.
[15] Not only does it possess excellent temporal modeling capabilities, but it can also leverage intermediate layer features to enhance cross-branch information exchange between video and text modalities, making it suitable for model compression and transfer learning.
[0005] Currently, in the field of Vision Transformer (ViT), Tiny-ViT... [5] The first attempt was made to transfer knowledge from the teacher model using a knowledge distillation mechanism, but it relied entirely on soft labels generated by the teacher while ignoring real labels and intermediate layer features, resulting in limited accuracy of the student model in real-world scenarios. To address this issue, Wang et al. [6] An improved strategy combining attention probes and teacher intermediate layer features is proposed to enhance student input using field data, but it still faces performance degradation due to cross-entropy loss. Future research could include coadvice. [7]PromptKD [8] and Touvron et al. [9] By introducing distillation tokens and real label tokens into the input and incorporating the cross-entropy constraint of the real labels, the robustness of knowledge transfer is effectively improved.
[0006] With the image-text alignment model CLIP [1] Following this breakthrough, researchers began exploring CLIP-based bimodal knowledge distillation methods. TinyClip
[10] ClipKD proposes two core technologies: affinity mimicry and weight inheritance. The former optimizes the knowledge transfer path by simulating intermodal interactions, while the latter accelerates student model training through parameter inheritance. However, its strict structural consistency requirements limit the flexibility of model adaptation.
[11] Building upon this foundation, a multi-dimensional distillation strategy involving relationships, features, gradients, and contrast paradigms is proposed and applied to the lightweight model MobileViT.
[12] The joint experiment verified the effectiveness of cross-modal knowledge transfer. In the same year, Apple's MobileClip...
[13] We innovatively utilize image description models and strong CLIP encoder sets to construct augmented datasets, providing a new data-driven paradigm for knowledge distillation.
[0007] However, existing research largely focuses on knowledge distillation of large image foudation models (IFMs) in the image domain, and its application in downstream tasks such as action recognition remains relatively limited. Although some works have attempted to transfer CLIP to the video domain through two-stage training (distillation-transfer),
[14] However, these methods either rely on complex adaptation structures and do not fully exploit the gains of bimodal interactions for temporal modeling. It is worth noting that the exploration of distilling knowledge directly from large video-text pairs (such as the visual foundation model VFM) to serve video recognition tasks is still in its infancy, and traditional transfer learning requires additional adaptation modules and is time-consuming.
[0008] Therefore, to address the dual challenges of reducing model complexity and achieving cross-modal semantic alignment between text and video, an end-to-end video multimodal distillation framework specifically designed for action recognition is needed. This framework can effectively reduce model complexity, enhance the expressive power of video text features, and improve detection accuracy and robustness.
[0009] References: [1] Radford A, Kim J W, Hallacy C, et al. Learning transferablevisual models from naturallanguage supervision. International conference onmachine learning. PMLR, 2021: 8748-8763. [2] Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, GaofengMeng, Jianlong Fu, Shiming Xiang, and Haibin Ling. Expanding language-imagepretrained models for general video recognition. In ECCV, pages 1–18.Springer, 2022. [3]Wang, Mengmeng et al. “ActionCLIP: A New Paradigm for Video ActionRecognition.”ArXiv abs / 2109.08472 (2021): n. pag. [4]Wu, Wenhao et al. “Bidirectional Cross-Modal Knowledge Explorationfor Video Recognition with Pre-trained Vision-Language Models.” 2023 IEEE / CVFConference on Computer Vision and Pattern Recognition (CVPR) (2022): 6620-6630. [5]Wu, K., Zhang, J., Peng, H., Liu, M., Xiao, B., Fu, J., & Yuan, L.(2022). Tinyvit: Fast pretraining distillation for small vision transformers.arXiv preprint arXiv:2207.10666. [6]Wang, J., Cao, M., Shi, S., Wu, B., & Yang, Y. (2022, May).Attention Probe: Vision Transformer Distillation in the Wild. ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP) (pp. 2220-2224). IEEE. [7]Ren, S., Gao, Z., Hua, T., Xue, Z., Tian, Y., He, S., & Zhao, H.(2022). Co-advise: Cross inductive bias distillation. In Proceedings of theIEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 16773-16782). [8]Li, Zheng et al. “PromptKD: Unsupervised Prompt Distillation forVision-Language Models.”ArXiv abs / 2403.02781 (2024): n. pag. [9]Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., & Jégou, H. (2021, July). Training data -efficient image transformers &distillation through attention. In International Conference on MachineLearning (pp. 10347-10357). PMLR.
[10] Kan Wu, Houwen Peng, Zhenghong Zhou, Bin Xiao, Mengchen Liu, LuYuan, Hong Xuan, Michael Valenzuela, Xi Stephen Chen, Xinggang Wang, et al.Tinyclip: Clip distillation via affinity mimicking and weight inheritance. InProceedings of the IEEE / CVF International Conference on Computer Vision,pages 21970–21980, 2023.
[11] Chuanguang Yang, Zhulin An, Libo Huang, Junyu Bi, Xinqiang Yu,Han Yang, Boyu Diao, Yongjun Xu. “CLIP-KD: An Empirical Study of CLIP ModelDistillation” Proceedings of the IEEE / CVF Conference on Computer Vision andPattern Recognition (CVPR), 2024, pp. 15952-15962.
[12] Mehta, Sachin and Mohammad Rastegari. “MobileViT: Light-weight,General-purpose, and Mobile-friendly Vision Transformer.”ArXiv abs / 2110.02178(2021): n. pag.
[13] Pavan Kumar Anasosalu Vasu, Hadi Pouransari, Fartash Faghri,Raviteja Vemulapalli, Oncel Tuzel. “MobileCLIP: Fast Image-Text Modelsthrough Multi-Modal Reinforced Training”, Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition (CVPR), 2024, pp.15963-15974
[14] Li, Kunchang, Yali Wang, Yizhuo Li, Yi Wang, Yinan He, Limin Wangand Yu Qiao. “Unmasked Teacher: Towards Training-Efficient Video FoundationModels.” 2023 IEEE / CVF International Conference on Computer Vision (ICCV)(2023): 19891-19903.
[15] Yu, Shaoqing et al. “VTR: Bidirectional Video-TextualTransmission Rail for CLIP-based Video Recognition.” 2024 IEEE InternationalConference on Multimedia and Expo (ICME) (2024): 1-6.
[16] Li, Kunchang et al. “UniFormer: Unified Transformer for EfficientSpatiotemporal Representation Learning.” ArXiv abs / 2201.04676 (2022): n. pag. Summary of the Invention
[0010] The purpose of this invention is to provide an action recognition method based on multimodal large model knowledge distillation, so as to enhance the expressive power of video text features, improve detection accuracy and robustness while reducing model complexity.
[0011] To achieve the above objectives, this invention provides a behavior recognition method based on multimodal large model knowledge distillation, comprising: Step S1: Provide an image preprocessing module and a prefix module, and provide a lightweight video text model as a student model; the lightweight video text model is a dual-branch lightweight multimodal model, which includes a text branch and a visual branch, as well as a hierarchical video text interaction module for enabling interaction between the text branch and the visual branch; the image preprocessing module and the prefix module are used to process video frames and text categories to input the student model; the student model is used to output the final text features and the final video features, process to obtain a category matrix and extract the text category of each sample; Step S2: Provide a teacher model and use the trained teacher model as a knowledge source; Step S3: Provide a transmembrane distillation framework, and train the student model using training samples and the transmembrane distillation framework to enable it to learn from the teacher model; Step S4: After the video frames to be tested are processed by the image preprocessing module, they are input into the trained student model. The final text features and final video features are output, the category matrix is obtained, and the text category of each video frame is extracted.
[0012] The lightweight video text model includes a text branch and a visual branch. The visual branch includes a local feature extraction network and a global feature extraction network. The local feature extraction network includes multiple layers of convolutional modules, and the global feature extraction network includes multiple layers of attention modules. Temporal location encoding is added between the multiple attention modules and the multiple convolutional modules. The local feature extraction network is connected to the image preprocessing module. The text branch includes a text embedding layer and a text feature extraction network. The text embedding layer is connected to a prefix module, and the local feature extraction network is connected to the image preprocessing module. The text embedding layer uses a word segmenter, and the text feature extraction network includes multiple text encoding layers connected in sequence, with each text encoding layer using a Transformer network architecture.
[0013] The layered video text interaction module includes a hybrid modal token located between the shallow attention module and the convolution module of the lightweight video text model, and a video text interaction unit located between the deep attention module and the convolution module of the lightweight video text model. The total number of layers of the text encoding layer and the attention module is M layers. The shallow layer consists of the first R-1 layers of the text encoding layer and the attention module, and the deep layer consists of the R-th to M-th layers of the text encoding layer and the attention module. R is a positive integer greater than 2, and M is a positive integer greater than 3.
[0014] Each text encoding layer's mixed-modal token is created by randomly initializing n learnable tokens, which are then combined to obtain a sequence of mixed-modal tokens. A linear transformation is performed on the sequence of text mixed-modal tokens using a learnable weight matrix to create a corresponding sequence of video mixed-modal tokens. The created hybrid modal tokens are concatenated with the outputs of each attention module and each text encoding layer along the sequence length dimension and then fed into subsequent attention modules and text encoding layers.
[0015] The video-text interaction unit for each layer is set as follows: Step A1: Use the deep attention module as the current layer to obtain the deep video features of the current layer. As input to the video text interaction unit, , where i represents the ordinal number of the attention module layer in the student model; Step A2: Deep features of the video Dimension transformation is performed, followed by local feature extraction via separable convolutional blocks and then flattening to obtain enhanced local spatiotemporal visual features. ; Step A3: Extract local spatiotemporal visual features Add the initial trainable temporal embedding information to incorporate the temporal context, thus obtaining local spatiotemporal visual features incorporating the temporal embedding information. ; Step A4: Incorporating local spatiotemporal visual features with embedded time information First, normalize the data, then feed it into the multi-head attention module, and insert residual connections to obtain enhanced global spatiotemporal visual features. ; Step A5: Transfer the video mixed-modal token of the current layer. As a supervisory signal, the video mixed-modal token With enhanced global spatiotemporal visual features The common inputs are fed into the cross-modal multi-head attention mechanism module to obtain the next layer's video hybrid modality token. Mix video tokens The next-level text mixed-modal token is obtained by mapping to the text space through a linear transformation. ; Step A6: Transfer the video mixed-modal token of the current layer and deep video features The concatenated sequences are then used as input to the attention module of the next layer to obtain the deep video features of the next layer. ;Text mixed modal token and the text features of the current layer After concatenation along the sequence length, the result is used as the input to the next text encoding layer to obtain the text features of the next layer. Step A7: Take the next layer as the new current layer and return to step A1, until the text features of the last layer are obtained. and deep video features , as the final feature of the text and the final feature of the video.
[0016] The teacher model is the VTR-B / 16 model.
[0017] A cross-membrane distillation framework is used to obtain the total loss function, which includes a spatial baseline knowledge distillation module, a temporal feature knowledge distillation module, and a text hierarchical feature knowledge distillation module.
[0018] The space reference knowledge distillation module is configured as follows: Step B1: Divide all layers of the visual branch of the student model into 4 major visual branch layers based on the channel dimension of the video features of the student model; extract the video features of each major visual branch layer of the student model. 'u' represents the ordinal number of the visual branch layer of the student model, and the output features of each visual branch layer of the teacher model are extracted. z represents the ordinal number of the visual branch layer of the teacher model; Step B2: Analyze the video features of the u-th visual branch of the student model's visual branch over a large layer. Perform feature rearrangement and concatenation to obtain student features And the video features of the z-th visual branch layer of the teacher model splicing to obtain teacher characteristics This is to achieve dimensional alignment between student and teacher characteristics; Step B3: Based on student characteristics and teacher characteristics To calculate the similarity matrix ; Step B4: For the similarity matrix Using the sequence mask matrix, the masked similarity matrix is obtained. ; Similarity matrix after masking Student characteristics We then perform weighted analysis to obtain student features aligned with the teacher model. Step B5: Determine teacher characteristics KL divergence calculations are performed on student features aligned with the teacher model to obtain the spatial benchmark knowledge distillation loss function. ; The time-series feature knowledge distillation module is set up as follows: Step C1: Align the output features of the last layer of the visual branch of the student model with the output features of the last layer of the visual branch of the teacher model to obtain the final spatiotemporal representation of the student model with consistent feature dimensions. The final spatiotemporal representation of the teacher model ; Step C2: Based on the final spatiotemporal representation of the student model The final spatiotemporal representation of the teacher model The time-series feature knowledge distillation loss function is obtained by calculating the average KL divergence of all batches. ; The hierarchical text knowledge distillation module is set up as follows: Step D1: Obtain the attention features of the j-th layer of the text encoding layer of the student model. and linear layer activation features Attention features of the k-th layer of the text branch layer in the teacher model and linear layer activation features , j represents the ordinal number of the text encoding layer in the student model, and k represents the ordinal number of the text branch layer in the teacher model; Step D2: Introduce a learnable alignment weight matrix for each layer. The alignment weight matrix Attention features of the j-th layer of the text encoding layer of the student model. and linear layer activation features These are respectively mapped to the attention features of the k-th layer of the text branch of the teacher model. and linear layer activation features Alignment characteristics; Step D3: Calculate the text attention distillation loss by calculating the KL divergence. and text linear layer distillation loss Finally, the summations yield the hierarchical text knowledge distillation loss function. .
[0019] During training, the formula for calculating the total loss function is: , in, For the hierarchical text knowledge distillation loss function, The loss function is the distillation function for time-series feature knowledge. For spatial benchmark knowledge distillation loss function, , , These are the weighting coefficients.
[0020] The image preprocessing module is configured to receive the original video frames and process them to obtain the input data of the visual branch of the student model. The processing of the input data of the visual branch of the student model specifically includes: firstly, normalizing the images of multiple video frames and adjusting the image size and resolution to meet the input requirements of the student model; and then expanding the input data through data augmentation technology. The prefix module adds the prefix "a video of {}" to the original text category to process the input data for the resulting text branch.
[0021] The behavior recognition method based on multimodal large model knowledge distillation of this invention employs a lightweight student model, which can significantly reduce model complexity while utilizing a cross-membrane distillation framework. Furthermore, it leverages a hierarchical video-text interaction module to enhance the alignment and representation capabilities of the model's video and text, thereby achieving efficient, high-precision, and robust behavior recognition and detection performance. In particular, the method provided by this invention exhibits good generalization ability and detection results in various scenarios. Attached Figure Description
[0022] Figure 1 This is a general framework diagram of an action recognition method based on multimodal large model knowledge distillation according to the present invention; Figure 2 Is it like this? Figure 1 The diagram shows a framework of a student model based on a behavior recognition method using multimodal large model knowledge distillation. Figure 3 Is it like this? Figure 2 The diagram shows the framework of the video-text interaction unit of the student model. Figure 4 is as follows Figure 2 The diagram shows the framework of the transmembrane distillation system for the student model. Figure 5 shows... Figure 2 The diagram shows the structure of the convolutional and attention modules in the student model. Detailed Implementation
[0023] Figure 1 This is a general framework diagram of an action recognition method based on multimodal large model knowledge distillation according to the present invention. It is used to reduce model complexity and achieve action recognition through cross-modal semantic alignment between text and video. The multimodal large model adopts a multimodal knowledge distillation framework and uses the VTR-B / 16 model.
[15] As a teacher model, this model not only possesses excellent temporal modeling capabilities but also leverages intermediate layer features to enhance cross-branch information exchange between video and text modalities.
[0024] For ease of subsequent description, the parameters are predefined as follows: R: Number of shallow coding layers; M: The total number of text encoding layers and attention modules in the student model, which is the total number of text branch layers and visual branch layers in the teacher model; Text-based mixed-modal token; Video mixed-modal token; The deep video features of the current layer, i.e., the output of the current layer of the attention module; i represents the ordinal number of the attention module layer in the student model; j represents the ordinal number of the text encoding layer in the student model; k represents the ordinal number of the text branch layer in the teacher model, and k=j in general; z represents the ordinal number of the visual branch layer of the teacher model; u represents the ordinal number of the visual branch of the student model; : The visual feature dimension of the u-th visual branch layer; : The output of the current layer of the j-th text encoding layer in the student model; : Video features of the large layer of the u-th visual branch of the student model; : Video features of the z-th visual branch layer of the teacher model; Teacher features, i.e., the total number of layers in the student model after SBKD feature rearrangement; Student features, i.e., the total number of layers in the teacher model after SBKD feature rearrangement; The similarity matrix W in SBKD; The similarity matrix W' after masking in SBKD; The final visual branch representation of the student model; The final visual branch representation of the teacher model; Attention features of the j-th layer of the text encoding layer in the student model; : The activation features of the j-th linear layer of the text encoding layer in the student model; Attention features of the k-th layer in the text branch of the teacher model Generally, k=j; The linear activation features of the k-th layer in the text branch of the teacher model; The learnable alignment weight matrix introduced in each layer of HTKD; , , The final weight coefficients for several distillation tests are then trained. n: Number of mixed tokens; : Channel dimension representing the textual features of the student model; : Channel dimension of video features in the teacher model; : Represents the channel dimension of the video features of the i-th layer attention module, [320, 512]; B: Represents the batch size, which represents the number of video samples input to the model each time; T: Represents the number of frames in the video sample; : Represents the length of the video feature sequence of the i-th attention module; N=4, where N represents the number of large layers of the visual branches in the student model; : The channel dimension of the video features of the student model in the u-th visual branch layer, i.e. ; The high video features of the student model in the u-th visual branch of the large layer; The width of the video features of the student model in the u-th visual branch layer; L is the sequence length of the video features of the teacher model, L=196; , , , These are all hyperparameters.
[0025] like Figure 1 As shown, the behavior recognition method based on multimodal large model knowledge distillation of the present invention includes: Step S1: Provide an image preprocessing module and a prefix module, and provide a lightweight video text model (also known as a MobileVTL model) as a student model; the lightweight video text model is a two-branch lightweight multimodal model, which includes a text branch and a visual branch, as well as a hierarchical video text interaction module for enabling interaction between the text branch and the visual branch; the image preprocessing module and the prefix module are used to process video frames and text categories to input the student model; the student model is used to output the final text features and the final video features, process to obtain a category matrix and extract the text category of each sample; The student model obtains text features with a size of [batch size, dimension], and video features with a size of [number of categories, dimension]. The text and video features are fused together via matrix multiplication to obtain a category matrix with a size of [batch size, number of categories]. This category matrix is a similarity matrix; by maximizing the number of categories dimension, the text category for each sample in the batch can be extracted. Contrastive learning involves training the class matrix obtained by multiplying the final text and video feature matrices to maximize the similarity between correct video-text pairs and minimize the similarity between incorrect video sample pairs, thereby making correct video-text pairs more discriminative.
[0026] The student model is a core component of the knowledge distillation framework. For example... Figure 1 and Figure 2 As shown, the lightweight video text model is a unique, lightweight, multimodal model with two branches for behavior recognition, designed to fully transfer knowledge from the teacher model. It utilizes a Uniformer architecture.
[16] The model serves as a baseline, with the classification head removed to serve as the visual branch, and a text branch added to extend it to the multimodal domain. The Uniformer model integrates convolutional neural networks and visual Transformers, thus combining the advantages of CNNs such as spatial inductive bias and robustness to data augmentation with the advantages of visual Transformers (ViT) such as adaptive input weighting and global processing, serving as local and global temporal modelers respectively. The text branch utilizes the CLIP-Transformer module with fewer parameters. [1] .
[0027] To better facilitate modal information interaction, this invention proposes a Hierarchical Video Text Interaction Module (HVTI). The HVTI includes a hybrid modal token (i.e., a video / text hybrid modal token) and a video text interaction unit, dynamically promoting the interaction of textual and video semantic features at lower and higher levels. The lower levels contain higher resolution and spatial information, while the higher levels contain richer semantic information. Therefore, for the lower levels, we use a hybrid modal token (MM token) to facilitate information interaction between textual and visual branches. For the higher levels, we use a video text interaction unit for further semantic interaction, using textual information to supervise visual encoding.
[0028] For the training phase, this invention proposes a cross-modal distillation framework (CMKD) that combines four distillation strategies. The distillation loss includes Spatial Base KD, Temporal Feature Distillation (TFD), and Text Hierarchical Distillation (THKD). The differences between the spatial features, text features, and temporal features of the teacher model and the student model are calculated hierarchically, enabling the student model to obtain better cross-modal representation capabilities.
[0029] The specific descriptions of the modules and structures involved in step S1 are as follows: (1) Image preprocessing module and prefix module The image preprocessing module is configured to receive raw video frames and process them to obtain the input data for the visual branch of the student model. The processing of the input data for the visual branch of the student model specifically includes: firstly, normalizing the images of multiple video frames and adjusting their size and resolution to suit the input requirements of the student model; then, expanding the input data through data augmentation techniques such as random rotation and cutmixup to simulate various possible real-world scenarios and improve the model's generalization ability.
[0030] The prefix module adds the prefix "a video of {}" to the original text category to process the input data of the obtained text branch, thereby simulating various possible real-world scenarios and improving the generalization ability of text features.
[0031] (2) Lightweight video text model like Figure 2As shown, the student model employs a lightweight video-text model, which includes a text branch and a visual branch. The visual branch includes local feature extraction networks and a global feature extraction network. The local feature extraction network comprises multiple layers of convolutional CNNs (S1 + S2), and the global feature extraction network comprises multiple layers of attention modules. Temporal location encoding is added between the attention modules and the convolutional CNNs to implement a visual Transformer. The text branch includes a text embedding layer and a text feature extraction network. The text embedding layer is connected to a prefix module, and the local feature extraction network is connected to an image preprocessing module. The text embedding layer uses a tokenizer, and the text feature extraction network comprises multiple sequentially connected text encoding layers, each employing a Transformer network architecture. The specific structures of the convolutional and attention modules are shown below. Figure 5 As shown.
[0032] (3) Layered Video Text Interaction Module (HVTI) Existing multimodal behavior recognition models do not adequately mine features in the intermediate layers of text branches, neglecting the potential correlation between text and visual modalities in these intermediate layers. Furthermore, student models, due to their limited number of parameters, have relatively weak temporal modeling capabilities and exhibit a certain gap in robustness compared to larger models. Therefore, this invention focuses on enhancing the information interaction between visual and text branches to improve the quality of temporal representation, proposing a hierarchical video-text interaction module.
[0033] like Figure 2 As shown, the hierarchical video-text interaction module includes a mixed-modal token (MM token) located between the shallow attention and convolutional modules of the lightweight video-text model, and a video-text interaction unit located between the deep attention and convolutional modules of the lightweight video-text model. Thus, in the shallow layers of the model (such as layers 1 to R-1), bidirectional cross-modal knowledge transfer between the attention module and the text encoder layer is achieved through the MM token; while in the deeper stages, due to the richer semantic information, a more complex video-text interaction mechanism is introduced to achieve bidirectional deep interaction: on the one hand, text information guides and enhances the temporal modeling ability of the visual branch; on the other hand, visual information is used to supervise the representation learning of the text branch.
[0034] The hybrid modality tokens include video hybrid modality tokens corresponding to each attention module in the shallow layer, text hybrid modality tokens corresponding to each text encoding layer in the shallow layer, and a learnable weight matrix for linearly transforming the sequence of text hybrid modality tokens into a sequence of video hybrid modality tokens.
[0035] Among them, the total number of layers of the text encoding layer of the student model and the total number of layers of the attention module are both M layers. The shallow layer is the first R - 1 layers (R < M) of the text encoding layer and the attention module. R is a positive integer greater than 2, and M is a positive integer greater than 3. Therefore, the text hybrid modality tokens of each text encoding layer are created by randomly initializing n learnable tokens, and the text hybrid modality tokens are denoted as , where the subscript t represents text, n represents the number of tokens, D t represents the channel dimension of the text features of the student model, and j represents the ordinal number of the text encoding layer of the student model. Finally, a sequence of text hybrid modality tokens is obtained . To promote the information interaction between the visual branch and the text branch, the present invention linearly transforms the sequence of text hybrid modality tokens through a learnable weight matrix to create a corresponding sequence of video hybrid modality tokens , and the video hybrid modality tokens , where the subscript v represents video, n represents the number of tokens, i represents the ordinal number of the attention module layer of the student model (whose value is equal to the ordinal number j of the text encoding layer of the student model), represents the channel dimension of the video features of the i-th layer attention module. As Figure 2 shown, the present invention splices the created hybrid modality tokens with the output of each attention module and the output of each text encoding layer in the sequence length dimension and sends them to the subsequent attention module and text encoding layer to achieve bidirectional cross-modal knowledge transfer.
[0036] To promote the deep bidirectional interaction between the video and text modalities, the present invention proposes a parallel multi-layer video text interaction unit (Video Text Interaction Unit, VTI). Each layer of the video text interaction unit corresponds to one of the deep text encoding layer and the attention module. The total number of layers of the text encoding layer and the attention module is both M layers, and the deep layer is the R-th layer to the M-th layer of the text encoding layer and the attention module.
[0037] The detailed structure of each layer of the video text interaction unit is as Figure 3 shown. Each layer of the video text interaction unit is set as: Step A1: Take the deep attention module as the current layer to obtain the deep video features of the current layer as the input of the video text interaction unit, where , B represents the batch size, which represents the number of video samples input to the model each time, T represents the number of frames of the video sample, represents the sequence length of the video features of the i-th layer attention module, The channel dimension representing the video features of the i-th layer attention module. 'i' represents the ordinal number of the attention module layer in the student model. The video mixed-modal token... The dimension is .
[0038] Step A2: Deep features of the video Dimension transformation is performed, and then local features are extracted and flattened through separable convolutional blocks (Divided Conv) to obtain enhanced local spatiotemporal visual features; Among them, the dimensions of deep video features are... First convert to Local spatiotemporal visual features The dimension is .
[0039] Step A3: Extract local spatiotemporal visual features With initial trainable temporal embedding information time_embedding By adding them together and incorporating the temporal context, local spatiotemporal visual features with embedded temporal information are obtained. ; Local spatiotemporal visual features incorporating time-embedded information The specific formula is: , For deep features of the video, For processing separable convolutional blocks, Embed information for time.
[0040] Step A4: Incorporating local spatiotemporal visual features with embedded time information First, normalize the data, then feed it into the multi-head attention module, and insert residual connections to obtain enhanced global spatiotemporal visual features. To enhance gradient flow and stabilize the training process; The specific formula for the enhanced global spatiotemporal visual features is as follows:
[0041] Among them, the local spatiotemporal visual features that incorporate time-embedded information are the inserted residual connections; For multi-head attention module processing, This is for normalization purposes.
[0042] Step A5: Transfer the video mixed-modal token of the current layer. As a supervisory signal, the video mixed-modal token With enhanced global spatiotemporal visual features The common inputs are fed into the Cross-Modal Multi-Head Attention (CrossMHA) module to obtain the next layer's video mixed-modality token. Mix video tokens The next-level text mixed-modal token is obtained by mapping to the text space through a linear transformation. ; Step A6: Transfer the video mixed-modal token of the current layer and deep video features The concatenated sequences are then used as input to the attention module of the next layer to obtain the deep video features of the next layer. ;Text mixed modal token The text features of the current layer are concatenated with the text features of the current layer in terms of sequence length and then used as the input of the text encoding layer of the next layer to obtain the text features of the next layer, thereby realizing the two-way information interaction between deep video and text. The specific operational formulas involved in steps A5 and A6 are shown below:
[0043]
[0044]
[0045] .
[0046] in, For deep features of the video, For text features, For video mixed-modal tokens, For text-based mixed-modal tokens, For the processing of cross-modal multi-head attention mechanism modules, For linear changes, For the attention module processing, For text encoding layer processing, This is a splicing process.
[0047] Step A7: Take the next layer as the new current layer and return to step A1, until the text features of the last layer are obtained. and deep video features , as the final feature of the text and the final feature of the video.
[0048] Thus, through a multi-layered stacked interaction mechanism, the Video Text Interaction Unit (VTI) can effectively achieve deep semantic alignment and information fusion between video and text modalities.
[0049] Step S2: Provide a teacher model and use the trained teacher model as a knowledge source; In this process, video frames and text categories are used as training samples to train the teacher model, resulting in a trained teacher model.
[0050] In this embodiment, the VTR model is selected as the teacher model. The VTR model is one of the leading multimodal behavior recognition models in the field, possessing powerful visual-text representation capabilities and temporal modeling performance. The teacher model is preferably the VTR-B / 16 model.
[0051] Step S3: Provide a transmembrane distillation framework, and train the student model using training samples and the transmembrane distillation framework to enable it to learn from the teacher model; (3) Transmembrane distillation framework (CMKD) The text above mentions a significant performance gap between the teacher and student models. To reduce this performance difference, such as... Figure 4 As shown, the cross-membrane distillation framework is used to obtain the total loss function. It includes the Spatial Baseline Knowledge Distillation Module (SBKD), the Temporal Feature Knowledge Distillation Module (TFKD), and the Text Hierarchical Feature Knowledge Distillation Module (THKD) to help the student model learn the teacher model during the training phase from three perspectives: image details, action coherence, and text content.
[0052] A. Spatial Reference Knowledge Distillation Module (SBKD) According to existing technologies, the middle and higher levels of the visual branch have a high number of input channels and rich semantic information, but low resolution and less spatial information; the lower levels are the opposite, with a low number of input channels, high input resolution, rich spatial information, and weaker semantic information. [8] Therefore, we proposed spatial benchmark distillation, which is specifically designed to capture teachers' spatial coding abilities across all layers from shallow to deep.
[0053] The Space Reference Knowledge Distillation (SBKD) module is configured as follows: Step B1: All layers of the visual branch of the student model are calculated based on the channel dimension of the video features of the student model. The model is divided into four visual branches; video features are extracted from each visual branch of the student model. 'u' represents the ordinal number of the visual branch layer of the student model, and the output features of each visual branch layer of the teacher model are extracted. z represents the ordinal number of the visual branch layer of the teacher model; Therefore, the first two layers of the four visual branches are convolutional modules (64 and 128 channels), responsible for capturing shallow spatial details (such as limb edges and action contours); the last two layers are attention modules (320 and 512 channels), responsible for capturing deep spatial semantics (such as relative limb positions and scene associations). The video features of each major layer in the visual branches... As a learning foundation for the teacher model, it enables the student model to learn from the teacher model.
[0054] The underlying formula for using the four visual branch layers as the learning foundation of the teacher model is as follows: (1) Where z represents the ordinal number of the visual branch layer in the teacher model, u represents the ordinal number of the major visual branch layer in the student model, M is the number of layers in the teacher model (the total number of text branch layers and visual branch layers in the teacher model is M), and N=4, where N represents the major visual branch layer number in the student model. These are the video features at the z-th layer of the visual branch of the teacher model. It represents the video features of the large layer in the u-th visual branch of the student model. It is the weight matrix for each layer.
[0055] Step B2: Analyze the video features of the u-th visual branch of the student model's visual branch over a large layer. Perform feature rearrangement and concatenation to obtain student features And the video features of the z-th layer of the teacher model splicing to obtain teacher characteristics This is to achieve dimensional alignment between student and teacher characteristics; Since the student model is a hybrid architecture of CNN-Transformer and the teacher model is a pure Transformer architecture, in order to overcome the feature differences caused by heterogeneity, and since each layer of the teacher model is a linear combination of multiple layers of the student model, SBKD aligns the features of the student model and the teacher model through feature rearrangement.
[0056] Feature rearrangement specifically includes: rearranging the video features of the u-th visual branch of the student model. From 3D features (size is ) is transformed into a 2D feature (size [BT, ...) through pixel rearrangement and flattening. , p×p]), to obtain the flattened student features, and then splice them together to obtain the student features.
[0057] Since the input and output of a CNN convolutional layer are 3D, they are first converted to 2D through pixel rearrangement. Assume the student model's visual branch has N layers, and the teacher model's text and visual branches each have M layers. The video features of the u-th visual branch layer of the student model... The size is B is the batch size, and T is the number of frames in the video sample. The channel dimension representing the video features of the student model in the u-th visual branch layer. For the high-order video features of the student model in the u-th visual branch, For the width of the video features of the student model in the large layer of the u-th visual branch, , The sequence length is given by the video features of the z-th layer of the teacher model. The size is B represents the batch size, T represents the number of frames in the video sample, and L is the sequence length of the video features in the teacher model, L=196. Let be the channel dimension of the video features for the teacher model. Therefore, feature reordering includes: rearranging the video features of the u-th visual branch of the student model. Mapped to the first intermediate feature through linear transformation Its size is Then, the size is changed to [BT, ] through pixel rearrangement. [,p,p], and then flattened to make the size become , The length of the sequence after feature rearrangement. p is the pixel rearrangement block size. The channel dimension of the video features for the teacher model; the flattened student features are then concatenated to obtain the student features. N represents the number of large layers of the visual branches of the student model, N=4.
[0058] The video features of the z-th layer of the teacher model Expand the number of layers and concatenate them along these layers to obtain the teacher characteristics. Its size is ,in =L.
[0059] Step B3: Based on student characteristics and teacher characteristics To calculate the similarity matrix Thus, similarity calculation and similarity matrix were achieved. Used to measure the similarity of characteristics among teachers and students at different levels.
[0060] Among them, teacher characteristics and transposed student characteristics The similarity matrix is obtained by matrix multiplication. Its size is .
[0061] Step B4: For the similarity matrix Using the sequence mask matrix, the masked similarity matrix is obtained. ; Similarity matrix after masking Student characteristics We then perform weighted analysis to obtain student features aligned with the teacher model. The principle of step B4 is that students' lower levels should learn from the teacher's lower levels, while students' higher levels should learn from all levels of the teacher. Based on this, the present invention designs a sequence mask.
[0062] Sequence mask matrix The calculation formula is as follows: (2) in, , These are hyperparameters, where u represents the ordinal number of the visual branch layer of the student model, and z represents the ordinal number of the visual branch layer of the teacher model.
[0063] In step S4, the sequence mask matrix and the similarity matrix are... Element-wise multiplication yields the masked similarity matrix. To ensure training convergence, this invention also modifies the masked similarity matrix. Perform softmax calculation.
[0064] In step S4, the masked similarity matrix is... With student characteristics Multiplying to achieve weighting, we obtain aligned student features; the size of the aligned student features is... .
[0065] Step B5: Determine teacher characteristics KL divergence calculations are performed on student features aligned with the teacher model to obtain the spatial benchmark knowledge distillation loss function. .
[0066] Teacher characteristics The size is
[0067] Spatial benchmark knowledge distillation loss function The calculation formula is: (3) in, Here is the KL divergence calculation function. The matrix is the sequence mask. This is a similarity matrix. For student characteristics, To align student characteristics, Characteristics of teachers.
[0068] B. Time-Series Feature Knowledge Distillation Module (TFKD) Temporal modeling capability is crucial for action recognition. Since the teacher model selected in this invention has strong temporal modeling capabilities, directly distilling the temporal features of the teacher model can minimize the temporal feature distance between students and teachers.
[0069] The Time-Series Feature Knowledge Distillation (TFKD) module is configured as follows: Step C1: Extract the output features of the last layer of the visual branch of the student model (i.e., the deep video features of the last layer). The output features of the last layer of the visual branch of the student model are aligned with those of the teacher model to obtain a representation of the last layer of the visual branch of the student model with consistent feature dimensions. The representation of the last layer of the visual branch of the teacher model ; Feature alignment can be achieved by linearly mapping the output features of the last layer of the visual branch of the student model in a dimension to align them with the output features of the last layer of the visual branch of the teacher model.
[0070] Step C2: Based on the final visual branch representation of the student model The final visual branch representation of the teacher model The time-series feature knowledge distillation loss function is obtained by calculating the average KL divergence of all batches. .
[0071] Therefore, in the Temporal Feature Knowledge Distillation (TFKD) module, the temporal feature knowledge distillation loss function... The calculation formula is: (4) Where B is the batch size, |B| is the total number of batches, b represents the batch ordinal number, and y represents the feature dimension ordinal number. The channel dimension representing the features of the teacher model. This represents the value of the y-th feature dimension in the final visual branch representation of the teacher model. In the final visual branch expression, the feature dimensions for students and teachers are consistent.
[0072] C. Hierarchical Text Knowledge Distillation Module (HTKD) In multimodal behavior recognition, the encoding ability of the text branch is crucial. The student model MobileVTL uses a smaller Transformer for its text branch. In order to reduce the difference between the text feature space and the teacher's text, this invention proposes hierarchical text knowledge distillation (HTKD), which learns the teacher's encoding ability from the low, middle and high layers respectively.
[0073] The Hierarchical Text Knowledge Distillation (HTKD) module is configured as follows: Step D1: Obtain the attention features of the j-th layer of the text encoding layer of the student model. and linear layer activation features Attention features of the k-th layer of the text branch layer in the teacher model and linear layer activation features , j represents the ordinal number of the text encoding layer in the student model, and k represents the ordinal number of the text branch layer in the teacher model; Step D2: Introduce a learnable alignment weight matrix for each layer. The alignment weight matrix Attention features of the j-th layer of the text encoding layer of the student model. and linear layer activation features These are respectively mapped to the attention features of the k-th layer of the text branch of the teacher model. and linear layer activation features Alignment characteristics; Step D3: Calculate the text attention distillation loss by calculating the KL divergence. and text linear layer distillation loss Finally, the summations yield the hierarchical text knowledge distillation loss function. .
[0074] Text attention distillation loss Text linear layer distillation loss The calculation formula is: (5) (6) Hierarchical text knowledge distillation loss function for: (7) The transmembrane distillation framework is used to obtain the total loss function, which is calculated during training using the following formula: , in, For the hierarchical text knowledge distillation loss function, The loss function is the distillation function for time-series feature knowledge. For spatial benchmark knowledge distillation loss function, , , These are the weighting coefficients.
[0075] Step S4: After the video frames to be tested are processed by the image preprocessing module, they are input into the trained student model. The final text features and final video features are output, the category matrix is obtained, and the text category of each video frame is extracted.
[0076] Through the implementation of the above embodiments, the behavior recognition method based on multimodal large model knowledge distillation of the present invention can significantly reduce model complexity by utilizing a cross-membrane distillation framework, while improving the alignment and representation capabilities of model video and text using a hierarchical video-text interaction module, thereby achieving efficient, high-precision, and robust behavior recognition and detection performance. In particular, the method provided by the present invention exhibits good generalization ability and detection effect in various scenarios.
[0077] The above description is merely a description of preferred embodiments of the present invention and is not intended to limit the scope of the present invention in any way. Any changes or modifications made by those skilled in the art based on the above disclosure shall fall within the protection scope of the claims.
Claims
1. A behavior recognition method based on multimodal large model knowledge distillation, characterized in that, include: Step S1: Provide an image preprocessing module and a prefix module, and provide a lightweight video text model as a student model; the lightweight video text model is a dual-branch lightweight multimodal model, which includes a text branch and a visual branch, as well as a hierarchical video text interaction module for enabling interaction between the text branch and the visual branch; the image preprocessing module and the prefix module are used to process video frames and text categories to input the student model; the student model is used to output the final text features and the final video features, process to obtain a category matrix and extract the text category of each sample; Step S2: Provide a teacher model and use the trained teacher model as a knowledge source; Step S3: Provide a transmembrane distillation framework, and train the student model using training samples and the transmembrane distillation framework to enable it to learn from the teacher model; Step S4: After the video frames to be tested are processed by the image preprocessing module, they are input into the trained student model. The final text features and final video features are output, the category matrix is obtained, and the text category of each video frame is extracted.
2. The behavior recognition method based on multimodal large model knowledge distillation according to claim 1, characterized in that, The lightweight video text model includes a text branch and a visual branch. The visual branch includes a local feature extraction network and a global feature extraction network. The local feature extraction network includes multiple layers of convolutional modules, and the global feature extraction network includes multiple layers of attention modules. Temporal location encoding is added between the multiple attention modules and the multiple convolutional modules. The local feature extraction network is connected to the image preprocessing module. The text branch includes a text embedding layer and a text feature extraction network. The text embedding layer is connected to a prefix module, and the local feature extraction network is connected to the image preprocessing module. The text embedding layer uses a word segmenter, and the text feature extraction network includes multiple text encoding layers connected in sequence, with each text encoding layer using a Transformer network architecture.
3. The behavior recognition method based on multimodal large model knowledge distillation according to claim 2, characterized in that, The layered video text interaction module includes a hybrid modal token located between the shallow attention module and the convolution module of the lightweight video text model, and a video text interaction unit located between the deep attention module and the convolution module of the lightweight video text model. The total number of layers of the text encoding layer and the attention module is M layers. The shallow layer consists of the first R-1 layers of the text encoding layer and the attention module, and the deep layer consists of the R-th to M-th layers of the text encoding layer and the attention module. R is a positive integer greater than 2, and M is a positive integer greater than 3.
4. The behavior recognition method based on multimodal large model knowledge distillation according to claim 3, characterized in that, Each text encoding layer's mixed-modal token is created by randomly initializing n learnable tokens, which are then combined to obtain a sequence of mixed-modal tokens. ; A linear transformation is performed on the sequence of text mixed-modal tokens using a learnable weight matrix to create the corresponding sequence of video mixed-modal tokens. The created hybrid modal tokens are concatenated with the outputs of each attention module and each text encoding layer along the sequence length dimension and then fed into subsequent attention modules and text encoding layers.
5. The behavior recognition method based on multimodal large model knowledge distillation according to claim 3, characterized in that, The video-text interaction unit for each layer is set as follows: Step A1: Use the deep attention module as the current layer to obtain the deep video features of the current layer. As input to the video text interaction unit, , where i represents the ordinal number of the attention module layer in the student model; Step A2: Deep features of the video Dimension transformation is performed, followed by local feature extraction via separable convolutional blocks and then flattening to obtain enhanced local spatiotemporal visual features. ; Step A3: Extract local spatiotemporal visual features Add the initial trainable temporal embedding information to incorporate the temporal context, thus obtaining local spatiotemporal visual features incorporating the temporal embedding information. ; Step A4: Incorporating local spatiotemporal visual features with embedded time information First, normalize the data, then feed it into the multi-head attention module, and insert residual connections to obtain enhanced global spatiotemporal visual features. ; Step A5: Transfer the video mixed-modal token of the current layer. As a supervisory signal, the video mixed-modal token With enhanced global spatiotemporal visual features The common inputs are fed into the cross-modal multi-head attention mechanism module to obtain the next layer's video hybrid modality token. Mix video tokens The next-level text mixed-modal token is obtained by mapping to the text space through a linear transformation. ; Step A6: Transfer the video mixed-modal token of the current layer and deep video features The concatenated sequences are then used as input to the attention module of the next layer to obtain the deep video features of the next layer. ; Text mixed modal token and the text features of the current layer After concatenation along the sequence length, the result is used as the input to the next text encoding layer to obtain the text features of the next layer. Step A7: Take the next layer as the new current layer and return to step A1, until the text features of the last layer are obtained. and deep video features , as the final feature of the text and the final feature of the video.
6. The behavior recognition method based on multimodal large model knowledge distillation according to claim 1, characterized in that, The teacher model is the VTR-B / 16 model.
7. The behavior recognition method based on multimodal large model knowledge distillation according to claim 1, characterized in that, A cross-membrane distillation framework is used to obtain the total loss function, which includes a spatial baseline knowledge distillation module, a temporal feature knowledge distillation module, and a text hierarchical feature knowledge distillation module.
8. The behavior recognition method based on multimodal large model knowledge distillation according to claim 7, characterized in that, The space reference knowledge distillation module is configured as follows: Step B1: Divide all layers of the visual branch of the student model into 4 major visual branch layers based on the channel dimension of the video features of the student model; extract the video features of each major visual branch layer of the student model. 'u' represents the ordinal number of the visual branch layer of the student model, and the output features of each layer of the teacher model are extracted. z represents the ordinal number of the visual branch layer of the teacher model; Step B2: Analyze the video features of the u-th visual branch of the student model's visual branch over a large layer. Perform feature rearrangement and concatenation to obtain student features And the video features of the z-th visual branch layer of the teacher model splicing to obtain teacher characteristics This is to achieve dimensional alignment between student and teacher characteristics; Step B3: Based on student characteristics and teacher characteristics To calculate the similarity matrix ; Step B4: For the similarity matrix Using the sequence mask matrix, the masked similarity matrix is obtained. ; Similarity matrix after masking Student characteristics We then perform weighted analysis to obtain student features aligned with the teacher model. Step B5: Determine teacher characteristics KL divergence calculations are performed on student features aligned with the teacher model to obtain the spatial benchmark knowledge distillation loss function. ; The time-series feature knowledge distillation module is set up as follows: Step C1: Align the output features of the last layer of the visual branch of the student model with the output features of the last layer of the visual branch of the teacher model to obtain the final visual branch representation of the student model with consistent feature dimensions. The final visual branch representation of the teacher model ; Step C2: Based on the final visual branch representation of the student model The final visual branch representation of the teacher model The time-series feature knowledge distillation loss function is obtained by calculating the average KL divergence of all batches. ; The hierarchical text knowledge distillation module is set up as follows: Step D1: Obtain the attention features of the j-th layer of the text encoding layer of the student model. and linear layer activation features Attention features of the k-th layer of the text branch layer in the teacher model and linear layer activation features , j represents the ordinal number of the text encoding layer in the student model, and k represents the ordinal number of the text branch layer in the teacher model; Step D2: Introduce a learnable alignment weight matrix for each layer. The alignment weight matrix Attention features of the j-th layer of the text encoding layer of the student model. and linear layer activation features These are respectively mapped to the attention features of the k-th layer of the text branch of the teacher model. and linear layer activation features Alignment characteristics; Step D3: Calculate the text attention distillation loss by calculating the KL divergence. and text linear layer distillation loss Finally, the summations yield the hierarchical text knowledge distillation loss function. .
9. The behavior recognition method based on multimodal large model knowledge distillation according to claim 7, characterized in that, During training, the formula for calculating the total loss function is: , in, For the hierarchical text knowledge distillation loss function, The loss function is the distillation function for time-series feature knowledge. For spatial benchmark knowledge distillation loss function, , , These are the weighting coefficients.
10. The behavior recognition method based on multimodal large model knowledge distillation according to claim 1, characterized in that, The image preprocessing module is configured to receive the original video frames and process them to obtain the input data of the visual branch of the student model. The processing of the input data of the visual branch of the student model specifically includes: firstly, normalizing the images of multiple video frames and adjusting the image size and resolution to meet the input requirements of the student model; and then expanding the input data through data augmentation technology. The prefix module adds the prefix "a video of {}" to the original text category to process the input data that yields the text branch.