Large Model-Driven Spatiotemporal Feature and Text-Enhanced Few-Shot Action Capture Method

Through the large-model-driven spatio-temporal features and text enhancement methods, video and text information are integrated to build a type prototype with strong generalization capabilities, solving the problem of insufficient motion capture accuracy in a learning environment with few samples in the existing technology, and achieving efficient and accurate motion capture effect.

CN119903479BActive Publication Date: 2025-06-13CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510388665.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-13
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing motion capture methods perform poorly in a learning environment with few samples and fail to effectively integrate text information, resulting in insufficient capture accuracy.

Method used

The space-time feature and text enhancement method driven by large-models are used to extract video features through visual encoder, and combined with CLIP text encoder and prototype construction modules, multimodal information is integrated to build a class prototype with strong generalization capabilities.

Benefits of technology

It significantly improves the accuracy of motion capture in the learning task of few samples, reduces the computational cost, and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119903479B_ABST
    Figure CN119903479B_ABST
Patent Text Reader

Abstract

The present invention discloses a large model-driven spatio-temporal feature and text-enhanced few-shot action capture method, belonging to the technical field of action capture and used for video action capture. It includes obtaining video data and performing preprocessing. The video data includes query video data to be subjected to action capture and support set video data with action labels. The preprocessed video data is input into a visual encoder to obtain the visual features of the video data. By synthesizing the category probability distributions twice, the action capture result of the query video data is obtained. The present invention realizes efficient spatio-temporal feature extraction through a temporal enhancement adapter and a spatio-temporal fusion adapter, enhancing the spatio-temporal modeling ability of video features; utilizes a multi-level attention mechanism to improve the fusion ability of text and video features, and constructs a class prototype with strong generalization ability; significantly improves the capture accuracy of the model in few-shot learning tasks, has fewer trainable parameters, and reduces the computational cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a large model-driven spatio-temporal feature and text-enhanced few-shot action capture method, belonging to the technical field of action capture. Background Art

[0002] In the field of action capture, few-shot learning is of great significance when dealing with limited labeled data of new categories. Traditional methods usually rely on large-scale data for training, resulting in high computational costs and limited generalization ability. Existing technologies still face challenges in efficiently extracting spatio-temporal features of videos and constructing class prototypes with strong generalization ability. Existing video action capture methods usually adopt a two-stage strategy. First is the feature extraction stage, where a pre-trained convolutional neural network is used to extract the spatial features of video frames. Second is the temporal modeling stage, where a recurrent neural network or a long short-term memory network is used to model the sequence information. Although these methods have achieved good results on standard datasets, in the few-shot scenario, due to the limitation of the amount of data, they often perform poorly. In addition, the training of these methods usually requires a large amount of computing resources, and there are significant time complexity problems when dealing with long videos. Some methods attempt to combine meta-learning and prototype learning to enhance the generalization ability of the model by constructing class prototypes. These methods use a small number of samples to learn feature prototypes that can represent the entire category, thereby improving the action capture ability on new categories. However, when constructing class prototypes, existing methods usually ignore the fusion of multi-modal information, especially the potential value of text information. As a rich semantic clue, text information can provide powerful assistance for video action capture. Summary of the Invention

[0003] The purpose of the present invention is to provide a large model-driven spatio-temporal feature and text-enhanced few-shot action capture method to solve the problem of insufficient capture accuracy caused by the lack of fusion of text information in existing action capture methods.

[0004] The large model-driven spatio-temporal feature and text-enhanced few-shot action capture method includes obtaining video data and performing preprocessing. The video data includes query video data to be subjected to action capture and support set video data with action labels. The preprocessed video data is input into a visual encoder to obtain the visual features of the video data;

[0005] Input the action tags of the support set into the CLIP text encoder to obtain the text features of the action tags of the support set. Calculate the similarity between the text features and the visual features of the query video data to obtain the first category probability distribution. Input the visual features and text features of the support set video data into the prototype construction module to obtain the prototype features of the support set video data. Input the visual features of the query video data into the prototype construction module to obtain the prototype features of the query video data. Calculate the time series similarity between the prototype features of the support set video data and the prototype features of the query video data to obtain the second category probability distribution.

[0006] Integrate the two category probability distributions to obtain the action capture result of the query video data.

[0007] Obtaining the visual features of the video data includes designing a temporal sequence enhancement adapter and a spatio-temporal fusion adapter, and alternately embedding the two adapters into the CLIP-ViT visual encoder.

[0008] The CLIP-ViT visual encoder includes N VIT Block layers. In the even layers, the temporal sequence enhancement adapter is embedded before and after the multi-head attention layer respectively. In the odd layers, the spatio-temporal fusion adapter is only embedded after the multi-head attention layer. VITBlock is a visual block.

[0009] The even layer sequentially sends the input information to the TE Adapter layer, the Layer Norm layer, the Multi-HeadAttention layer, the feature fusion layer, the TE Adapter layer, the Layer Norm layer, the MLP layer, and the feature fusion layer, and then outputs. The output of the first TE Adapter layer is provided with a shortcut connection to the first feature fusion layer, and the output of the second TE Adapter layer is provided with a shortcut connection to the second feature fusion layer.

[0010] The TE Adapter layer sequentially includes an FC Down layer, a Gelu layer, an FC Up layer, and a feature fusion layer. The input of the TEAdapter layer is provided with a shortcut connection to the feature fusion layer.

[0011] The TE Adapter is a temporal sequence enhancement adapter, Layer Norm is layer normalization, Multi-HeadAttention is multi-head attention, MLP is a multi-layer perceptron, FC Down is fully connected downsampling, FC Up is fully connected upsampling, and Gelu is an activation function.

[0012] The odd-numbered layers sequentially send the input information to the Layer Norm layer, the Multi-Head Attention layer, the feature fusion layer, the STF Adapter layer, the Layer Norm layer, the MLP layer, and the feature fusion layer, and then output; the input of the spatio-temporal fusion adapter is provided with a shortcut connection to the first feature fusion layer, and the output of the STF Adapter layer is provided with a shortcut connection to the second feature fusion layer;

[0013] The STF Adapter layer sequentially includes an FC Down layer, a Max-Pool layer, a 3D-Conv layer, a Sigmiod layer, a neuron layer, an FC Up layer, and a feature fusion layer. The input of the STF Adapter layer is provided with a shortcut connection to the feature fusion layer, and the output of the FC Down layer is provided with a shortcut connection to the neuron layer;

[0014] The STF Adapter is a spatio-temporal fusion adapter, Max-Pool is max pooling, 3D-Conv is three-dimensional convolution, and Sigmiod is an activation function.

[0015] The similarity calculation includes performing a max pooling operation on the visual features of the query video data, and then calculating the cosine similarity between the visual features and the text features of the query video data to obtain the first category probability distribution.

[0016] The prototype construction module includes performing a Concat operation on the visual feature Video Features and the text feature Txext Features as the query q. The Concat is for fusion. The text feature is input to the Repeat layer, and then undergoes the first feature fusion with the visual feature as the key k and the value v. The q, k, and v are input to the Multi-Head Attention layer, and then undergo the second feature fusion with the result of the Concat operation. Then it is input to the feed-forward neural network FFN, and then undergoes the third feature fusion with the output of the second feature fusion. The obtained result is used as the second query q. The result of the first feature fusion is used as the second key k and the second value v. The second query q, the second key k, and the second value v are input to the second Multi-Head Attention layer, and then undergo the fourth feature fusion with the result of the third feature fusion. Then it is input to the feed-forward neural network FFN, and then undergoes the fifth feature fusion with the output of the fourth feature fusion. Then it is input to the MLP layer, and then undergoes the sixth feature fusion with the output of the fifth feature fusion, and then outputs.

[0017] Obtaining the motion capture result of the query video data includes:

[0018] ;

[0019] ;

[0020] Wherein, is the motion capture result of querying video data, is the first category probability distribution, is the second category probability distribution, is an adjustable hyperparameter.

[0021] Compared with the prior art, the present invention has the following beneficial effects: The present invention realizes efficient spatio-temporal feature extraction through a temporal enhancement adapter and a spatio-temporal fusion adapter, enhancing the spatio-temporal modeling ability of video features; using a multi-level attention mechanism, improving the fusion ability of text and video features, and constructing a class prototype with strong generalization ability; significantly improving the capture accuracy of the model in few-shot learning tasks, with fewer trainable parameters and reducing the computational cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a schematic structural diagram of a CLIP-ViT visual encoder;

[0023] Figure 2 is a schematic structural diagram of an even-numbered layer;

[0024] Figure 3 is a schematic structural diagram of an odd-numbered layer;

[0025] Figure 4 is a schematic structural diagram of a prototype construction module;

[0026] Figure 5 is the influence of hyperparameters on the accuracy of the 5-way 1-shot task on the Kinetics-400 (K400) dataset;

[0027] Figure 6 is the influence of hyperparameters on the accuracy of the 5-way 1-shot task on the SSv2-Small dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0029] Large Model-Driven Spatiotemporal Feature and Text-Augmented Few-Shot Action Capture Method, including obtaining video data and performing preprocessing. The video data includes query video data to be captured for actions and support set video data with action labels. The preprocessed video data is input into a visual encoder to obtain visual features of the video data;

[0030] The action labels of the support set are input into a CLIP text encoder to obtain text features of the action labels of the support set. The text features and the visual features of the query video data are subjected to similarity calculation to obtain a first category probability distribution. The visual features and the text features of the support set video data are input into a prototype construction module to obtain prototype features of the support set video data. The visual features of the query video data are input into the prototype construction module to obtain prototype features of the query video data. The prototype features of the support set video data and the prototype features of the query video data are subjected to time series similarity calculation to obtain a second category probability distribution;

[0031] Integrate the two category probability distributions to obtain the action capture result of the query video data.

[0032] Obtaining the visual features of the video data includes designing a temporal enhancement adapter and a spatiotemporal fusion adapter, and alternately embedding the two adapters into a CLIP-ViT visual encoder.

[0033] The CLIP-ViT visual encoder is as Figure 1 , including N VIT Block layers. In even layers, the temporal enhancement adapter is embedded before and after the multi-head attention layer respectively. In odd layers, the spatiotemporal fusion adapter is only embedded after the multi-head attention layer. A VIT Block is a visual block.

[0034] The even layer is as Figure 2 , and the input information is successively sent to a TE Adapter layer, a Layer Norm layer, a Multi-Head Attention layer, a feature fusion layer, a TE Adapter layer, a Layer Norm layer, an MLP layer, and a feature fusion layer, and then output. The output of the first TE Adapter layer is provided with a shortcut connection to the first feature fusion layer, and the output of the second TE Adapter layer is provided with a shortcut connection to the second feature fusion layer;

[0035] The TE Adapter layer successively includes an FC Down layer, a Gelu layer, an FC Up layer, and a feature fusion layer. The input of the TE Adapter layer is provided with a shortcut connection to the feature fusion layer;

[0036] The TE Adapter is a temporal enhancement adapter, Layer Norm is layer normalization, Multi-HeadAttention is multi-head attention, MLP is a multi-layer perceptron, FC Down is fully connected downsampling, FC Up is fully connected upsampling, and Gelu is an activation function.

[0037] The odd-numbered layers, such as Figure 3 , sequentially send the input information to the Layer Norm layer, the Multi-Head Attention layer, the feature fusion layer, the STF Adapter layer, the Layer Norm layer, the MLP layer, and the feature fusion layer, and then output; the input of the spatio-temporal fusion adapter is provided with a shortcut connection connecting to the first feature fusion layer, and the output of the STF Adapter layer is provided with a shortcut connection connecting to the second feature fusion layer;

[0038] The STF Adapter layer sequentially includes an FC Down layer, a Max-Pool layer, a 3D-Conv layer, a Sigmiod layer, a neuron layer, an FC Up layer, and a feature fusion layer. The input of the STF Adapter layer is provided with a shortcut connection connecting to the feature fusion layer, and the output of the FC Down layer is provided with a shortcut connection connecting to the neuron layer;

[0039] The STF Adapter is a spatio-temporal fusion adapter, Max-Pool is max pooling, 3D-Conv is three-dimensional convolution, and Sigmiod is an activation function.

[0040] The similarity calculation includes performing a max pooling operation on the visual features of the query video data, and then calculating the cosine similarity between the visual features and the text features of the query video data to obtain the first category probability distribution.

[0041] The prototype construction module, such as Figure 4, including performing a Concat operation on visual features (Video Features) and text features (TxextFeatures) as query q, where Concat is fusion. Input the text features into a repeat layer (Repeat), then perform the first feature fusion with the visual features as key k and value v. Input q, k, and v into a Multi-HeadAttention layer, then perform the second feature fusion with the result of the Concat operation, then input it into a feed-forward neural network (FFN), and then perform the third feature fusion with the output of the second feature fusion. The obtained result is used as the second query q, the result of the first feature fusion is used as the second key k and the second value v. Input the second query q, the second key k, and the second value v into the second Multi-Head Attention layer, then perform the fourth feature fusion with the result of the third feature fusion, then input it into a feed-forward neural network (FFN), and then perform the fifth feature fusion with the output of the fourth feature fusion, then input it into an MLP layer, and then perform the sixth feature fusion with the output of the fifth feature fusion, and then output.

[0042] Obtaining the motion capture result of the query video data includes:

[0043] ;

[0044] ;

[0045] Wherein, is the motion capture result of the query video data, is the first category probability distribution, is the second category probability distribution, is an adjustable hyperparameter.

[0046] Tables 1 and 2 provide a detailed comparative experiment between the present invention and the few-shot motion capture method FSAR. The pre-trained ViT-B / 16 model is mainly used, and a fair comparison is made with other methods using large-scale pre-trained models to comprehensively evaluate the performance and advantages of the present invention. In Table 1, the complete visual backbone network is fine-tuned as a whole, and parameter-efficient fine-tuning is performed on the visual backbone network. In Table 2, the complete visual backbone network is fine-tuned as a whole, and parameter-efficient fine-tuning is performed on the visual backbone network.

[0047] Table 1. Effects of each method on the time-series related dataset (unit: %)

[0048] ;

[0049] In Table 1, the "Method" column lists various existing motion capture methods. The pre-trained models used are the existing INet-RN50 and the CLIP ViT-B / 16 of the present invention. SSv2-Small and SSv2-Full are two datasets, and 1-shot and 5-shot are two training tasks.

[0050] Table 2. Effects of Each Method on Spatially Related Datasets (Unit: %)

[0051] ;

[0052] In Table 2, the "Method" column lists various existing motion capture methods. The pre-trained models used are the existing INet-RN50 and the CLIP ViT-B / 16 of the present invention. HMDB51, UCF101, and Kinetics are three datasets, and 1-shot and 5-shot are two training tasks.

[0053] From the results of the time modeling related datasets shown in Table 1, the present invention far exceeds other methods on the SSv2-Small dataset. Specifically, in the 1-shot task, the present invention is 1.3% higher than MA-CLIP, and in the 5-shot task, it is 2.2% higher. This shows that the present invention not only ensures the efficiency of parameters but also outperforms other methods based on parameter-efficient fine-tuning (PEFT) in terms of performance. For the SSv2-Full dataset, the performance of the present invention in the 1-shot task is comparable to that of the full fine-tuning CLIP-FSAR method. In addition, in the 5-shot task, the present invention also shows competitiveness. For datasets dominated by spatial information, such as HMDB51, UCF101, and Kinetics, the results shown in Table 2 indicate that the present invention can achieve performance comparable to that of the full fine-tuning method. Although the present invention is slightly inferior to the full fine-tuning method in some tasks, its advantage lies in parameter efficiency. Different from the full fine-tuning method that needs to optimize the entire model, the present invention only fine-tunes a small number of parameters, thus effectively reducing the consumption of computing resources while maintaining performance comparable to that of the full fine-tuning model.

[0054] To evaluate the impact of each module in the framework, ablation experiments were conducted under the 5-way 1-shot task of the SSv2-Small dataset. The experimental results are summarized in Table 3. Starting from the baseline model with a frozen backbone network and no learnable modules, the accuracy was 38.0%. When three modules - the Temporal Enhancement Adaptation Module (TEA), the Spatio-Temporal Fusion Adaptation Module (STFA), and the Text Enhancement Prototype Module (TEPM) were introduced separately, the performance was improved by 15.4%, 14.2%, and 14.1% respectively. When combining any two modules, the performance was further improved. The accuracy of the combination of TEA and TEPM reached 59.2%, indicating that temporal modeling and prototype construction have complementary advantages in the capture task. Finally, when all three modules were introduced, the model achieved the best performance of 61.4%, a 23.4% improvement compared to the baseline. These results verified the effectiveness of each module and showed that their synergy played a key role in the overall performance improvement.

[0055] Table 3. Performance comparison under different settings of the 5-way 1-shot task on the SSv2-Small dataset

[0056] ;

[0057] Adjust the hyperparameters to control and the relative contributions between, and analyze their impact on the final classification accuracy. As Figure 5 shown, in the SSv2-Small dataset dominated by temporal information, as the hyperparameter increases from 0, the accuracy rises rapidly and reaches a peak of 61.5% when the hyperparameter is 0.5 and the hyperparameter is 0.6. When the hyperparameter is further increased to 1, the accuracy drops sharply. This indicates that in tasks dominated by temporal features, the balanced contribution of the two probabilities is optimal, and overemphasizing a certain component will lead to a performance decline. In contrast, in the Figure 6 K400 dataset dominated by spatial information shown, the accuracy gradually improves as the hyperparameter increases, reaching the highest value of 93.8% when the hyperparameter is 0.7, and then showing a slight decline. This shows that in tasks mainly characterized by spatial information, increasing the weight of the video-text matching probability significantly contributes to performance improvement, but excessive focus on a certain probability may lead to a slight performance decline. For the sake of consistency, all experiments in this invention use a hyperparameter of 0.6.

[0058] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A large model-driven spatiotemporal feature and text-enhanced few-sample motion capture method, characterized in that: The method includes obtaining and preprocessing video data, wherein the video data includes query video data to be motion captured and support set video data with action labels, and inputting the preprocessed video data into a visual encoder to obtain visual features of the video data; Input the action labels of the support set into the CLIP text encoder to obtain the text features of the action labels of the support set, perform similarity calculation on the text features and the visual features of the query video data, and obtain the first category probability distribution; input the visual features and text features of the support set video data into the prototype construction module to obtain the prototype features of the support set video data, input the visual features of the query video data into the prototype construction module to obtain the prototype features of the query video data, perform time series similarity calculation on the prototype features of the support set video data and the prototype features of the query video data, and obtain the second category probability distribution; Combining the two category probability distributions, we can obtain the motion capture results of the query video data. The prototype construction module includes performing a Concat operation on the visual feature Video Features and the text feature Txext Features as a query q, wherein the Concat is a fusion, inputting the text feature into the repeating layer Repeat, and then performing the first feature fusion with the visual feature as the key k and the value v, inputting q, k, and v into the Multi-Head Attention layer, and then performing a second feature fusion with the result of the Concat operation, and then inputting the feedforward neural network FFN, and then performing a third feature fusion with the output of the second feature fusion, and the result obtained is used as the second query q, and the result of the first feature fusion is used as the second key k and the second value v, and the second query q, the second key k, and the second value v are input into the second Multi-Head Attention layer, and then performing a fourth feature fusion with the result of the third feature fusion, and then inputting the feedforward neural network FFN, and then performing a fifth feature fusion with the output of the fourth feature fusion, and then inputting the MLP layer, and then performing a sixth feature fusion with the output of the fifth feature fusion, and then outputting.

2. The large model driven spatiotemporal feature and text enhanced few-sample motion capture method according to claim 1, characterized in that: Obtaining the visual features of video data includes designing a temporal enhancement adapter and a spatiotemporal fusion adapter, and alternately embedding the two adapters into the CLIP-ViT visual encoder.

3. The large model driven spatiotemporal feature and text enhanced few-sample motion capture method according to claim 2, characterized in that: The CLIP-ViT visual encoder consists of N VIT Block layers. In the even-numbered layers, the temporal enhancement adapter is embedded before and after the multi-head attention layer. In the odd-numbered layers, the spatiotemporal fusion adapter is embedded only after the multi-head attention layer. VITBlock is the visual block.

4. The large model driven spatiotemporal feature and text enhanced few-sample motion capture method according to claim 3, characterized in that: The even-numbered layers sequentially send input information to the TE Adapter layer, the Layer Norm layer, the Multi-HeadAttention layer, the feature fusion layer, the TE Adapter layer, the Layer Norm layer, the MLP layer, the feature fusion layer, and then output; the output of the first TE Adapter layer is provided with a shortcut connection to the first feature fusion layer, and the output of the second TE Adapter layer is provided with a shortcut connection to the second feature fusion layer; The TE Adapter layer includes an FC Down layer, a Gelu layer, an FC Up layer and a feature fusion layer in sequence, and the input of the TE Adapter layer is provided with a shortcut connection to the feature fusion layer; The TE Adapter is a timing enhancement adapter, Layer Norm is layer normalization, Multi-Head Attention is multi-head attention, MLP is a multi-layer perceptron, FC Down is a fully connected downsampling, FC Up is a fully connected upsampling, and Gelu is an activation function.

5. The large model driven spatiotemporal feature and text enhanced few-sample motion capture method according to claim 4, characterized in that: The odd-numbered layers send the input information to the Layer Norm layer, the Multi-Head Attention layer, the feature fusion layer, the STF Adapter layer, the Layer Norm layer, the MLP layer, the feature fusion layer in sequence, and then output it; the input of the spatiotemporal fusion adapter is provided with a shortcut connection to the first feature fusion layer, and the output of the STF Adapter layer is provided with a shortcut connection to the second feature fusion layer; The STF Adapter layer includes an FC Down layer, a Max-Pool layer, a 3D-Conv layer, a Sigmiod layer, a neuron layer, an FC Up layer and a feature fusion layer in sequence. The input of the STF Adapter layer is provided with a quick connection connected to the feature fusion layer, and the output of the FCDown layer is provided with a quick connection connected to the neuron layer. The STF Adapter is a spatiotemporal fusion adapter, Max-Pool is maximum pooling, 3D-Conv is a three-dimensional convolution, and Sigmiod is an activation function.

6. The large model driven spatiotemporal feature and text enhanced few-sample motion capture method according to claim 5, characterized in that: The similarity calculation includes performing a maximum pooling operation on the visual features of the query video data, and then performing a cosine similarity calculation on the visual features and text features of the query video data to obtain a first category probability distribution.

7. The large model driven spatiotemporal feature and text enhanced few-sample motion capture method according to claim 6, characterized in that: The motion capture results of the query video data include: ; ; In the formula, is the motion capture result of querying video data, is the first category probability distribution, is the second category probability distribution, is a tunable hyperparameter.

Citation Information

Patent Citations

  • Method and system for judging dribbling foul action based on AI

    CN118155292A

  • Video question-answering method and system based on keyword perception multi-modal attention

    WO2023035610A1