Small sample action recognition method based on multi-granularity spatial-temporal feature enhancement

Through the multi-grained spatiotemporal feature enhancement method and VMamba network architecture, data sparsity and generalization problems in small sample action recognition are solved, and efficient video action recognition is achieved.

CN120543931APending Publication Date: 2025-08-26LANZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510647550.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing small sample action recognition methods have low overfitting and generalization of the model under the limitation of data sparseness, making it difficult to achieve efficient recognition.

Method used

Using multi-grained spatial and temporal feature enhancement method, through data set division, video frame segmentation, feature extraction, multi-grained spatial coding and timing encoding, combined with the VMamba network architecture, action prototypes are generated and similarity matching is performed, and the model is optimized using loss function.

Benefits of technology

It improves the accuracy of small sample video action recognition, overcomes the problems of sample sparseness limitation and low generalization, and improves the recognition performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543931A_ABST
    Figure CN120543931A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample action recognition method based on multi-granularity spatial-temporal feature enhancement, and the method comprises the steps: segmenting video data, and randomly segmenting a video into a plurality of video frames through FFMPEG; feature extraction: extracting visual features of the video frames by using a pre-trained neural network; video feature multi-granularity space coding: coding the video features by using a block-level coding module, a window-level coding module and a frame-level coding module, and extracting static visual information in the video; video feature time sequence encoding: encoding by using a Vmamba-based time sequence encoding module, learning time sequence information of actions in the video, and generating a feature prototype of each action category; combining prototype matching, and calculating the matching degree of the queried video and the supported video prototype, so as to obtain the action category of the queried video. According to the method, the inherent problems of sample sparsity limitation, model over-fitting and low generalization of a small sample video action recognition task are solved, and the accuracy of video action recognition is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a small sample action recognition method based on multi-granularity spatiotemporal feature enhancement. Background Art

[0002] Small-sample action recognition is a challenging research area in computer vision. Traditional action recognition methods typically rely on large amounts of labeled data to train models. However, in practical applications, obtaining sufficient labeled samples is often difficult. Therefore, achieving efficient action recognition with limited data has become a pressing challenge. In recent years, multi-granularity models have gradually emerged in computer vision tasks, integrating feature information at different granularities to improve model accuracy and efficiency. Furthermore, with the continuous development of deep learning technology, various new network architectures and algorithms have emerged, providing new approaches for small-sample action recognition.

[0003] VMamba, a novel network architecture, is based on a stack of Visual State Space (VSS) blocks with a 2D Selective Sweep (SS2D) module. It boasts linear time complexity and excels in a variety of visual perception tasks. Therefore, combining multi-granularity spatial and VMamba temporal enhancement techniques for small-sample action recognition is expected to improve both the model's recognition performance and generalization capabilities. Summary of the Invention

[0004] In order to overcome the inherent sample sparsity limitations, model overfitting and low generalization problems of existing small-sample action recognition methods, the present invention proposes a small-sample action recognition method based on multi-granularity spatiotemporal feature enhancement to effectively improve the accuracy of video action recognition.

[0005] A small sample action recognition method based on multi-granularity spatiotemporal feature enhancement includes the following steps:

[0006] S1: Dataset partitioning: Based on the data partitioning strategy of the small sample learning framework, the dataset is divided into a training set and a test set, where no coexisting action types exist in the training set and the test set. The training set and the test set are then divided into a support sample set and a query sample set, respectively. The support sample set and the query sample set are input simultaneously in each round of training and testing the model.

[0007] S2: Video data segmentation: For the input support sample set and query sample set video data, FFMPEG is used to randomly cut the video data into multiple video frames, and a sparse sampling strategy is used to filter some representative key frames from the frame set to obtain the support sample set and query sample set video frame sequence sets;

[0008] S3: Feature extraction, using a neural network pre-trained on a large-scale dataset to extract visual features of the video frame to obtain support features and query features. The calculation of feature extraction is shown in formula (1) and formula (2):

[0009]

[0010] in, and F q To support the features of the concentrated c-type action videos and the query video features, δ(g) and It is the normalization function and feature extraction function;

[0011] S4: Multi-granularity spatial coding of video features, encoding the visual features using block-level, window-level, and frame-level coding modules, including the following steps:

[0012] S41: Block-level spatial feature encoding, the video frame feature map output in step S3 and F q Divide each into multiple blocks of the same size, and then flatten the set into a feature vector along the width or height direction;

[0013] S42: Splice a randomly initialized learnable token T at the beginning or end of the block feature vector p , the token is used to obtain the patch set features of the support set and query set in the spatial information of the video frame

[0014] S43: In A position code is attached to record the relative position of each element in the patch set;

[0015] S44: Input the features output from step S43 into the block-level encoding module based on the attention mechanism for encoding, so as to achieve fine-grained learning of video information. The calculation of the block-level encoding module is shown in formula (3) to formula (5).

[0016]

[0017] Among them, η(g) is the softmax function, Att(g) and Matt(g) are the self-attention mechanism and multi-head self-attention mechanism functions, is the self-attention head function, W is the learnable weight matrix, and then the token set is taken from the output feature map

[0018] S45, window-level spatial feature encoding, The adjacent m patches in the image are combined into a window to obtain multiple window-level features. Concatenate a randomly initialized learnable token T for each window w , used to learn the spatial information within the window and avoid the fine-grained division of the patch level from destroying the integrity of the action subject; The input window-level encoding module encodes each window and learns the target information within the window area; the window position is shifted to the lower right. pixels, reorganize the window features and input them into the window-level encoding module to encode each window and learn the spatial information between the original windows; extract the Token set from the output features

[0019] S46: Frame-level spatial feature encoding, concatenating learnable tokens and inputting them into the window-level encoding module to learn the original window information; After the pooling operation, the frame-level feature encoding module is directly input for encoding to extract the global information of the video frame and obtain the window-level feature token set. According to the token set T at the block level, window level and frame level p 、T w and T f , constituting the feature token set of each frame of the video, combining all frame features of each video to form a three-dimensional feature including channel dimension, spatial dimension and time dimension;

[0020] S47: Video feature temporal coding, using the Vmamba-based temporal coding module to encode in the horizontal and vertical dimensions, learning the intra-frame spatial context relationship and inter-frame temporal dependency relationship of the video, including the following steps:

[0021] S471: Encode the video feature token set output in step S4 from the horizontal dimension, i.e., the time dimension, to extract temporal motion information between video frames;

[0022] S472: Encode the video feature token set output in step S4 from the vertical dimension, i.e., the spatial dimension, to extract the relative position relationship between the objects in the video frame;

[0023] S473: Generate a prototype that supports videos of different action types in the set;

[0024] S5: Prototype matching, measuring the similarity of window-level, block-level, and frame-level prototypes respectively, combining them to generate a relationship matrix, and taking the action type with the smallest distance from the support set as the type of the query video. The specific calculation is shown in formulas (6) and (7);

[0025]

[0026]

[0027] Among them, ρ(g) is the cosine similarity function, O q and are the prototypes of query samples and samples of category c in support samples, respectively. |C| is the number of labels in each training round, r c O q and The matching probability.

[0028] S6: Identify the action category of the query video and use the annotation labels to construct a loss function, including patch-level loss and window-level loss. The specific calculation formula is shown in (8).

[0029]

[0030] in, and is the block-level and window-level loss, CE(g) is the cross entropy loss function, σ(g) is the softmax activation function, W p and W w is the learnable weight matrix, Y s is the real action label.

[0031] Furthermore, the large-scale dataset in step S3 is ImageNet, and the pre-trained neural network is Resnet.

[0032] Furthermore, the similarities of the window-level, block-level and frame-level prototypes in step S5 are cosine similarities.

[0033] The beneficial effect of the present invention is that, compared with the existing technology, the present invention overcomes the sample sparsity limitations inherent in small sample video action recognition tasks, the problems of model overfitting and low generalization, and effectively improves the accuracy of video action recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a flow chart of a small sample action recognition method based on multi-granularity spatiotemporal feature enhancement of the present invention;

[0035] Figure 2 This is a system structure diagram of a small sample action recognition method based on multi-granularity spatiotemporal feature enhancement of the present invention. DETAILED DESCRIPTION

[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0037] See Figure 1 and Figure 2 The present invention provides a small sample action recognition method based on multi-granularity spatiotemporal feature enhancement, comprising the steps of:

[0038] S1: Dataset partitioning: Based on the data partitioning strategy of the small sample learning framework, the dataset is divided into a training set and a test set, where no coexisting action types exist in the training set and the test set. The training set and the test set are then divided into a support sample set and a query sample set, respectively. The support sample set and the query sample set are input simultaneously in each round of training and testing the model.

[0039] S2: Video data segmentation: For the input support sample set and query sample set video data, FFMPEG is used to randomly cut the video data into multiple video frames, and a sparse sampling strategy is used to filter some representative key frames from the frame set to obtain the support sample set and query sample set video frame sequence sets;

[0040] S3: Feature extraction, using a neural network pre-trained on a large-scale dataset to extract visual features of the video frame to obtain support features and query features. The calculation of feature extraction is shown in formula (1) and formula (2):

[0041]

[0042] in, and F q To support the features of the concentrated c-type action videos and the query video features, δ(g) and It is the normalization function and feature extraction function;

[0043] S4: Multi-granularity spatial coding of video features, encoding the visual features using block-level, window-level, and frame-level coding modules, including the following steps:

[0044] S41: Block-level spatial feature encoding, the video frame feature map output in step S3 and F q Divide each into multiple blocks of the same size, and then flatten the set into a feature vector along the width or height direction;

[0045] S42: Splice a randomly initialized learnable token T at the beginning or end of the block feature vector p , the token is used to obtain the patch set features of the support set and query set in the spatial information of the video frame

[0046] S43: In A position code is attached to record the relative position of each element in the patch set;

[0047] S44: Input the features output from step S43 into the block-level encoding module based on the attention mechanism for encoding, so as to achieve fine-grained learning of video information. The calculation of the block-level encoding module is shown in formula (3) to formula (5).

[0048]

[0049] Among them, η(g) is the softmax function, Att(g) and Matt(g) are the self-attention mechanism and multi-head self-attention mechanism functions, is the self-attention head function, W is the learnable weight matrix, and then the token set is taken from the output feature map

[0050] S45, window-level spatial feature encoding, The adjacent m patches in the image are combined into a window to obtain multiple window-level features. Concatenate a randomly initialized learnable token T for each window w , used to learn the spatial information within the window and avoid the fine-grained division of the patch level from destroying the integrity of the action subject; The input window-level encoding module encodes each window and learns the target information within the window area; the window position is shifted to the lower right. pixels, reorganize the window features and input them into the window-level encoding module to encode each window and learn the spatial information between the original windows; extract the Token set from the output features

[0051] S46: Frame-level spatial feature encoding, concatenating learnable tokens and inputting them into the window-level encoding module to learn the original window information; After the pooling operation, the frame-level feature encoding module is directly input for encoding to extract the global information of the video frame and obtain the window-level feature token set. According to the token set T at the block level, window level and frame level p 、T w and T f , constituting the feature token set of each frame of the video, combining all frame features of each video to form a three-dimensional feature including channel dimension, spatial dimension and time dimension;

[0052] S47: Video feature temporal coding, using the Vmamba-based temporal coding module to encode in the horizontal and vertical dimensions, learning the intra-frame spatial context relationship and inter-frame temporal dependency relationship of the video, including the following steps:

[0053] S471: Encode the video feature token set output in step S4 from the horizontal dimension, i.e., the time dimension, to extract temporal motion information between video frames;

[0054] S472: Encode the video feature token set output in step S4 from the vertical dimension, i.e., the spatial dimension, to extract the relative position relationship between the objects in the video frame;

[0055] S473: Generate a prototype that supports videos of different action types in the set;

[0056] S5: Prototype matching, measuring the similarity of window-level, block-level, and frame-level prototypes respectively, combining them to generate a relationship matrix, and taking the action type with the smallest distance from the support set as the type of the query video. The specific calculation is shown in formulas (6) and (7);

[0057]

[0058] Among them, ρ(g) is the cosine similarity function, O q and are the prototypes of query samples and samples of category c in support samples, respectively. |C| is the number of labels in each training round, r c O q and The matching probability.

[0059] S6: Identify the action category of the query video and use the annotation labels to construct a loss function, including patch-level loss and window-level loss. The specific calculation formula is shown in (8).

[0060]

[0061] in, and is the block-level and window-level loss, CE(g) is the cross entropy loss function, σ(g) is the softmax activation function, W p and W w is the learnable weight matrix, Y s is the real action label.

[0062] In step S3, the large-scale dataset is ImageNet, and the pre-trained neural network is Resnet.

[0063] The similarity of the window-level, block-level, and frame-level prototypes in step S5 is cosine similarity. The above is a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A small sample action recognition method based on multi-granularity spatiotemporal feature enhancement, characterized by: Including steps: S1: Dataset partitioning: Based on the data partitioning strategy of the small sample learning framework, the dataset is divided into a training set and a test set, where no coexisting action types exist in the training set and the test set. The training set and the test set are then divided into a support sample set and a query sample set, respectively. The support sample set and the query sample set are input simultaneously in each round of training and testing the model. S2: Video data segmentation: For the input support sample set and query sample set video data, FFMPEG is used to randomly cut the video data into multiple video frames, and a sparse sampling strategy is used to filter some representative key frames from the frame set to obtain the support sample set and query sample set video frame sequence sets; S3: Feature extraction, using a neural network pre-trained on a large-scale dataset to extract visual features of the video frame to obtain support features and query features. The calculation of feature extraction is shown in formula (1) and formula (2): in, and F q To support the features of the concentrated c-type action videos and the query video features, δ(g) and It is the normalization function and feature extraction function; S4: Multi-granularity spatial coding of video features, encoding the visual features using block-level, window-level, and frame-level coding modules, including the following steps: S41: Block-level spatial feature encoding, the video frame feature map F output in step S3 c s and F q Divide each into multiple blocks of the same size, and then flatten the set into a feature vector along the width or height direction; S42: Splice a randomly initialized learnable token T at the beginning or end of the block feature vector p , the token is used to obtain the patch set features of the support set and query set in the spatial information of the video frame S43: In A position code is added to record the relative position of each element in the patch set; S44: The features output from step S43 are input into the block-level encoding module based on the attention mechanism for encoding, so as to achieve fine-grained learning of video information. The calculation of the block-level encoding module is shown in formula (3) to formula (5): Among them, η(g) is the softmax function, Att(g) and Matt(g) are the self-attention mechanism and multi-head self-attention mechanism functions, is the self-attention head function, W is the learnable weight matrix, and then the token set T is taken from the output feature map s p , S45, window-level spatial feature encoding, The adjacent m patches in the image are combined into a window to obtain multiple window-level features. Concatenate a randomly initialized learnable token T for each window w , used to learn the spatial information within the window and avoid the fine-grained division of the patch level from destroying the integrity of the action subject; The input window-level encoding module encodes each window and learns the target information within the window area; the window position is shifted to the lower right. pixels, reorganize the window features and input them into the window-level encoding module to encode each window and learn the spatial information between the original windows; extract the Token set T from the output features s p , S46: Frame-level spatial feature encoding, concatenating learnable tokens and inputting them into the window-level encoding module to learn the original window information; After the pooling operation, the frame-level feature encoding module is directly input for encoding to extract the global information of the video frame and obtain the window-level feature token set T s f , According to the token set T at the block level, window level and frame level p 、T w and T f , constituting the feature token set of each frame of the video, combining all frame features of each video to form a three-dimensional feature including channel dimension, spatial dimension and time dimension; S47: Video feature temporal coding, using the Vmamba-based temporal coding module to encode in the horizontal and vertical dimensions, learning the intra-frame spatial context relationship and inter-frame temporal dependency relationship of the video, including the following steps: S471: For the video feature token set output in step S4, encode it from the horizontal dimension, i.e., the time dimension, to extract temporal motion information between video frames; S472: Encode the video feature token set output in step S4 from the vertical dimension, i.e., the spatial dimension, to extract the relative position relationship between the objects in the video frame; S473: Generate a prototype that supports videos of different action types in the set; S5: Prototype matching, which measures the similarity of window-level, block-level, and frame-level prototypes respectively, and combines them to generate a relationship matrix. The action type with the smallest distance in the support set is used as the type of the query video. The specific calculation is shown in formulas (6) and (7): Among them, ρ(g) is the cosine similarity function, O q and are the prototypes of query samples and the prototypes of category c samples in support samples, respectively. |C| is the number of labels in each training round, r c O q and The matching probability. S6: Identify the action category of the query video and use the annotation labels to construct a loss function, including patch-level loss and window-level loss. The specific calculation formula is shown in (8): in, and is the block-level and window-level loss, CE(g) is the cross entropy loss function, σ(g) is the softmax activation function, W p and W w is the learnable weight matrix, Y s is the real action label.

2. The small sample action recognition method based on multi-granularity spatiotemporal feature enhancement according to claim 1 is characterized in that: The large-scale dataset in step S3 is ImageNet, and the pre-trained neural network is Resnet.

3. The small sample action recognition method based on multi-granularity spatiotemporal feature enhancement according to claim 1 is characterized in that: The similarity of the window-level, block-level and frame-level prototypes in step S5 is cosine similarity.