A Zero-Shot Action Recognition Method Based on Hybrid Contrastive Learning

By optimizing and aligning video and text features in the public feature space, the problem of limited performance of existing zero-sample action recognition methods is solved, and more efficient recognition performance and broader application prospects are achieved.

CN115050096BActive Publication Date: 2025-06-17NANJING UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210629941.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-06
Publication Date
2025-06-17
Estimated Expiration
2042-06-06

AI Technical Summary

Technical Problem

The existing zero-sample action recognition method relies on existing video features and label word vector representations, and the data stream structure obtained by visual features and text features is inconsistent, resulting in limited performance.

Method used

The end-to-end zero-sample action recognition framework based on hybrid contrast learning is adopted to optimize and align different modal features in the common feature space, and improve recognition performance through positive example sample generation methods of video and text.

Benefits of technology

By optimizing and aligning different modal features in the public feature space, the performance of zero-sample action recognition is improved, and a wider application prospect and practical value are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115050096B_ABST
    Figure CN115050096B_ABST
Patent Text Reader

Abstract

The present invention discloses a zero-shot action recognition method based on hybrid contrastive learning, and the steps are as follows: 1. First, randomly divide the video categories; 2. Generate positive example samples for the videos and the text descriptions of the categories respectively; 3. Preprocess the videos to obtain the spatio-temporal features of the videos; 4. Learn the video embedding features related to the task; 5. Learn the text embedding features related to the task; 6. Optimize the embedding features of the videos and the text simultaneously; 7. Align the embedding features of the videos and the text in the common feature space; 8. Train the model; 9. Test the model. The present invention proposes an end-to-end zero-shot action recognition framework and a zero-shot action recognition method based on hybrid contrastive learning, and realizes the generalization of information between different modalities by simultaneously optimizing and aligning different modality features in the common feature space. To specifically achieve this goal, the present invention also proposes a method for generating positive example samples for videos and text respectively for this hybrid contrastive learning method, and finally improves the performance of zero-shot action recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention designs a zero - shot action recognition method based on hybrid contrast learning, especially an end - to - end zero - shot action recognition framework, a method for simultaneously optimizing and aligning different - modality features in a common feature space, and a method for generating positive - example samples for videos and texts respectively. Background Art

[0002] Thanks to the development of deep learning, especially the development of 3D convolutional neural networks, traditional human action recognition tasks have been widely studied in the field of computer vision and achieved great success. However, as the application scope of human action recognition becomes wider and the amount of video data increases day by day, the limitations of this task are gradually revealed. Traditional action recognition tasks rely on a large amount of labeled data covering all categories, bringing extremely high video collection and manual annotation costs. Therefore, the new task of zero - shot action recognition, which aims to use the learned knowledge to recognize unseen classes, has been proposed. Specifically, in this task of zero - shot action recognition, only videos belonging to visible classes and their label information can be accessed during training, but it is required to recognize the video categories belonging to unseen classes during testing. Compared with traditional human action recognition tasks, it has a broader application prospect and practical value in reality.

[0003] Currently, most zero - shot action recognition methods are based on existing video features and label word vector representations, and establish a mapping relationship between them or explore a generative method. However, the features used by these methods are not specific to the zero - shot contrast learning task, and since visual features and text features are often obtained based on different networks and representation spaces, the manifold structures of their data are also inconsistent, resulting in a great limitation on the performance of the method. Summary of the Invention

[0004] To solve the above problems, the present invention proposes a zero - shot action recognition method based on hybrid contrast learning, based on an end - to - end zero - shot action recognition framework, and realizes the generalization of information between different modalities by simultaneously optimizing and aligning different - modality features in a common feature space. To specifically achieve this goal, the present invention also proposes methods for generating positive - example samples for videos and texts respectively for this hybrid contrast learning method, and finally improves the performance of zero - shot action recognition.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] A zero - shot action recognition method based on hybrid contrast learning, characterized in that the method comprises the following steps:

[0007] Step (1): Obtain a benchmark dataset including video data and corresponding label description statements, and divide it into a training set and a test set;

[0008] Step (2): For the video data in step (1), generate positive samples of video features by randomly selecting video segments at different frame rates; for the corresponding label description statements, generate positive samples of text features by randomly masking several words in the label description statements;

[0009] Step (3): Construct an action recognition model, which includes a video processing module, a video feature extraction module, and a text feature extraction module;

[0010] In the video processing module, use the pre-trained network model r(2+1)d to extract spatio-temporal features from the segments of the input video randomly selected at different frame rates, and obtain the spatio-temporal features of each segment;

[0011] In the video feature extraction module, according to the spatio-temporal features of each segment, output the video embedding features of each segment;

[0012] In the text feature extraction module, use the pre-trained language model BERT to extract features from the label description statements after randomly masking several words in the label description statements, and obtain the text embedding features;

[0013] Step (4): Based on the training set obtained in step (2), jointly train the video feature extraction module and the text feature extraction module;

[0014] Step (5): For the test set obtained in step (2), use the trained text feature extraction module to extract the text embedding features of the label description statements after randomly masking several words in the label description statements; sequentially input the segments of the video data randomly selected at different frame rates into the video processing module and the trained video feature extraction module to obtain the video embedding features of each segment, and take the average value of the video embedding features of each segment as the video embedding feature of the video data; based on the nearest neighbor classification method, compare the video embedding feature of the video data with the corresponding text embedding feature of the test set to predict its category.

[0015] Furthermore, the specific implementation of step (1) is as follows:

[0016] For a benchmark dataset with N categories, randomly divide it into a visible class set, that is, the training set S, and an unseen class set, that is, the test set U, each accounting for 50% according to the categories, where

[0017] Further, the video feature extraction module includes a position embedding module, a transformer module, a first fully connected layer, a first non-linear activation layer, a second fully connected layer, a second non-linear activation layer, a third fully connected layer, and a first L2 normalization layer in a cascaded BERT model.

[0018] Further, the text feature extraction module includes a pre-trained BERT language model, a fourth fully connected layer, a third non-linear activation layer, a fifth fully connected layer, a fourth non-linear activation layer, a sixth fully connected layer, and a second L2 normalization layer in a cascaded manner.

[0019] Further, in step (4), the video feature extraction module and the text feature extraction module are jointly trained, and the specific implementation is as follows:

[0020] Step 6-1: Let the video embedding features and the text embedding features respectively pass through a projection module to obtain video projection features and text projection features where P is the dimension of the projection feature space, C is the number of segments selected in a video, and M is the dimension of the embedding features of each segment;

[0021] Step 6-2: Apply the self-supervised contrastive learning loss function SimCLR to the video projection features and the text projection features respectively in their respective projection spaces.

[0022] Step 6-3: Apply the self-supervised contrastive learning loss function SimCLR to the video embedding features and the text embedding features simultaneously to achieve feature optimization and alignment in the common feature space.

[0023] Further, the projection module is composed of a seventh fully connected layer, a fifth non-linear activation layer, an eighth fully connected layer, and a third L2 normalization layer in a cascaded manner.

[0024] Further, the self-supervised contrastive learning loss function SimCLR in step 6-2 has the following formula:

[0025]

[0026] where (f, f + ) and (f, f - ) are positive example sample pairs and positive and negative example sample pairs respectively, is the set of negative samples, sim(·,·) is the similarity function, and τ is an adjustable temperature parameter.

[0027] Furthermore, the self-supervised contrastive learning loss function SimCLR in step 6-3 is as follows:

[0028]

[0029] where (V, T + ) are video and text sample pairs that are positive examples of each other, and (V, T - ) are video and text sample pairs that are negative examples of each other.

[0030] Furthermore, the total loss of jointly training the video feature extraction module and the text feature extraction module is as follows:

[0031] L = L v + L t + L vt (5)

[0032] where the loss function for optimizing the video embedding features is:

[0033]

[0034] The loss function for optimizing the text embedding features is:

[0035]

[0036] In the formula, and are the positive example sample pairs and the positive and negative example sample pairs of the video features respectively, is the set of negative samples of the video features; and are the positive example sample pairs and the positive and negative example sample pairs of the text features respectively, is the set of negative samples of the text features.

[0037] Furthermore, step (5) is specifically implemented as follows:

[0038] Use the trained text feature extraction module to extract the text embedding features where is the number of categories of unseen classes; through the trained video feature extraction module, obtain the video embedding features of each segment, and take the average of the video embedding features of each segment to obtain the video embedding feature of this video data The category of this video data is obtained by the nearest neighbor classification method:

[0039]

[0040] In the formula, y U represents this video data.

[0041] The beneficial effects of the present invention are as follows:

[0042] 1) The present invention proposes an end-to-end training framework applicable to zero-shot action recognition tasks;

[0043] 2) The present invention designs a hybrid contrastive learning loss to simultaneously optimize and align the feature representations of different modalities in a common feature space;

[0044] 3) The present invention proposes a method for generating positive example samples of videos and texts for this task. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a schematic diagram of the positive example generation module for videos and texts of the present invention, where (a) is a video and (b) is a text;

[0046] Figure 2 is a test flow chart of the present invention;

[0047] Figure 3 is an overall flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0049] As Figure 3 shown, a zero-shot action recognition method based on hybrid contrastive learning is as follows:

[0050] Step (1): Randomly divide the training set and the test set. For a benchmark dataset with N categories, randomly divide the N video categories into a visible class set, i.e., the training set, and an unseen class set, i.e., the test set, where

[0051] Step (2): Construct a method for generating positive example samples of videos and texts for this task, as shown in Figure 1 (a) and (b) in.

[0052] Step 2-1: For a video, according to the total number of frames of the video, randomly extract video segments at different frame rates to generate positive example samples of video features. The total number of frames of a video is F, and the required input number of frames of the video model is f. If then extract the spatio-temporal features of consecutive frames; if then there is The probability of extracting spatio-temporal features of consecutive frames is The probability of extracting spatio-temporal features with a one-frame interval; if Then there is The probability of extracting spatio-temporal features with a one-frame interval is The probability of extracting spatio-temporal features with a two-frame interval.

[0053] Step 2-2: For a sentence of label text description, according to the total number of words in the sentence, randomly mask several words to generate positive example samples of text features. The total number of words in a description sentence is W. If 0 < W ≤ 20, then randomly mask 1 word; if 20 < W ≤ 40, then randomly mask 1 or 2 words with the same probability; if 40 < W ≤ 60, then randomly mask 1, 2 or 3 words with the same probability; if W > 60, then randomly mask 1, 2, 3 or 4 words with the same probability.

[0054] Step (3): For the original video v, first randomly extract video segments at different frame rates, and then use the r(2+1)d model pre-trained on the kinetics dataset to extract the spatio-temporal features of the video segments where C is the number of segments selected in a video, and M′ = 512 is the feature dimension of each segment.

[0055] Step (4): Construct a video feature extraction module, which includes a position embedding model in a standard BERT model, a transformer module, a fully connected layer with 2048 nodes, a non-linear activation layer, a fully connected layer, a non-linear activation layer, a fully connected layer, and an L2 normalization layer applied to the feature dimension. Pass the above spatio-temporal features through this video feature extraction module to obtain video embedding features where M = 2048 is the embedding feature dimension of each segment.

[0056] Step (5): Construct a text feature extraction module, which includes a BERT language model pre-trained on a large corpus, a fully connected layer with 2048 nodes, a non-linear activation layer, a fully connected layer, a non-linear activation layer, a fully connected layer, and an L2 normalization layer applied to the feature dimension. Pass the text description sentences of the categories through this text feature extraction module to obtain text embedding features where C is the number of mutually positive example samples generated from one sentence, and M = 2048 is the text embedding feature dimension of each sample.

[0057] Step (6): Model training

[0058] During the training process, the video feature extraction module and the text feature extraction module are jointly trained end-to-end, using a mixed self-supervised contrastive learning loss to optimize different modal feature spaces and achieve feature optimization and alignment of different modal features in the common feature space.

[0059] Step 6-1: Let the video embedding feature and the text embedding feature respectively pass through a projection module to obtain the video projection feature and the text projection feature where P = 512 is the dimension of the projection feature space. The projection module consists of a fully connected layer with 1024 nodes, a non-linear activation layer, a fully connected layer with 512 nodes, and an L2 normalization layer applied to the feature dimension.

[0060] Step 6-2: Apply the self-supervised contrastive learning loss function SimCLR to the video projection feature and the text projection feature respectively in their respective projection spaces. The formula is as follows:

[0061]

[0062] where (f, f + ) and (f, f - ) are the positive example sample pairs and positive and negative example sample pairs respectively, is the set of negative samples, sim is the similarity function, and τ is an adjustable temperature parameter.

[0063] Among them, the loss function for optimizing the video embedding feature is:

[0064]

[0065] The loss function for optimizing the text embedding feature is:

[0066]

[0067] In the formula, and are the positive example sample pair and positive and negative example sample pair of the video feature respectively, is the set of negative samples of the video feature; and are the positive example sample pair and positive and negative example sample pair of the text feature respectively, is the set of negative samples of the text feature.

[0068] Step 6-3: At the same time, for the video embedding feature and the text embedding feature Apply the self-supervised contrastive learning loss function SimCLR to perform optimization and alignment operations on it simultaneously. The specific formula is as follows:

[0069]

[0070] Among them, (V, T + ) are video and text sample pairs that are positive examples of each other, and (V, T - ) are video and text sample pairs that are negative examples of each other.

[0071] Step 6-4: The total loss of the joint training is as follows:

[0072] L = L v + L t + L vt (5)

[0073] Step 7): Model testing, see Figure 2 .

[0074] Use the text feature extraction module in step (5) of the trained model to extract the representation features of the description text of the test category label in the common feature space where is the number of classes of unseen classes, and M = 2048 is the dimension of the common feature space. Pass the test video through the video feature extraction module in step (4) of the trained model to obtain its representation in the common feature space and take the average of the features of each segment to obtain the final feature representation of the video The test category label is obtained through the nearest neighbor classification method:

[0075]

[0076] In this article, specific examples are used to elaborate on the principles and implementation methods of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A zero-shot action recognition method based on hybrid contrast learning, characterized in that, The method includes the following steps: Step (1): Obtain a benchmark dataset including video data and corresponding label description statements, and divide it into a training set and a test set; Step (2): For the video data in step (1), generate positive samples of video features by randomly selecting video segments at different frame rates; for the corresponding label description statements, generate positive samples of text features by randomly masking several words in the label description statements; Step (3): Construct an action recognition model, which includes a video processing module, a video feature extraction module, and a text feature extraction module; In the video processing module, use the pre-trained network model r(2+1)d to extract spatio-temporal features from the segments of the input video randomly selected at different frame rates, and obtain the spatio-temporal features of each segment; In the video feature extraction module, output the video embedding features of each segment according to the spatio-temporal features of each segment; In the text feature extraction module, use the pre-trained language model BERT to extract features from the label description statements after randomly masking several words in the label description statements, and obtain text embedding features; Step (4): Based on the training set obtained in step (2), jointly train the video feature extraction module and the text feature extraction module; Step (5): For the test set obtained in step (2), use the trained text feature extraction module to extract the text embedding features of the label description statements after randomly masking several words in the label description statements; sequentially input the segments of the video data randomly selected at different frame rates into the video processing module and the trained video feature extraction module to obtain the video embedding features of each segment, and take the average value of the video embedding features of each segment as the video embedding feature of the video data; based on the nearest neighbor classification method, compare the video embedding feature of the video data with the corresponding text embedding feature of the test set to predict its category.

2. The zero-shot action recognition method based on hybrid contrast learning according to claim 1, characterized in that, The specific implementation of step (1) is as follows: For a benchmark dataset with N categories, it is randomly divided into a visible class set, i.e., the training set S, and an unseen class set, i.e., the test set U, each accounting for 50% according to the categories, where 3. The zero-shot action recognition method based on hybrid contrast learning according to claim 1, characterized in that, The video feature extraction module includes a position embedding module, a transformer module, a first fully connected layer, a first non-linear activation layer, a second fully connected layer, a second non-linear activation layer, a third fully connected layer, and a first L2 normalization layer in the cascaded BERT model.

4. The zero-shot action recognition method based on hybrid contrast learning according to claim 1, characterized in that, The text feature extraction module includes a cascaded pre-trained BERT language model, a fourth fully connected layer, a third non-linear activation layer, a fifth fully connected layer, a fourth non-linear activation layer, a sixth fully connected layer, and a second L2 normalization layer.

5. The zero-shot action recognition method based on hybrid contrast learning according to claim 1, characterized in that, In step (4), the specific implementation of jointly training the video feature extraction module and the text feature extraction module is as follows: Step 6-1: Make the video embedding feature and the text embedding feature respectively pass through a projection module to obtain the video projection feature and the text projection feature where P is the dimension of the projection feature space, C is the number of segments selected in a video, and M is the embedding feature dimension of each segment; Step 6-2: Apply the self-supervised contrastive learning loss function SimCLR to the video projection features and the text projection features respectively in their respective projection spaces; Step 6-3: Apply the self-supervised contrastive learning loss function SimCLR to the video embedding features and the text embedding features simultaneously to achieve feature optimization and alignment in the common feature space.

6. The zero-shot action recognition method based on hybrid contrast learning according to claim 5, characterized in that, The projection module is cascaded by a seventh fully connected layer, a fifth non-linear activation layer, an eighth fully connected layer, and a third L2 normalization layer.

7. The zero-shot action recognition method based on hybrid contrast learning according to claim 5, characterized in that, The self-supervised contrastive learning loss function SimCLR in step 6-2 is as follows: where \((f,f + )\) and \((f,f - )\) are the positive example sample pair and the positive and negative example sample pair respectively, is the set of negative samples, sim(·,·) is the similarity function, and τ is an adjustable temperature parameter.

8. The zero-shot action recognition method based on hybrid contrast learning according to claim 5, characterized in that, The self-supervised contrastive learning loss function SimCLR in step 6-3 is as follows: where (V, T + ) are video and text sample pairs that are positive examples of each other, and (V, T - ) are video and text sample pairs that are negative examples of each other.

9. The zero-shot action recognition method based on hybrid contrast learning according to claim 5, characterized in that, The total loss of jointly training the video feature extraction module and the text feature extraction module is as follows: L = L v + L t + L vt (5) Among them, the loss function for optimizing the video embedding feature is: The loss function for optimizing the text embedding feature is: Wherein, and are the positive example sample pair and the positive and negative example sample pair of video features respectively, is the set of negative samples of video features; and are the positive example sample pair and the positive and negative example sample pair of text features respectively, is the set of negative samples of text features.

10. A zero-shot action recognition method based on hybrid contrastive learning according to claim 2, wherein, The specific implementation of step (5) is as follows: Extract the text embedding features using the trained text feature extraction module where is the number of classes of unseen classes; through the trained video feature extraction module, obtain the video embedding features of each segment, and take the average of the video embedding features of each segment to obtain the video embedding feature of the video data The class of the video data is obtained by the nearest neighbor classification method: where y I represents the video data.

Citation Information

Patent Citations

  • Zero-sample text classification method based on label extension

    CN113723106A