A Semi-Supervised Short Video Classification Method Based on Semantic Consistency

The semi-supervised short video classification method enhances classification accuracy by leveraging intra-video frame consistency through a neural network with spatial and temporal attention, addressing the high annotation costs and resource demands of existing methods.

CN116340569BActive Publication Date: 2025-07-15TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310086713.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-09
Publication Date
2025-07-15
Estimated Expiration
2043-02-09

AI Technical Summary

Technical Problem

The existing deep learning methods require a large number of labels in short video classification, resulting in high training costs. However, the existing semi-supervised learning methods are less studied in the video field, making it difficult to effectively use semantic consistency information between video frames for efficient classification.

Method used

The semi-supervised learning method based on semantic consistency is adopted to learn video frame spatial features through a two-dimensional residual network, and the importance of frames is learned in the time dimension using the attention mechanism, and data enhancement and consistency loss optimization are carried out by constructing keyframe sequences to improve classification accuracy.

Benefits of technology

With a small number of labels, the inherent semantic information of the video is effectively mined, the classification accuracy of short videos is improved, the high computing cost of the three-dimensional convolutional network is avoided, and the model has higher robustness and classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116340569B_ABST
    Figure CN116340569B_ABST
Patent Text Reader

Abstract

The present invention discloses a semi-supervised short video classification method based on semantic consistency, including: utilizing the correlation of adjacent video frames, extracting key frames from labeled data and unlabeled data at equal time intervals T starting from frames t0 and t0+τ respectively to obtain two frame sequences, and performing strong data augmentation and standard data augmentation respectively; building a neural network composed of a spatial feature learning module, a temporal attention fusion module, and a classifier module; taking the same number of labeled samples and unlabeled samples after standard data augmentation, splicing them and inputting them into the neural network, and calculating the classification loss of the labeled part; taking the same number of labeled samples and unlabeled samples after strong data augmentation, splicing them and inputting them into the neural network to obtain a prediction output, and calculating the consistency loss for the prediction distributions obtained by inputting the same sample into the network after these two different data augmentation processes; jointly using the classification loss and the consistency loss for the optimization training of the neural network; inputting the video samples into the optimized neural network to output corresponding prediction scores, and obtaining the final video classification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of short video action and event classification, and particularly to a semi-supervised short video classification method based on semantic consistency. Background Art

[0002] With the popularization of mobile intelligent devices and the rapid development of network technology, short videos have gradually replaced traditional pictures and texts as important information carriers in people's daily lives. Compared with traditional videos, these short videos themselves have the characteristics of short duration, complex and diverse content, and fragmented information. As one of the sub-tasks of video content understanding, short video action classification mainly solves the understanding and modeling of human actions in short videos, and is the basis for understanding complex video events. Related representation learning also plays an important role in downstream applications such as video classification and recommendation, copyright protection, and content review. In recent years, thanks to the continuous improvement of the performance of computing devices and the rapid development and implementation of artificial intelligence-related technologies, deep learning has been widely used in the field of video content understanding. Some deep algorithms have gradually replaced traditional machine learning and can learn more complete deep representations of videos, and the specific performance is also more excellent. However, most methods of deep learning require a large number of labels for training, and the annotation of a large number of videos requires a large amount of labor cost. Therefore, semi-supervised learning that only requires a small number of labels for training has important practical significance.

[0003] In video analysis, deep methods are mainly divided into three categories: methods based on two-stream networks, methods based on 3D convolutional neural networks, and methods based on 2D convolutional neural networks. The method based on the optical flow network mainly learns temporal dimension information by extracting video optical flow features. The method based on the 3D convolutional network directly extends the convolutional kernel in the temporal dimension and synchronously learns spatio-temporal information through convolution. The method based on the 2D convolutional network uses two-dimensional convolution to learn the spatial features of video frames, and then mines temporal dimension information by performing temporal modeling on video frames. Although the first two methods have achieved good results, the calculation of optical flow and the calculation and training of three-dimensional networks require huge costs; the method of extracting key frame features based on the two-dimensional convolutional network needs to consider more the interaction and fusion of temporal dimension information.

[0004] Currently, the research on semi-supervised learning mainly focuses on the field of images, and there is relatively little actual research in the field of videos. The mainstream semi-supervised learning algorithms are mainly divided into two types: methods based on pseudo-labels and methods based on consistency regularization. The method based on pseudo-labels is committed to improving the confidence of generating pseudo-labels, while the method based on regularization optimizes learning by perturbation to promote the decision boundary to be set in the low-density region. The semi-supervised algorithms on images mainly mine the invariance of model predictions for the same picture under different sample augmentation perturbations. Since a video itself consists of a large number of video frames, there is a high degree of semantic correlation and interaction between video frames. Mining internal supervision signals in videos is the development direction of semi-supervised learning on videos.

[0005] Although these existing methods have achieved good results in video classification, video semi-supervised learning is still in the development stage. Therefore, it is meaningful to propose a short video learning method based on semi-supervised learning, deeply mine and utilize the semantic consistency information between frames in the video, and at the same time perform sufficient information interaction and fusion between video frames to extract high-order semantic information for classification. Summary of the Invention

[0006] The present invention provides a semi-supervised short video classification method based on semantic consistency. The input video of the present invention learns the spatial features of video frames through a two-dimensional residual network, and uses an attention mechanism to learn the importance of each video frame for the video in the time dimension, so as to obtain video-level high-level semantic information; by introducing semi-supervised learning, using the similarity characteristics of adjacent frames of short videos, two key frame sequences of the same short video are constructed, and while learning the supervision signal of labeled samples, the network output prediction distributions of these two frame sequences after being enhanced with two different intensities of data are constrained, so that the model learns more robust video-level semantics and improves the classification accuracy in the case of only a small number of labeled samples. See the following description for details:

[0007] A semi-supervised short video classification method based on semantic consistency, the method comprising:

[0008] Using the correlation of adjacent video frames, key frames are extracted from the labeled data and unlabeled data at equal time intervals T starting from t0 and t0 + τ frames respectively to obtain two frame sequences, and strong data augmentation and standard data augmentation are performed respectively;

[0009] Build a neural network composed of a spatial feature learning module, a temporal attention fusion module, and a classifier module;

[0010] The labeled samples and unlabeled samples after standard data augmentation are spliced with the same number and then input into the neural network to calculate the classification loss of the labeled part; the labeled samples and unlabeled samples after strong data augmentation are spliced with the same number and then input into the neural network, and the consistency loss is calculated for the prediction distribution output by the network;

[0011] The classification loss and the consistency loss are jointly used for the optimization training of the neural network; the video samples are input into the optimized neural network to output corresponding prediction scores, and the final video classification result is obtained.

[0012] Wherein, the building of the neural network composed of a spatial feature learning module, a temporal attention fusion module, and a classifier module is:

[0013] The spatial feature learning module uses a residual network for spatial feature encoding, and inputs the spatial feature sequence of video frames into the temporal attention fusion module;

[0014] The temporal attention fusion module concatenates the position information encoding in the learned video frame spatial feature sequence for learning the temporal relationships of each key frame, and at the same time adds the category information encoding for fusing the temporal semantic information of the entire sequence;

[0015] Learn the degree of correlation between key frames within the sequence, calculate the correlation matrix of each video frame in the sequence with all video frames in the sequence, and obtain a frame feature sequence with global attention information;

[0016] Use the frame feature sequence to average and fuse the video frame information to obtain the video-level semantic representation, and jointly input the semantic features and the learned category encoding information into the classifier module composed of fully connected layers to output the final classification result.

[0017] The so-called strong data augmentation and standard data augmentation are as follows:

[0018] Extract N frames at equal time intervals T starting from t0 and t0 + τ for all video samples to obtain two key frame sequence sets for each sample and

[0019] where, F i is the i-th frame image of the video sample. For the images in X w perform the same standard data augmentation operation, and perform multiple strong data augmentation operations on X s using the RandAugment algorithm.

[0020] The so-called adding the category information encoding for fusing the temporal semantic information of the entire sequence is as follows:

[0021] Input the sample X m into the residual network f(·) to learn the spatial appearance features of each frame and obtain the video feature matrix n is the number of frames, and d is the output feature dimension of the network f(·);

[0022] Concatenate a randomly initialized category information encoding bit at the first dimension position of the video feature matrix to obtain the feature matrix representing the video and then add the randomly initialized position encoding information to obtain the final feature matrix.

[0023] The classification loss is:

[0024] The labeled samples after standard data augmentation and the unlabeled samples Take the same number of spliced inputs and feed them into the neural network. The neural network outputs the predicted distribution and obtains the result. Use cross-entropy to calculate the error loss L between the predicted result of the labeled part and the actual label. cls 。

[0025] The consistency loss is as follows: Use the JS divergence to calculate the consistency loss L between the predicted distributions of the network for the same sample after two different data augmentation processes. cons ;

[0026]

[0027] Among them, KL represents the KL divergence, and the calculation method is as follows:

[0028]

[0029] Finally, the two losses are jointly used for the training and optimization of the network. The total loss is as follows:

[0030] L = L cls + λ·L cons

[0031] Among them, λ is an adjustable parameter.

[0032] The beneficial effects of the technical solution provided by the present invention are as follows:

[0033] 1. Different from supervised learning, introducing semi-supervised learning can utilize the characteristics of the data itself to mine deep features when only a small amount of data is labeled, improving the short video classification accuracy; this method utilizes the correlation between internal frames of the same video to construct key frame sequences for semi-supervised consistency learning, which can mine the internal semantic information of the video to a greater extent compared to image enhancement of the image itself, providing a new idea for semi-supervised learning in the video field;

[0034] 2. This method uses an end-to-end learning idea, uses the visual modality with the richest semantic expression of the video for learning classification, and the model is based on the residual network optimized on the existing large-scale image dataset and fine-tuned on the video classification task to achieve feature learning of short videos; compared with the three-dimensional convolutional neural network, this method's model avoids problems such as the need for a large-scale dataset for training and the difficulty of training from scratch to a certain extent;

[0035] 3. Different from the traditional recurrent neural network for learning in the time dimension, this method is based on an attention mechanism fusion strategy to learn the degree of association similarity between each frame and all frames in the video, stimulating the model to understand the cross-time global action event information. The additional positional encoding can help the model learn the temporal information between key frames in the video, while the category encoding can help fuse the semantic information of the entire sequence. Finally, the learned category encoding and feature mean are jointly used for classification, making the discrimination result more robust. Brief Description of the Drawings

[0036] Figure 1 This is the overall flowchart of the semi-supervised short video classification method based on semantic consistency provided by the present invention. Detailed Implementation Manner

[0037] To make the objectives, technical solutions, and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.

[0038] The embodiments of the present invention provide a semi-supervised short video classification method based on semantic consistency. Refer to Figure 1 This method includes the following steps:

[0039] Step 101: Divide the short video dataset into a training set, a validation set, and a test set, and divide the samples in the training set into labeled data and unlabeled data according to a ratio.

[0040] In the embodiments of the present invention, the Flickr dataset is collected through an open-source interface. After screening, it includes 20 action and event categories such as football, playing the piano, running, and wedding, with a total of 18,924 video samples. The sample duration in the dataset is 3 to 30 seconds, and the size is 224×224. The data is divided according to a ratio of 4:1 for the training set and the test set. In the training set, 10% of the samples in each category are taken as labeled data, and the remaining samples are taken as unlabeled data.

[0041] Before inputting into the network, all video samples start from frame t0 = 0 and extract 16 frames of images at equal intervals; at the same time, starting from frame t0 + τ, extract 16 frames of images again at equal intervals, where τ = 16. Two neighboring sequences and of the same sample are obtained, where N = 16. For the neighboring sequence X w perform standard data augmentation (including: random cropping and random flipping) operations, and for the neighboring sequence X s use the RandAugment algorithm to perform 4 times of strong data augmentation operations.

[0042] Among them, the RandAugment algorithm is an image data augmentation algorithm open-sourced by Google and is well-known to those skilled in the art.

[0043] Step 102: Build a neural network, including: spatial feature learning and encoding, temporal attention fusion module, and classifier, which consists of three parts;

[0044] Among them, the spatial feature learning and encoding part is the part before the last fully-connected layer of the Resnet-34 residual network, and the output is 512-dimensional spatial features. This part of the network is initialized with the pre-trained weights of ImageNet and fine-tuned on this basis. The input x of the temporal attention fusion module has a size of [16, 512], that is, 16 frames of 512-dimensional features. The class encoding clsToken with a size of [1, 512] is concatenated on the first dimension of x to obtain a size of [17, 512]. At the same time, it is added element-wise to the position encoding positions with a size of [17, 512], and finally the size of x is [17, 512]. At this time, the input x is passed through three fully-connected layers respectively to obtain the Queries, Keys, and Values matrices. The Queries matrix and the Keys matrix are multiplied to obtain the similarity of each frame relative to other frames in the entire video. Then, the Softmax function is used for normalization to obtain the corresponding attention matrix. The attention matrix passes through a dropout layer and is multiplied by the Values matrix. Finally, after passing through a linear layer, the feature output is obtained. Then, the frame features that fuse the global temporal attention information in the obtained feature matrix are averaged, and concatenated with the class feature encoding learned in the first dimension of the matrix. Two fully-connected layers are used as the classifier for classification. Finally, after normalization using the Softmax function, the final prediction output is obtained.

[0045] The calculation formula is as follows:

[0046]

[0047] Among them, Q, K, and V represent the Queries, Keys, and Values matrices respectively, and d k is the matrix dimension.

[0048] Step 103: Model training;

[0049] Take the same number of labeled samples and unlabeled samples after the standard data augmentation in Step 101 and concatenate them as a batch and input them into the neural network F(·) in Step 102. Calculate the error between the predicted output of the labeled part of the network samples and the actual label y i using the cross-entropy loss to construct the classification error At the same time, take the same number of labeled samples and unlabeled samples after the strong data augmentation in Step 101 and concatenate them as a batch and input them into the network in Step 102 again. Calculate the predicted output of the network through forward propagation. Use the JS divergence loss constraint to construct the consistency error between the two predicted distributions The final loss function Loss = L is obtained by jointly using the classification error and the consistency error cls + λ·L cons , and the network is optimized using backpropagation. JSD is the JS divergence, is the sample for standard data augmentation, is the sample for strong data augmentation, F(·) is the neural network in step 102, and CE is the cross entropy.

[0050] In the embodiment of the present invention, batch_size is set to 8, the model is optimized using stochastic gradient descent (SGD), the learning rate is set to 0.01, and the learning rate decay uses a multi-stage descent strategy, decaying to 0.1 of the original at the 10th, 30th, 50th, and 75th rounds of training, and a total of 80 rounds of training are performed. After the training is completed, the optimized model and parameters are saved.

[0051] Step 104: Input the video sample X into the optimized neural network to output the corresponding prediction score, and obtain the final video classification result.

[0052] In the embodiment of the present invention, except for those with special specifications for the models of each device, the models of other devices are not limited, and any device that can perform the above functions can be used.

[0053] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0054] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A semi-supervised short video classification method based on semantic consistency, characterized in that, The method includes: Utilizing the correlation of adjacent video frames, key frames are extracted from labeled data and unlabeled data at equal time intervals T starting from frames t0 and t0+τ respectively to obtain two frame sequences, and strong data augmentation and standard data augmentation are performed respectively; Construct a neural network composed of a spatial feature learning module, a temporal attention fusion module, and a classifier module; The same number of labeled samples and unlabeled samples after standard data augmentation are concatenated and input into the neural network, and the classification loss of the labeled part is calculated; the same number of labeled samples and unlabeled samples after strong data augmentation are concatenated and input into the neural network to obtain a prediction output, and the consistency loss is calculated for the prediction distributions obtained by inputting the same sample into the network after these two different data augmentation processes; The classification loss and the consistency loss are jointly used for the optimization training of the neural network; the video samples are input into the optimized neural network to output corresponding prediction scores, and the final video classification result is obtained; Among them, the construction of the neural network composed of a spatial feature learning module, a temporal attention fusion module, and a classifier module is as follows: The spatial feature learning module uses a residual network for spatial feature encoding and inputs the video frame spatial feature sequence into the temporal attention fusion module; The temporal attention fusion module concatenates position information encoding in the learned video frame spatial feature sequence for the learning of the temporal relationship of each key frame, and at the same time adds category information encoding for the fusion of the temporal semantic information of the entire sequence; Learn the degree of correlation between key frames within the sequence, calculate the association matrix of each video frame in the sequence with all video frames in the sequence, and obtain a frame feature sequence with global attention information; Use the frame feature sequence to calculate the mean to fuse the video frame information, obtain the video-level semantic representation, and jointly input the semantic feature and the learned category encoding information into the classifier module composed of fully connected layers to output the final classification result.

2. The semi-supervised short video classification method based on semantic consistency according to claim 1, characterized in that, The performing of strong data augmentation and standard data augmentation respectively is as follows: Extract N frames from all video samples at equal time intervals T starting from t0 and t0 + τ, obtaining two key frame sequence sets for each sample and Among them, F i is the i-th frame image of the video sample. For the images in X w , the same standard data augmentation operations are performed. For X s , the RandAugment algorithm is used to perform multiple strong data augmentation operations.

3. A semi-supervised short video classification method based on semantic consistency according to claim 2, characterized in that The adding of category information encoding for the fusion of the temporal semantic information of the entire sequence is as follows: Input sample X m Input the residual network f(·) to learn the spatial appearance features of each frame and obtain the video feature matrix n is the number of frames, and d is the output feature dimension of the network f(·); Concatenate a randomly initialized class information encoding bit at the first dimension position of the video feature matrix Obtain the feature matrix representing the video Add the randomly initialized position encoding information Obtain the final feature matrix.

4. A semi-supervised short video classification method based on semantic consistency according to claim 1, characterized in that, The classification loss is: The labeled samples after standard data augmentation and the unlabeled samples Take the same number of them, splice them and input them into the neural network. The neural network outputs the predicted distribution and obtains the result. Use cross-entropy to calculate the error loss L between the predicted result of the labeled part and the actual label cls .

5. A semi-supervised short video classification method based on semantic consistency according to claim 4, characterized in that, The consistency loss is: The consistency loss L between the prediction distributions of the network for the same sample after two different data augmentation processes is calculated using the JS divergence. cons ; Among them, KL represents the KL divergence, and the calculation method is as follows: Finally, the two losses are jointly used for the training optimization of the network, and the total loss is as follows: L = L cls + λ·L cons Among them, λ is an adjustable parameter.

Citation Information

Patent Citations

  • Semi-supervised video classification method and system based on neighbor consistency and comparative learning

    CN115311605A

  • Weakly supervised video activity detection method and system based on iterative learning

    US20220189209A1