Training Method and Device for Video Feature Extraction Network
Through self-supervised learning and comparative learning, the video feature extraction network is trained, and the problem that training video-related algorithm models rely on a large amount of labeled data is solved, and efficient video feature extraction and classification is achieved.
Patent Information
- Application Number
- CN202210530591.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-05-16
AI Technical Summary
In the prior art, training video-related algorithm models requires a large amount of marked training data, which leads to high cost of manual labeling and making it difficult to efficiently extract video feature.
The self-supervised learning strategy is adopted to obtain video frame sequences of different image parameters of the sample video, and to train the video feature extraction network using contrast learning and cross-entropy loss functions, combined with the out-of-order weight prediction network, reduce the dependence on the annotated data.
With a small amount of labeled data, efficient training of video feature extraction network is realized, the accuracy of video classification is improved, and it can be used for video retrieval.
Smart Images

Figure CN114821443B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the technical field of video processing, and particularly to a training method for a video feature extraction network, and a video classification method based on the video feature extraction network. One or more embodiments of this specification also relate to a training device for a video feature extraction network, a video classification device based on the video feature extraction network, a computing device, and a computer-readable storage medium. Background Art
[0002] With the application and development of multimedia technology, more and more services rely on technologies such as video classification and video understanding. Generally speaking, when training algorithm models related to videos, a large amount of labeled training data needs to be prepared. Since video is an important content carrier and contains a large number of features, large-scale manual annotation of videos in the business field is time-consuming and laborious. Therefore, how to reduce the training cost of training algorithm models related to videos is an urgent problem to be solved currently. Summary of the Invention
[0003] In view of this, the embodiments of this specification provide a training method for a video feature extraction network, and a video classification method based on the video feature extraction network. One or more embodiments of this specification also relate to a training device for a video feature extraction network, a video classification device based on the video feature extraction network, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0004] According to the first aspect of the embodiments of this specification, a training method for a video feature extraction network is provided, including:
[0005] Obtain a sample video, and obtain a first video frame sequence, a second video frame sequence, and a sequence shuffled sample weight of the sample video according to the sample video, wherein the image parameters of the first video frame sequence and the second video frame sequence are different;
[0006] Input the first video frame sequence and the second video frame sequence into the to-be-trained feature extraction network respectively, obtain a first extraction result output by the to-be-trained feature extraction network, and input the first video frame sequence and the second video frame sequence into the reference feature extraction model respectively, obtain a second extraction result output by the reference feature extraction model;
[0007] Obtain a comparison coding result, and calculate a first loss value based on the comparison coding result, the first extraction result, and the second extraction result;
[0008] Input the first extraction result into the out-of-order weight prediction network to obtain the sequence out-of-order prediction weight corresponding to the sample video output by the out-of-order weight prediction network, and calculate a second loss value based on the sequence out-of-order prediction weight and the sequence out-of-order sample weight;
[0009] Adjust the network parameters of the to-be-trained feature extraction network according to the first loss value and the second loss value until the training stop condition is reached, and obtain a trained video feature extraction network.
[0010] According to the second aspect of the embodiments of the present specification, a video classification method based on a video feature extraction network is provided, including:
[0011] Obtain a video to be classified;
[0012] Input the video to be classified into a video feature extraction network trained by the training method of the video feature extraction network to obtain a target feature extraction result corresponding to the video to be classified output by the video feature extraction network;
[0013] Input the target feature extraction result into a classifier to obtain a classification result of the video to be classified output by the classifier.
[0014] According to the third aspect of the embodiments of the present specification, a training device for a video feature extraction network is provided, including:
[0015] An acquisition module, configured to acquire a sample video, and obtain a first video frame sequence, a second video frame sequence and a sequence out-of-order sample weight of the sample video according to the sample video, where image parameters of the first video frame sequence and the second video frame sequence are different;
[0016] An input module, configured to input the first video frame sequence and the second video frame sequence into the to-be-trained feature extraction network respectively to obtain a first extraction result output by the to-be-trained feature extraction network, and input the first video frame sequence and the second video frame sequence into a reference feature extraction model respectively to obtain a second extraction result output by the reference feature extraction model;
[0017] A first calculation module, configured to obtain a comparison encoding result, and calculate a first loss value based on the comparison encoding result, the first extraction result and the second extraction result;
[0018] A second calculation module, configured to input the first extraction result into the out-of-order weight prediction network to obtain the sequence out-of-order prediction weight corresponding to the sample video output by the out-of-order weight prediction network, and calculate a second loss value based on the sequence out-of-order prediction weight and the sequence out-of-order sample weight;
[0019] An adjustment module, configured to adjust network parameters of the to-be-trained feature extraction network according to the first loss value and the second loss value until a training stop condition is reached, and obtain a trained video feature extraction network.
[0020] According to a fourth aspect of the embodiments of the present specification, there is provided a video classification device based on a video feature extraction network, including:
[0021] An acquisition module, configured to acquire a to-be-classified video;
[0022] An input module, configured to input the to-be-classified video into a video feature extraction network obtained by training through a training method of a video feature extraction network, and obtain a target feature extraction result corresponding to the to-be-classified video output by the video feature extraction network;
[0023] An obtaining module, configured to input the target feature extraction result into a classifier, and obtain a classification result of the to-be-classified video output by the classifier.
[0024] According to a fifth aspect of the embodiments of the present specification, there is provided a computing device, including a memory, a processor, and computer instructions stored on the memory and executable on the processor. When the processor executes the computer instructions, the steps of the training method of the video feature extraction network and the video classification method based on the video feature extraction network are implemented.
[0025] According to a sixth aspect of the embodiments of the present specification, there is provided a computer-readable storage medium storing computer instructions, and when the computer instructions are executed by a processor, the steps of the training method of the video feature extraction network and the video classification method based on the video feature extraction network are implemented.
[0026] According to a seventh aspect of the embodiments of the present specification, there is provided a computer program, and when the computer program is executed on a computer, the computer is made to execute the steps of the training method of the video feature extraction network and the video classification method based on the video feature extraction network.
[0027] The training method of the video feature extraction network provided in this specification includes: obtaining a sample video, and obtaining a first video frame sequence, a second video frame sequence and a sequence scrambling sample weight of the sample video according to the sample video, wherein the image parameters of the first video frame sequence and the second video frame sequence are different; respectively inputting the first video frame sequence and the second video frame sequence into the feature extraction network to be trained to obtain a first extraction result output by the feature extraction network to be trained, and respectively inputting the first video frame sequence and the second video frame sequence into a reference feature extraction model to obtain a second extraction result output by the reference feature extraction model; obtaining a comparison coding result, and calculating a first loss value based on the comparison coding result, the first extraction result and the second extraction result; inputting the first extraction result into a scrambling weight prediction network to obtain a sequence scrambling prediction weight corresponding to the sample video output by the scrambling weight prediction network, and calculating a second loss value based on the sequence scrambling prediction weight and the sequence scrambling sample weight; adjusting network parameters of the feature extraction network to be trained according to the first loss value and the second loss value until a training stop condition is reached, and obtaining a trained video feature extraction network.
[0028] In one embodiment of this specification, by obtaining a first video frame sequence, a second video frame sequence and a sequence scrambling sample weight of a sample video according to the sample video, extracting the first video frame sequence and the second video frame sequence according to the feature extraction network to be trained to obtain a first extraction result, and extracting the first video frame sequence and the second video frame sequence by a reference feature extraction model to obtain a second extraction result, calculating a first loss value based on the comparison coding result, the first extraction result and the second extraction result, calculating a second loss value according to the sequence scrambling prediction weight output by the scrambling weight prediction network and the sequence scrambling sample weight, and adjusting parameters of the video feature extraction network based on the first loss value and the second loss value, a trained video feature extraction network is obtained. It realizes the training of the video feature extraction network without a large amount of labeled data, and can achieve the purpose of high accuracy with a small amount of labeled data. Description of the Drawings
[0029] Figure 1 is a flowchart of a training method of a video feature extraction network provided by an embodiment of this specification;
[0030] Figure 2 is a flowchart of a processing procedure of a video classification method based on a video feature extraction network provided by an embodiment of this specification;
[0031] Figure 3 is a schematic structural diagram of a training device of a video feature extraction network provided by an embodiment of this specification;
[0032] Figure 4 It is a schematic structural diagram of a video classification device based on a video feature extraction network provided by an embodiment of this specification;
[0033] Figure 5 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0034] Numerous specific details are set forth in the following description to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0035] The terms used in one or more embodiments of this specification are merely for the purpose of describing specific embodiments and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to any and all possible combinations of one or more of the associated listed items.
[0036] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first can also be referred to as the second, and similarly, the second can also be referred to as the first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining".
[0037] First, the noun terms related to one or more embodiments of this specification are explained.
[0038] Self-supervised learning: A training method that can be trained on unlabeled data and enables the model to have the ability to extract features.
[0039] Feature learning: A method for training the feature extraction ability of a model, the purpose of which is to obtain a feature extraction model with strong generalization ability to assist the training of downstream tasks, and generally can be supervised or self-supervised.
[0040] Contrastive learning: A form of feature learning, which is trained by constructing sets of positive sample pairs and negative sample pairs and using similarity comparison.
[0041] At present, more and more businesses rely on technologies such as video classification and video understanding. Generally speaking, training algorithm models related to videos requires a large amount of training data with labeled tags. However, manually labeling videos in large-scale business fields is time-consuming and laborious, resulting in too high costs for training data. Therefore, reducing the labeled training data for businesses has very high value.
[0042] Based on this, in this specification, a training method for a video feature extraction network is provided, which learns visual features using a self-supervised training strategy to obtain a video feature extraction network; a video classification method based on the video feature extraction network. This specification also relates to a training device for a video feature extraction network, a video classification device based on the video feature extraction network, a computing device, a computer-readable storage medium, and a computer program, which will be described in detail one by one in the following embodiments.
[0043] Figure 1 The flowchart shows a training method for a video feature extraction network provided according to an embodiment of this specification, including steps 102 to 110.
[0044] Step 102: Obtain a sample video, and obtain a first video frame sequence, a second video frame sequence, and a sequence scrambling sample weight of the sample video according to the sample video, where the image parameters of the first video frame sequence and the second video frame sequence are different.
[0045] Among them, the sample video can be understood as a video for training. The sample video is not labeled with corresponding classification tags. In practical applications, the sample video can be a landscape video, a person video, etc.; the first video frame sequence and the second video frame sequence can be understood as video frame sequences sampled from the same sample video, and the first video frame sequence and the second video frame sequence include video frames of the sample video; the sequence scrambling sample weight can be understood as the weight calculated according to the sample video. Substantially, the sequence scrambling sample weight refers to the frame order chaos degree classification weight of the video frame sequence. The sequence scrambling sample weight is used as the chaos degree classification label corresponding to the sample video for subsequent calculation of the cross-entropy loss function.
[0046] In practical applications, the image parameters of the first video frame sequence and the second video frame sequence are different because the first video frame sequence and the second video frame sequence are obtained through different methods of image data augmentation of the video frame sequence of the sample video, resulting in different image parameters of the first video frame sequence and the second video frame sequence. Image data augmentation is to artificially expand on a limited data set to generate more equivalent data. It can effectively make up for the deficiencies of existing training data, prevent the model from overfitting, and enhance the generalization ability of the model.
[0047] In an embodiment of this specification, a sample video A is obtained, and a first video frame sequence a1, a second video frame sequence a2 of the sample video, and a sequence disordered sample weight are obtained according to the sample video A.
[0048] Specifically, obtaining the first video frame sequence, the second video frame sequence, and the sequence disordered sample weight of the sample video according to the sample video includes:
[0049] Performing video sampling on the sample video according to a preset sampling frequency to generate an initial video frame sequence;
[0050] Shuffling the order of video frames in the initial video frame sequence to obtain a shuffled video frame sequence corresponding to the initial video frame sequence;
[0051] Generating the first video frame sequence, the second video frame sequence, and the sequence disordered sample weight of the sample video based on the initial video frame sequence and the shuffled video frame sequence.
[0052] Among them, the preset sampling frequency can be understood as a preset sampling interval. For example, the preset sampling frequency is 2, that is, the sampling interval is set to 2. When the sample video is a 64-frame video, after interval sampling according to the preset sampling frequency, a video frame sequence with a length of 32 can be obtained. The shuffled video frame sequence can be understood as a video frame sequence obtained by shuffling the initial video frame sequence. For example, the video frames in the initial video frame sequence are sorted according to [1-2-3...32], and after randomly shuffling the order of video frames in the initial video frame sequence, the order of video frames in the obtained shuffled video frame sequence can be [3-6-30...4], etc.
[0053] In practical applications, after obtaining the initial video frame sequence and the shuffled video frame sequence, the first video frame sequence and the second video frame sequence can be obtained based on the initial video frame sequence and the shuffled video frame sequence, and the degree of difference between the shuffled video frame sequence and the initial video frame sequence can be calculated through the edit distance formula as the sequence disordered sample weight, that is, the chaos degree classification label. It should be noted that the shuffling method needs to ensure that the edit distance conforms to a uniform distribution, which is beneficial to the training of the video frame sequence disorder degree classification.
[0054] In an embodiment of this specification, following the above example, the preset sampling frequency is 2, interval sampling is performed on the sample video A to obtain an initial video frame sequence, and the length of the initial video frame sequence is 16. The order of the initial video frame sequence is randomly shuffled to obtain a shuffled video frame sequence, and then the first video frame sequence a1 and the second video frame sequence a2 are generated based on the initial video frame sequence and the shuffled video frame sequence, and the sequence disordered sample weight is calculated.
[0055] Specifically, generating the first video frame sequence, the second video frame sequence, and the sequence scrambling sample weight of the sample video based on the initial video frame sequence and the scrambled video frame sequence includes:
[0056] Synthesizing a video frame sequence to be enhanced according to the initial video frame sequence and the scrambled video frame sequence;
[0057] Adjusting the image parameters of the video frame sequence to be enhanced according to a first preset enhancement rule to obtain a first video frame sequence;
[0058] Adjusting the image parameters of the video frame sequence to be enhanced according to a second preset enhancement rule to obtain a second video frame sequence;
[0059] Calculating the sequence scrambling sample weight corresponding to the sample video according to the initial video frame sequence and the scrambled video frame sequence.
[0060] Among them, the video frame sequence to be enhanced can be understood as the video frame sequence obtained by merging the initial video frame sequence and the scrambled video frame sequence in time series. The video sequence to be enhanced is used for subsequent image data enhancement to obtain the first video frame sequence and the second video frame sequence. Adjusting the image parameters of the video frame sequence to be enhanced can be understood as performing data enhancement on the video frame sequence to be enhanced. Data enhancement includes methods such as blurring, equalization, color interference, and noise. Randomly select two different enhancement methods to perform image enhancement on the video frame sequence to be enhanced, so as to obtain the first video frame sequence and the second video frame sequence. Since the first video frame sequence and the second video frame sequence are obtained through different data enhancement methods, the image parameters of the first video frame sequence and the second video frame sequence are different.
[0061] In practical applications, after merging the initial video frame sequence and the scrambled video frame sequence in time series to obtain the video frame sequence to be enhanced, two different sets of image data enhancement parameters can be randomly obtained using a random seed, and data enhancement is performed on the video frame sequence to be enhanced according to the two different sets of image data enhancement parameters to obtain the first video frame sequence and the second video frame sequence.
[0062] In an embodiment of this specification, following the above example, the initial video frame sequence and the scrambled video frame sequence are merged in time series to obtain the video frame sequence to be enhanced, and two different sets of image data enhancement parameters are obtained using a random seed, including adjusting the short side of the video frame sequence to be enhanced to 128 pixel values and then adjusting the scale to shear to 112 pixel values to obtain the first video frame sequence a1; flipping the video sequence to be enhanced horizontally to obtain the second video frame sequence a2. Use the edit distance formula to calculate the degree of difference between the initial video frame sequence and the scrambled video frame sequence, and use the calculation result as the sequence scrambling sample weight, that is, the chaos degree classification label.
[0063] Step 104: Input the first video frame sequence and the second video frame sequence into the feature extraction network to be trained respectively, obtain the first extraction result output by the feature extraction network to be trained, input the first video frame sequence and the second video frame sequence into the reference feature extraction model respectively, and obtain the second extraction result output by the reference feature extraction model.
[0064] Among them, the feature extraction network to be trained can be a 3D convolutional network, and the reference feature extraction model can be a video translation model (video transformer). Inputting the first video frame sequence and the second video frame sequence into the feature extraction network to be trained respectively can obtain the first extraction result output by the feature extraction network to be trained. The first extraction result includes the result extracted from the first video frame sequence and the result extracted from the second video frame sequence. Inputting the first video frame sequence and the second video frame sequence into the reference feature extraction model respectively can obtain the second extraction result output by the reference feature extraction model. The second extraction result includes the result extracted from the first video frame sequence and the result extracted from the second video frame sequence.
[0065] In practical applications, since both the first extraction result and the second extraction result are obtained by extracting from the video frame sequences of the same sample video, although the network structures of the feature extraction network to be trained and the reference feature extraction model are different, that is, through different extraction forms, the semantic information represented by the extraction results is the same, both representing the visual features of the sample video. Therefore, the similarity between the first extraction result and the second extraction result is relatively high.
[0066] In an embodiment of this specification, following the above example, input the first video frame sequence a1 and the second video frame sequence a2 into the feature extraction network to be trained respectively to obtain the first extraction result output by the feature extraction network to be trained. Input the first video frame sequence a1 and the second video frame sequence a2 into the reference feature extraction model respectively to obtain the second extraction result output by the reference feature extraction model.
[0067] Step 106: Obtain the alignment coding result, and calculate the first loss value based on the alignment coding result, the first extraction result, and the second extraction result.
[0068] Among them, the alignment coding result can be understood as the coding result obtained by coding the first video frame sequence and the second video frame sequence. The first loss value can be understood as the contrastive learning loss value, that is, the loss value calculated by the contrastive learning loss function.
[0069] In practical applications, during the training process based on a sample video, the obtained comparison encoding result can be the comparison encoding result obtained previously by encoding the corresponding first video frame sequence and second video frame sequence according to other sample videos. And when obtaining the comparison encoding result, all the previously generated encoding results will be obtained as the training data used in this training, that is, the comparison encoding result.
[0070] In an embodiment of this specification, the comparison encoding result is obtained, and the contrast learning loss value is calculated based on the comparison encoding result, the first extraction result, and the second extraction result.
[0071] Specifically, obtaining the comparison encoding result includes:
[0072] Read the comparison encoding result queue;
[0073] Obtain the comparison encoding result from the comparison encoding result queue.
[0074] Among them, the comparison encoding result queue can be understood as a queue that stores the encoding results obtained by encoding in each training iteration. In each training iteration, the generated encoding results will be stored in a queue structure. In subsequent training, the comparison encoding result is obtained from the comparison encoding structure queue.
[0075] In practical applications, during this training process, the first video frame sequence and the second video frame sequence are encoded to obtain the encoding result of this training, and the encoding result of this training is stored in the comparison encoding result queue. This comparison encoding result queue follows the first-in, first-out principle, and the new encoding result generated in each iteration replaces the old encoding result.
[0076] In an embodiment of this specification, the comparison encoding result queue is read, and all the comparison encoding results in the comparison encoding result queue are taken out.
[0077] Specifically, the method further includes:
[0078] Input the first video frame sequence and the second video frame sequence into a dynamic encoder to obtain the comparison encoding result corresponding to the sample video output by the dynamic encoder;
[0079] Add the comparison encoding result corresponding to the sample video to the comparison encoding result queue.
[0080] Among them, the dynamic encoder (momentum encoder) can be understood as an encoder that performs momentum encoding. Input the first video frame sequence and the second video frame sequence into the dynamic encoder, and the dynamic encoder will output feature vectors respectively. The obtained feature vectors are used as the comparison encoding result and added to the comparison encoding result queue.
[0081] In practical applications, adding the comparison coding results generated in this training to the comparison coding result queue will replace the old comparison coding results at the tail of the comparison coding result queue, ensuring that the comparison coding results in the comparison coding result queue are consistent with the comparison coding results of this training.
[0082] In an embodiment of this specification, following the previous example, the first video frame sequence a1 and the second video frame sequence a2 are input into the dynamic encoder to obtain the comparison coding result b output by the dynamic encoder, and the comparison coding result b is added to the comparison coding result queue.
[0083] Specifically, calculating the first loss value based on the comparison coding result, the first extraction result, and the second extraction result includes:
[0084] Constructing a set of positive sample pairs and negative sample pairs according to the comparison coding result, the first extraction result, and the second extraction result;
[0085] Calculating the first loss value according to the set of positive sample pairs and the set of negative sample pairs.
[0086] Among them, the positive sample pair can be understood as being constructed based on the first extraction result and the second extraction result. The positive sample pair includes: the feature vector extracted from the first video frame sequence in the first extraction result and the feature vector extracted from the second video frame sequence in the second extraction result; the negative sample pair can be understood as being constructed based on the extraction result and the comparison coding result. The negative sample pair includes: the feature vector extracted from the video frame sequence in the extraction result and the feature vector in the comparison coding result. It should be noted that the feature vectors in the extraction results of the same set of positive sample pairs and negative sample pairs are the same, that is, they include the feature vectors extracted from the same video frame sequence. In practical applications, since there can be multiple comparison coding results, one positive sample pair can correspond to multiple negative sample pairs.
[0087] In an embodiment of this specification, following the previous example, a set of positive sample pairs and negative sample pairs is constructed according to the comparison coding result obtained from the comparison coding result queue, the first extraction result, and the second extraction result, and the first loss value is calculated according to the set of positive sample pairs and the set of negative sample pairs.
[0088] Specifically, constructing a set of positive sample pairs and negative sample pairs according to the comparison coding result, the first extraction result, and the second extraction result includes:
[0089] Constructing positive sample pairs according to the first extraction result and the second extraction result;
[0090] Constructing the first set of negative sample pairs corresponding to the first extraction result according to the first extraction result and the comparison coding result;
[0091] Construct a second negative sample pair set corresponding to the second extraction result according to the second extraction result and the comparison coding result;
[0092] Calculate a first loss value according to the positive sample pair, the first negative sample pair set, and the second negative sample pair set.
[0093] In practical applications, a positive sample pair is constructed according to the first extraction result and the second extraction result, a first negative sample pair set is constructed according to the first extraction result and multiple comparison coding results, a second negative sample pair set is constructed according to the second extraction result and multiple comparison coding results, and a first loss value is calculated according to the positive sample pair and the negative sample pairs in the first negative sample pair set and the negative sample pairs in the second negative sample pair set.
[0094] In an embodiment of this specification, a positive sample pair is constructed according to the first extraction result and the second extraction result, a first negative sample pair set corresponding to the first extraction result is constructed according to the first extraction result and the comparison coding result, a second negative sample pair combination corresponding to the second extraction result is constructed according to the second extraction result and the comparison coding result, and a contrast learning function is calculated based on the positive sample pair, the negative sample pairs in the first negative sample pair set, and the negative sample pairs in the second negative sample pair set to obtain a first loss value.
[0095] Specifically, calculating the first loss value according to the positive sample pair, the first negative sample pair set, and the second negative sample pair set includes:
[0096] Substitute the positive sample pair, the first negative sample pair set, and the second negative sample pair set into a contrast learning loss function to obtain a first loss value.
[0097] In practical applications, substitute the positive sample pair, the negative sample pairs in the first negative sample pair set, and the negative sample pairs in the second negative sample pair set into the contrast learning loss function respectively to obtain a first loss value. Based on the first loss value, subsequent adjustment of the network parameters of the feature extraction network to be trained can be performed.
[0098] Specifically, the first extraction result includes a first feature vector extracted from a first video frame sequence and a second feature vector extracted from a second video frame sequence, and the second extraction result includes a third feature vector extracted from the first video frame sequence and a fourth feature vector extracted from the second video frame sequence;
[0099] Constructing a positive sample pair according to the first extraction result and the second extraction result includes:
[0100] Construct a first positive sample pair according to the first feature vector and the third feature vector;
[0101] Construct a second positive sample pair according to the second feature vector and the fourth feature vector.
[0102] In practical applications, after inputting the first video frame sequence and the second video frame sequence into the feature extraction network to be trained respectively, the feature extraction network to be trained will output two feature vectors. After inputting the first video frame sequence and the second video frame sequence into the reference feature extraction model respectively, the reference feature extraction model will also output two feature vectors. When constructing positive sample pairs, positive sample pairs will be constructed according to the feature vectors extracted from the same video frame sequence. For example, input the first video frame sequence into the feature extraction network to be trained and the reference feature extraction model respectively, to obtain the feature vector 1 output by the feature extraction network to be trained and the feature vector 2 output by the reference feature extraction model. Input the second video frame sequence into the feature extraction network to be trained and the reference feature extraction model respectively, to obtain the feature vector 3 output by the feature extraction network to be trained and the feature vector 4 output by the reference feature extraction model. Then, a positive sample pair can be constructed according to the feature vector 1 and the feature vector 2, and a positive sample pair can be constructed according to the feature vector 3 and the feature vector 4.
[0103] In an embodiment of this specification, the first extraction result includes a first feature vector extracted from the first video frame sequence a1 and a second feature vector obtained from the second video frame sequence a2. The second extraction result includes a third feature vector extracted from the first video frame sequence a1 and a fourth feature vector obtained from the second video frame sequence a2. Construct a first positive sample pair according to the first feature vector and the third feature vector, and construct a second positive sample pair according to the second feature vector and the fourth feature vector.
[0104] Specifically, constructing the first negative sample pair set corresponding to the first extraction result according to the first extraction result and the comparison coding result, and constructing the second negative sample pair set corresponding to the second extraction result according to the second extraction result and the comparison coding result, includes:
[0105] Construct a first negative sample pair set according to the first feature vector, the third feature vector and the comparison coding result;
[0106] Construct a second negative sample pair set according to the second feature vector, the fourth feature vector and the comparison coding result.
[0107] Among them, the negative sample pairs in the first negative sample pair set are sample pairs constructed according to the first feature vector and the comparison coding result or the second feature vector and the comparison coding result. Since there are multiple comparison coding results, the first and second feature vectors can construct negative sample pairs with multiple comparison coding results. Similarly, the negative sample pairs in the second negative sample pair set are sample pairs constructed according to the third feature vector and the comparison coding result or the fourth feature vector and the comparison coding result, and the third and fourth feature vectors can construct negative sample pairs with multiple comparison coding results.
[0108] In practical applications, when calculating the first loss value based on the positive sample pairs, the first negative sample pair set, and the second negative sample pair set, positive samples and negative sample pairs constructed by the same feature vector are used to calculate the first loss value. For example, feature vector 1 and feature vector 2 construct positive sample pair 1, feature vector 1 and the comparison coding result construct negative sample pair 1, feature vector 2 and the comparison coding result construct negative sample pair 2, and the positive sample pair 1 is used to calculate the contrast learning loss value with the negative sample pair 1 or the negative sample pair 2.
[0109] In an embodiment of this specification, following the above example, multiple first negative sample pairs are constructed according to the first feature vector in the first extraction result and multiple comparison coding results, and multiple first negative sample pairs are constructed according to the third feature vector in the second extraction result and multiple comparison coding results. The first negative sample pair set is constructed based on all the first negative sample pairs; multiple second negative sample pairs are constructed according to the second feature vector in the first extraction result and multiple comparison results, and multiple second negative sample pairs are constructed according to the fourth feature vector in the second extraction result and multiple comparison coding results. The second negative sample pair set is constructed based on all the second negative sample pairs.
[0110] Step 108: Input the first extraction result into the disordered weight prediction network to obtain the sequence disorder prediction weight of the sample video output by the disordered weight prediction network, and calculate the second loss value based on the sequence disorder prediction weight and the sequence disorder sample weight.
[0111] Among them, the disordered weight prediction network can be a classification head network. The disordered weight prediction network is used to predict the sequence disorder prediction weight of the video frame sequence, that is, to predict the degree of disruption of the video frame sequence. And according to the sequence disorder sample weight, the cross-entropy loss function can be calculated to obtain the second loss value.
[0112] In practical applications, the feature vector output by the feature extraction network to be trained is input into the classification head network, and after normalization processing (softmax) in the classification head network, the sequence disorder prediction weight is output.
[0113] In an embodiment of this specification, following the previous example, the feature vector 1 and the feature vector 2 in the first extraction result are input into the disordered weight prediction network to obtain the sequence disorder prediction weight output by the disordered weight prediction network. Based on the sequence disorder prediction weight and the sequence disorder sample weight, a second loss value is calculated, and based on the second loss value, the network parameters of the feature extraction network to be trained are adjusted.
[0114] Specifically, calculating the second loss value based on the sequence disorder prediction weight and the sequence disorder sample weight includes:
[0115] Substitute the sequence disorder prediction weight and the sequence disorder sample weight into the cross-entropy loss function to obtain the second loss value.
[0116] In practical applications, the sequence disorder prediction weight and the sequence disorder sample weight will be substituted into the cross-entropy loss function to obtain the second loss function.
[0117] In an embodiment of this specification, the sequence disorder prediction weight and the sequence disorder sample weight output by the disordered weight prediction network are substituted into the cross-entropy loss function to calculate and obtain the second loss value.
[0118] Step 110: Adjust the network parameters of the feature extraction network to be trained according to the first loss value and the second loss value until the training stop condition is reached, and obtain the trained video feature extraction network.
[0119] Among them, the training stop condition can be that the first loss value and the second loss value reach the preset loss threshold, the training iteration times reach the preset number of rounds, etc. The specific training stop condition can be determined according to the actual situation. The trained video feature extraction network can be understood as that the video feature extraction network has been trained to the model network convergence state.
[0120] In practical applications, after obtaining the first loss value and the second loss value each time training, the network parameters of the feature extraction network to be trained are adjusted according to the first loss value and the second loss value. When the number of training times reaches the preset number of times, the training of the video feature extraction network is stopped, or when both the first loss value and the second loss value reach the corresponding preset loss threshold, the training of the video feature extraction network is stopped.
[0121] In an embodiment of this specification, the network parameters of the feature extraction network to be trained are adjusted according to the first loss value and the second loss value, and the training is stopped after the training iteration rounds reach the preset number of rounds, and the trained video feature extraction network is obtained.
[0122] In another embodiment of this specification, the network parameters of the feature extraction network to be trained are adjusted according to the first loss value and the second loss value. When the first loss value reaches the first preset loss threshold and the second loss value reaches the second preset loss threshold, the training is stopped and the trained video feature extraction network is obtained.
[0123] Specifically, adjusting the network parameters of the feature extraction network to be trained according to the first loss value and the second loss value includes:
[0124] Performing gradient backpropagation according to the first loss value to adjust the network parameters of the feature extraction network to be trained;
[0125] Adjusting the network parameters of the feature extraction network to be trained according to the second loss value.
[0126] In practical applications, the first loss value is used for gradient backpropagation to adjust the network parameters of the feature extraction network to be trained.
[0127] In one embodiment of this specification, the first loss value is used for gradient backpropagation to adjust the network parameters of the feature extraction network to be trained, improving the similarity of positive sample pairs and reducing the similarity of negative sample pairs. The network parameters of the feature extraction network to be trained are adjusted according to the second loss value, so as to achieve the purpose of optimizing the 3D convolutional network parameters.
[0128] During the training iteration process, the model parameters of the reference feature extraction model and the model parameters of the dynamic encoder are also adjusted based on the first loss value and the second loss value. Specifically, the method further includes:
[0129] Adjusting the model parameters of the reference feature extraction model and the model parameters of the dynamic encoder according to the first loss value;
[0130] Adjusting the network parameters of the scrambled weight prediction network according to the second loss value.
[0131] In one embodiment of this specification, the model parameters of the reference feature extraction model and the model parameters of the dynamic encoder are adjusted according to the first loss value, and the network parameters of the scrambled weight prediction network are adjusted according to the second loss value.
[0132] A training method for a video feature extraction network provided in this specification includes: obtaining a sample video, and obtaining a first video frame sequence, a second video frame sequence, and a sequence scrambling sample weight of the sample video according to the sample video, where image parameters of the first video frame sequence and the second video frame sequence are different; respectively inputting the first video frame sequence and the second video frame sequence into a to-be-trained feature extraction network to obtain a first extraction result output by the to-be-trained feature extraction network, and respectively inputting the first video frame sequence and the second video frame sequence into a reference feature extraction model to obtain a second extraction result output by the reference feature extraction model; obtaining a comparison encoding result, and calculating a first loss value based on the comparison encoding result, the first extraction result, and the second extraction result; inputting the first extraction result into a scrambling weight prediction network to obtain a sequence scrambling prediction weight corresponding to the sample video output by the scrambling weight prediction network, and calculating a second loss value based on the sequence scrambling prediction weight and the sequence scrambling sample weight; adjusting network parameters of the to-be-trained feature extraction network according to the first loss value and the second loss value until a training stop condition is reached, and obtaining a trained video feature extraction network. By respectively inputting the first video frame sequence and the second video frame sequence of the sample video into the to-be-trained feature extraction network and the reference feature extraction model, and combining two training tasks of contrast learning and cross-entropy learning, the training of the to-be-trained feature extraction network is completed. It is realized that a large number of unlabeled sample videos are used as training data, and visual features are learned by using a self-supervised training strategy to obtain a trained video feature extraction network. The video feature extraction network can be applied to downstream classification tasks, can effectively improve the accuracy of video classification, and can also be directly used for video feature extraction and then video retrieval.
[0133] The following combines the attached Figure 2 , taking the application of the video classification method based on the video feature extraction network provided in this specification in video classification as an example, to illustrate the video classification method based on the video feature extraction network. Among them, Figure 2 FIG. shows a processing procedure flowchart of a video classification method based on a video feature extraction network provided in an embodiment of this specification, and the specific steps include Step 202 to Step 206.
[0134] Step 202: Obtain a video to be classified.
[0135] Among them, the video to be classified is a video waiting to be classified, which can be obtained from the Internet or pre-stored in a device to obtain the video to be classified.
[0136] Step 204: Input the video to be classified into the video feature extraction network obtained by any of the above video feature extraction network training methods, and obtain the target feature extraction result corresponding to the video to be classified output by the video feature extraction network.
[0137] In the embodiment provided in this specification, taking the video to be classified as a landscape video as an example, the landscape video includes an island in the sea. Input the landscape video into the pre-trained video feature extraction network. The pre-trained video feature extraction network is trained to output the video features of the landscape video according to the input landscape video.
[0138] Step 206: Input the target feature extraction result into the classifier, and obtain the classification result of the video to be classified output by the classifier.
[0139] In the embodiment provided in this specification, input the target feature extraction result output by the video feature extraction network into the classifier, and the classifier determines the classification result of the landscape video.
[0140] A video classification method based on a video feature extraction network provided in this specification includes: obtaining a video to be classified, inputting the video to be classified into the video feature extraction network obtained by any of the above video feature extraction network training methods, obtaining the target feature extraction result corresponding to the video to be classified output by the video feature extraction network, inputting the target feature extraction result into the classifier, and obtaining the classification result of the video to be classified output by the classifier. Extract the classification feature vector of the video to be classified through the trained video feature extraction network, input the classification feature vector into the classifier, and determine the category of the video to be classified according to the classification result output by the classifier.
[0141] Corresponding to the above method embodiment, this specification also provides an embodiment of a training device for a video feature extraction network. Figure 3 The structural schematic diagram of a training device for a video feature extraction network provided by an embodiment of this specification is shown. As Figure 3 shown, the device includes:
[0142] An acquisition module 302, configured to acquire a sample video, and obtain a first video frame sequence, a second video frame sequence, and a sequence scrambling sample weight of the sample video according to the sample video, where image parameters of the first video frame sequence and the second video frame sequence are different;
[0143] The input module 304 is configured to input the first video frame sequence and the second video frame sequence into the feature extraction network to be trained, respectively, to obtain a first extraction result output by the feature extraction network to be trained, and to input the first video frame sequence and the second video frame sequence into a reference feature extraction model, respectively, to obtain a second extraction result output by the reference feature extraction model;
[0144] A first calculation module 306 is configured to obtain a comparison coding result, and calculate a first loss value based on the comparison coding result, the first extraction result, and the second extraction result;
[0145] A second calculation module 308 is configured to input the first extraction result into a disordered weight prediction network, obtain a sequence disorder prediction weight corresponding to the sample video output by the disordered weight prediction network, and calculate a second loss value based on the sequence disorder prediction weight and the sequence disorder sample weight;
[0146] The adjustment module 310 is configured to adjust the network parameters of the feature extraction network to be trained according to the first loss value and the second loss value until the training stop condition is reached to obtain a trained video feature extraction network.
[0147] Optionally, the acquisition module 302 is further configured to:
[0148] Sampling the sample video according to a preset sampling frequency to generate an initial video frame sequence;
[0149] Disrupting the order of video frames in the initial video frame sequence to obtain a disordered video frame sequence corresponding to the initial video frame sequence;
[0150] A first video frame sequence, a second video frame sequence and a sequence-shuffled sample weight of the sample video are generated based on the initial video frame sequence and the shuffled video frame sequence.
[0151] Optionally, the acquisition module 302 is further configured to:
[0152] Synthesize a to-be-enhanced video frame sequence according to the initial video frame sequence and the out-of-order video frame sequence;
[0153] Adjusting the image parameters of the to-be-enhanced video frame sequence according to a first preset enhancement rule to obtain a first video frame sequence;
[0154] Adjusting the image parameters of the to-be-enhanced video frame sequence according to a second preset enhancement rule to obtain a second video frame sequence;
[0155] The weight of the sequence out-of-order samples corresponding to the sample video is calculated according to the initial video frame sequence and the out-of-order video frame sequence.
[0156] Optionally, the first calculation module 306 is further configured to:
[0157] Read the comparison encoding result queue;
[0158] Obtain the comparison encoding result from the comparison encoding result queue.
[0159] Optionally, the first calculation module 306 is further configured to:
[0160] Construct a set of positive sample pairs and negative sample pairs according to the comparison encoding result, the first extraction result, and the second extraction result;
[0161] Calculate a first loss value according to the set of positive sample pairs and the set of negative sample pairs.
[0162] Optionally, the first calculation module 306 is further configured to:
[0163] Construct positive sample pairs according to the first extraction result and the second extraction result;
[0164] Construct a set of first negative sample pairs corresponding to the first extraction result according to the first extraction result and the comparison encoding result;
[0165] Construct a set of second negative sample pairs corresponding to the second extraction result according to the second extraction result and the comparison encoding result;
[0166] Calculate a first loss value according to the positive sample pairs, the set of first negative sample pairs, and the set of second negative sample pairs.
[0167] Optionally, the first calculation module 306 is further configured to:
[0168] The first extraction result includes a first feature vector extracted from a first video frame sequence and a second feature vector extracted from a second video frame sequence, and the second extraction result includes a third feature vector extracted from the first video frame sequence and a fourth feature vector extracted from the second video frame sequence;
[0169] Construct a first positive sample pair according to the first feature vector and the third feature vector;
[0170] Construct a second positive sample pair according to the second feature vector and the fourth feature vector.
[0171] Optionally, the first calculation module 306 is further configured to:
[0172] Construct a set of first negative sample pairs according to the first feature vector, the third feature vector, and the comparison encoding result;
[0173] Construct a second set of negative sample pairs based on the second eigenvector, the fourth eigenvector, and the comparison coding result.
[0174] Optionally, the adjustment module 310 is further configured to:
[0175] Perform gradient backpropagation according to the first loss value to adjust the network parameters of the to-be-trained feature extraction network.
[0176] Adjust the network parameters of the to-be-trained feature extraction network according to the second loss value.
[0177] Optionally, the device further includes: The addition module is configured to:
[0178] Input the first video frame sequence and the second video frame sequence into the dynamic encoder to obtain the comparison coding result corresponding to the sample video output by the dynamic encoder.
[0179] Add the comparison coding result corresponding to the sample video to the comparison coding result queue.
[0180] Optionally, the device further includes: The adjustment sub-module is configured to:
[0181] Adjust the model parameters of the reference feature extraction model and the model parameters of the dynamic encoder according to the first loss value.
[0182] Adjust the network parameters of the scrambled weight prediction network according to the second loss value.
[0183] A training device for a video feature extraction network provided in this specification includes: an acquisition module configured to acquire a sample video and obtain a first video frame sequence, a second video frame sequence, and a sequence scrambling sample weight of the sample video according to the sample video, where image parameters of the first video frame sequence and the second video frame sequence are different; an input module configured to respectively input the first video frame sequence and the second video frame sequence into a to-be-trained feature extraction network to obtain a first extraction result output by the to-be-trained feature extraction network, and respectively input the first video frame sequence and the second video frame sequence into a reference feature extraction model to obtain a second extraction result output by the reference feature extraction model; a first calculation module configured to obtain a comparison encoding result and calculate a first loss value based on the comparison encoding result, the first extraction result, and the second extraction result; a second calculation module configured to input the first extraction result into a scrambling weight prediction network to obtain a sequence scrambling prediction weight corresponding to the sample video output by the scrambling weight prediction network, and calculate a second loss value based on the sequence scrambling prediction weight and the sequence scrambling sample weight; an adjustment module configured to adjust network parameters of the to-be-trained feature extraction network according to the first loss value and the second loss value until a training stop condition is reached, and obtain a trained video feature extraction network. By respectively inputting the first video frame sequence and the second video frame sequence of the sample video into the to-be-trained feature extraction network and the reference feature extraction model, and combining two training tasks of contrast learning and cross-entropy learning, the training of the to-be-trained feature extraction network is completed. It is realized that a large number of unlabeled sample videos are used as training data, and a self-supervised training strategy is adopted to learn visual features, and a trained video feature extraction network is obtained. The video feature extraction network can be applied to downstream classification tasks, can effectively improve the accuracy of video classification, and can also be directly used for video feature extraction and then retrieve videos.
[0184] The above is a schematic solution of a training device for a video feature extraction network in this embodiment. It should be noted that the technical solution of the training device for the video feature extraction network and the technical solution of the above video feature extraction network training method belong to the same concept. For the details not described in detail in the technical solution of the training device for the video feature extraction network, reference can be made to the description of the technical solution of the above video feature extraction network training method.
[0185] Corresponding to the above method embodiment, this specification also provides an embodiment of a video classification device based on a video feature extraction network. Figure 4 FIG. shows a structural schematic diagram of a video classification device based on a video feature extraction network provided in an embodiment of this specification. As Figure 4 shown, the device includes:
[0186] An acquisition module 402, configured to acquire a video to be classified;
[0187] An input module 404, configured to input the video to be classified into a video feature extraction network obtained by training through a training method of a video feature extraction network, and obtain a target feature extraction result corresponding to the video to be classified output by the video feature extraction network;
[0188] An obtaining module 406, configured to input the target feature extraction result into a classifier, and obtain a classification result of the video to be classified output by the classifier.
[0189] A video classification device based on a video feature extraction network provided in this specification includes: an acquisition module, configured to acquire a video to be classified; an input module, configured to input the video to be classified into a video feature extraction network obtained by a training method of a video feature extraction network, and obtain a target feature extraction result corresponding to the video to be classified output by the video feature extraction network; an obtaining module, configured to input the target feature extraction result into a classifier, and obtain a classification result of the video to be classified output by the classifier. By extracting a classification feature vector of the video to be classified through a trained video feature extraction network, and inputting the classification feature vector into the classifier, the category of the video to be classified is determined according to the classification result output by the classifier.
[0190] The above is a schematic solution of a video classification device based on a video feature extraction network according to this embodiment. It should be noted that the technical solution of the video classification device based on the video feature extraction network and the technical solution of the above video classification method based on the video feature extraction network belong to the same concept. For the details not described in detail in the technical solution of the video classification device based on the video feature extraction network, reference can be made to the description of the technical solution of the above video classification method based on the video feature extraction network.
[0191] Figure 5 FIG. shows a structural block diagram of a computing device 500 according to an embodiment of this specification. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 through a bus 530, and a database 550 is used to store data.
[0192] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., Network Interface Card (NIC)), such as an IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0193] In one embodiment of this specification, the above components of the computing device 500 and Figure 5 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 5 the block diagram of the computing device shown is only for illustrative purposes and is not a limitation on the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0194] The computing device 500 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., tablet computer, personal digital assistant, laptop computer, notebook computer, netbook, etc.), a mobile phone (e.g., smartphone), a wearable computing device (e.g., smartwatch, smart glasses, etc.) or other types of mobile devices, or a stationary computing device such as a desktop computer or PC. The computing device 500 can also be a mobile or stationary server.
[0195] Wherein, when the processor 520 executes the computer instructions, it implements the steps of the training method of the video feature extraction network and the video classification method based on the video feature extraction network.
[0196] The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above training method of the video feature extraction network and the video classification method based on the video feature extraction network belong to the same concept. For the details not described in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above training method of the video feature extraction network and the video classification method based on the video feature extraction network.
[0197] One embodiment of this specification also provides a computer-readable storage medium, which stores computer instructions that, when executed by a processor, implement the steps of the training method of the video feature extraction network and the video classification method based on the video feature extraction network as described above.
[0198] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above-mentioned training method of the video feature extraction network and the video classification method based on the video feature extraction network belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above-mentioned training method of the video feature extraction network and the video classification method based on the video feature extraction network.
[0199] An embodiment of this specification also provides a computer program. When the computer program is executed on a computer, the computer is made to execute the steps of the above-mentioned training method of the video feature extraction network and the video classification method based on the video feature extraction network.
[0200] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solutions of the above-mentioned training method of the video feature extraction network and the video classification method based on the video feature extraction network belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the descriptions of the technical solutions of the above-mentioned training method of the video feature extraction network and the video classification method based on the video feature extraction network.
[0201] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0202] The computer instructions include computer program code, which may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0203] It should be noted that, for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0204] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0205] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, according to the content of the embodiments of this specification, many modifications and changes can be made. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A training method for a video feature extraction network, comprising: Obtaining a sample video, and obtaining a first video frame sequence, a second video frame sequence and a sequence scrambling sample weight of the sample video according to the sample video, wherein image parameters of the first video frame sequence and the second video frame sequence are different. Obtaining the first video frame sequence, the second video frame sequence and the sequence scrambling sample weight of the sample video according to the sample video includes: performing video sampling on the sample video according to a preset sampling frequency to generate an initial video frame sequence; scrambling the order of video frames in the initial video frame sequence to obtain a scrambled video frame sequence corresponding to the initial video frame sequence; synthesizing a to-be-enhanced video frame sequence according to the initial video frame sequence and the scrambled video frame sequence; adjusting image parameters of the to-be-enhanced video frame sequence according to a first preset enhancement rule to obtain a first video frame sequence; adjusting image parameters of the to-be-enhanced video frame sequence according to a second preset enhancement rule to obtain a second video frame sequence; calculating the sequence scrambling sample weight corresponding to the sample video according to the initial video frame sequence and the scrambled video frame sequence; Inputting the first video frame sequence and the second video frame sequence into a to-be-trained feature extraction network respectively to obtain a first extraction result output by the to-be-trained feature extraction network, and inputting the first video frame sequence and the second video frame sequence into a reference feature extraction model respectively to obtain a second extraction result output by the reference feature extraction model; Obtaining a comparison coding result, and calculating a first loss value based on the comparison coding result, the first extraction result and the second extraction result; Inputting the first extraction result into a scrambled weight prediction network to obtain a sequence scrambling prediction weight corresponding to the sample video output by the scrambled weight prediction network, and calculating a second loss value based on the sequence scrambling prediction weight and the sequence scrambling sample weight; Adjusting network parameters of the to-be-trained feature extraction network according to the first loss value and the second loss value until a training stop condition is reached, to obtain a trained video feature extraction network.
2. The method according to claim 1, wherein obtaining a comparison coding result includes: Reading a comparison coding result queue; Obtaining a comparison coding result from the comparison coding result queue.
3. The method according to claim 1, wherein calculating a first loss value based on the comparison coding result, the first extraction result and the second extraction result includes: Constructing a set of positive sample pairs and negative sample pairs according to the comparison coding result, the first extraction result and the second extraction result; Calculating a first loss value according to the set of positive sample pairs and negative sample pairs.
4. The method according to claim 3, wherein constructing a set of positive sample pairs and negative sample pairs according to the comparison coding result, the first extraction result and the second extraction result includes: Constructing positive sample pairs according to the first extraction result and the second extraction result; Constructing a first set of negative sample pairs corresponding to the first extraction result according to the first extraction result and the comparison coding result; Construct a second negative sample pair set corresponding to the second extraction result according to the second extraction result and the comparison coding result; Calculate a first loss value according to the positive sample pair, the first negative sample pair set, and the second negative sample pair set.
5. The method according to claim 4, wherein the first extraction result includes a first feature vector extracted from a first video frame sequence and a second feature vector extracted from a second video frame sequence, and the second extraction result includes a third feature vector extracted from the first video frame sequence and a fourth feature vector extracted from the second video frame sequence; Constructing a positive sample pair according to the first extraction result and the second extraction result includes: Construct a first positive sample pair according to the first feature vector and the third feature vector; Construct a second positive sample pair according to the second feature vector and the fourth feature vector.
6. The method according to claim 5, constructing a first negative sample pair set corresponding to the first extraction result according to the first extraction result and the comparison coding result, and constructing a second negative sample pair set corresponding to the second extraction result according to the second extraction result and the comparison coding result, includes: Construct a first negative sample pair set according to the first feature vector, the third feature vector, and the comparison coding result; Construct a second negative sample pair set according to the second feature vector, the fourth feature vector, and the comparison coding result.
7. The method according to claim 1, adjusting the network parameters of the to-be-trained feature extraction network according to the first loss value and the second loss value, includes: Perform gradient backpropagation according to the first loss value to adjust the network parameters of the to-be-trained feature extraction network; Adjust the network parameters of the to-be-trained feature extraction network according to the second loss value.
8. The method according to claim 1, the method further includes: Input the first video frame sequence and the second video frame sequence into a dynamic encoder to obtain a comparison coding result corresponding to the sample video output by the dynamic encoder; Add the comparison coding result corresponding to the sample video to the comparison coding result queue.
9. The method according to claim 8, the method further includes: Adjust the model parameters of the reference feature extraction model and the model parameters of the dynamic encoder according to the first loss value; Adjust the network parameters of the shuffled weight prediction network according to the second loss value.
10. A video classification method based on a video feature extraction network, includes: Obtain a video to be classified; Input the video to be classified into a video feature extraction network trained according to any one of claims 1-9 to obtain a target feature extraction result corresponding to the video to be classified output by the video feature extraction network; Input the target feature extraction result into a classifier to obtain a classification result of the video to be classified output by the classifier.
11. A computing device, including a memory, a processor, and computer instructions stored on the memory and executable on the processor, the processor implements the steps of the method according to any one of claims 1-9 or 10 when executing the computer instructions.
12. A computer-readable storage medium stores computer-executable instructions, and when the computer instructions are executed by a processor, the steps of the method according to any one of claims 1-9 or 10 are implemented.
Citation Information
Patent Citations
Feature extraction network training method and device, terminal equipment and medium
CN114419489A
JP0013418A1