Method and apparatus for training a video classification model
By generating virtual video samples through label conversion and fusion processing of video samples, the problem of insufficient generalization ability of models in small sample video classification is solved, and high-accuracy classification is achieved on new class video samples.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG E COMMERCE BANK CO LTD
- Filing Date
- 2023-05-04
- Publication Date
- 2026-07-21
AI Technical Summary
In the process of classifying videos with few samples, existing technologies are unable to effectively improve the generalization ability and robustness of the model, especially when the number of new class video samples is limited, resulting in insufficient classification accuracy.
By performing label conversion on the paired labels of test video samples and training video samples to generate category labels, and using fusion parameters to fuse the test video samples and training video samples to generate virtual video samples, these virtual video samples are input into the video classification model for training and adjustment, thereby improving the model's generalization ability and robustness.
Under small sample conditions, data augmentation techniques generate more virtual video samples, which improves the classification accuracy and robustness of the video classification model and enhances the model's generalization ability on new types of video samples.
Smart Images

Figure CN122435508A_ABST
Abstract
Description
[0001] This patent application is a divisional application of Chinese patent application No. 2023105077742, filed on May 4, 2023, entitled "Training Method and Apparatus for Video Classification Model". Technical Field
[0002] This document relates to the field of data processing technology, and in particular to a training method and apparatus for a video classification model. Background Technology
[0003] With the development of network technology, information networks have become an important part of life, and users requesting and interacting with services online has become a mainstream trend. More and more services or capabilities are being provided to users online. This also places higher demands on the service capabilities of service providers. How to effectively process data during the service provision process to obtain more reliable and accurate service results is an increasingly important focus for both service providers and users. Summary of the Invention
[0004] This specification provides one or more embodiments of a method for training a video classification model. The method includes: sampling a video set to obtain test video samples and at least one training video sample, and generating paired labels for the test video samples and each training video sample; performing label conversion based on the paired labels of the test video samples and each training video sample to obtain model parameters and category labels for the test video samples and the at least one training video sample; fusing the test video samples and each training video sample according to the fusion parameters and the category labels to obtain at least one virtual video sample; inputting the at least one virtual video sample into a video classification model constructed based on the model parameters for video classification, and adjusting the parameters of the video classification model based on the classification results of each virtual video sample to obtain a video classification model.
[0005] This specification provides one or more embodiments of a video classification processing method, comprising: sampling a target video set to obtain a test video and at least one training video; fusing the test video and each training video according to fusion parameters to obtain at least one virtual video corresponding to the test video; inputting the at least one virtual video into a video classification model for video classification to obtain classification results for each virtual video; and calculating the video classification result of the test video based on the classification results of each virtual video. The video classification model is obtained by training a video classification model based on model parameters using at least one virtual video sample; the model parameters are obtained after label conversion based on the paired labels of the test video sample and each training video sample.
[0006] This specification provides one or more embodiments of a training apparatus for a video classification model, comprising: a paired label generation module configured to sample a video set to obtain test video samples and at least one training video sample, and generate paired labels for the test video samples and each training video sample; a label conversion module configured to perform label conversion based on the paired labels of the test video samples and each training video sample to obtain model parameters and category labels for the test video samples and the at least one training video sample; a fusion processing module configured to perform fusion processing on the test video samples and each training video sample according to the fusion parameters and the category labels to obtain at least one virtual video sample; and a parameter adjustment module configured to input the at least one virtual video sample into a video classification model constructed based on the model parameters for video classification, and adjust the parameters of the video classification model based on the classification results of each virtual video sample to obtain a video classification model.
[0007] This specification provides one or more embodiments of a video classification processing apparatus, comprising: a sampling module configured to sample a target video set to obtain a test video and at least one training video; a fusion processing module configured to fuse the test video and each training video according to fusion parameters to obtain at least one virtual video corresponding to the test video; a video classification module configured to input the at least one virtual video into a video classification model for video classification to obtain classification results for each virtual video; and a result calculation module configured to calculate the video classification result of the test video based on the classification results of each virtual video. The video classification model is obtained by training a video classification model based on model parameters using at least one virtual video sample; the model parameters are obtained after label conversion based on the paired labels of the test video sample and each training video sample.
[0008] This specification provides one or more embodiments of a training device for a video classification model, comprising: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: sample a video set to obtain test video samples and at least one training video sample, and generate paired labels for the test video samples and each training video sample; perform label conversion based on the paired labels of the test video samples and each training video sample to obtain model parameters and category labels for the test video samples and the at least one training video sample; perform fusion processing on the test video samples and each training video sample according to fusion parameters and the category labels to obtain at least one virtual video sample; input the at least one virtual video sample into a video classification model constructed based on the model parameters for video classification, and adjust the parameters of the video classification model based on the classification results of each virtual video sample to obtain a video classification model.
[0009] This specification provides one or more embodiments of a video classification processing device, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: sample a target video set to obtain a test video and at least one training video; fuse the test video and each training video according to fusion parameters to obtain at least one virtual video corresponding to the test video; input the at least one virtual video into a video classification model for video classification to obtain classification results for each virtual video; and calculate the video classification result of the test video based on the classification results of each virtual video. The video classification model is obtained by training a video classification model based on model parameters using at least one virtual video sample; the model parameters are obtained after label conversion based on paired labels of the test video sample and each training video sample.
[0010] This specification provides one or more embodiments of a storage medium for storing computer-executable instructions that, when executed by a processor, implement the following process: sampling a video set to obtain test video samples and at least one training video sample, and generating paired labels for the test video samples and each training video sample; performing label conversion based on the paired labels of the test video samples and each training video sample to obtain model parameters and category labels for the test video samples and the at least one training video sample; fusing the test video samples and each training video sample according to the fusion parameters and the category labels to obtain at least one virtual video sample; inputting the at least one virtual video sample into a video classification model constructed based on the model parameters for video classification, and adjusting the parameters of the video classification model based on the classification results of each virtual video sample to obtain a video classification model.
[0011] This specification provides one or more embodiments of another storage medium for storing computer-executable instructions that, when executed by a processor, implement the following process: sampling a target video set to obtain a test video and at least one training video; fusing the test video and each training video according to fusion parameters to obtain at least one virtual video corresponding to the test video; inputting the at least one virtual video into a video classification model for video classification to obtain classification results for each virtual video; and calculating the video classification result of the test video based on the classification results of each virtual video. The video classification model is obtained by training a video classification model based on model parameters using at least one virtual video sample; the model parameters are obtained after label conversion based on the paired labels of the test video sample and each training video sample. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 A schematic diagram illustrating an implementation environment provided for one or more embodiments of this specification; Figure 2 A flowchart illustrating a training method for a video classification model provided in one or more embodiments of this specification; Figure 3A flowchart illustrating a training method for a video classification model applied to training and testing scenarios, provided for one or more embodiments of this specification; Figure 4 A flowchart illustrating a video classification processing method provided in one or more embodiments of this specification; Figure 5 A schematic diagram of an embodiment of a training device for a video classification model provided in one or more embodiments of this specification; Figure 6 A schematic diagram of an embodiment of a video classification processing device provided in one or more embodiments of this specification; Figure 7 A schematic diagram of the structure of a training device for a video classification model provided in one or more embodiments of this specification; Figure 8 This is a schematic diagram of the structure of a video classification processing device provided for one or more embodiments of this specification. Detailed Implementation
[0013] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.
[0014] like Figure 1 As shown, in one or more embodiments of this specification, the implementation environment includes the PyTorch deep learning framework.
[0015] In this implementation environment, two phases are executed: meta-training and meta-testing. In the meta-training phase, a query video sample set and a support video sample set are constructed based on the base class video sample set. The query video samples are fused using the support video samples to obtain a number of virtual video samples corresponding to the number of support video samples. In this way, the data of the query video samples is augmented, and more virtual video samples are obtained to train the video classification model to be trained.
[0016] During the meta-testing phase, query videos and support video sets are constructed based on the new type of video sample set. Data augmentation is performed on the query videos based on the support video set to obtain multiple virtual videos. The classification results of each virtual video are then used by a video classification model to calculate the classification results of the query videos, thereby improving the accuracy of the obtained classification results of the query videos.
[0017] like Figure 1 As shown, in the meta-training phase, the query video sample set Q is obtained from a large set of labeled base class video samples. n and supporting video sample set S n Taking a query video sample and two support video samples as an example, the query video sample is fused with support video sample 1 to obtain virtual video sample 1, and the query video sample is fused with support video sample 2 to obtain virtual video sample 2. Based on virtual video sample 1 and virtual video sample 2, the video classification model to be trained is trained to obtain the trained video classification model. During the meta-testing phase, a new class of videos, triggered by a small amount of labeled data, is used to obtain the query video set Q. n and supporting video set S n Taking a query video and two supporting videos as an example, the query video and supporting video 1 are fused to obtain virtual video 1, and the query video and supporting video 2 are fused to obtain virtual video 2. The similarity between virtual video 1 and supporting video 1, and between virtual video 1 and supporting video 2, are calculated, as are the similarity between virtual video 2 and supporting video 1, and between virtual video 2 and supporting video 2. Figure 1 As shown, the values above the arrows represent similarity scores; the similarity score of the query video is calculated based on each similarity score.
[0018] This specification provides one or more embodiments of a training method for a video classification model, as follows: This embodiment provides a training method for a video classification model. It starts by performing label transformation on the paired labels of test video samples and training video samples, obtaining model parameters during the label transformation process. By fusing the parameters and the category labels of the test video samples and training video samples obtained from the label transformation, the test video samples and training video samples are fused to obtain virtual video samples. Based on the classification results of the virtual video samples by the video classification model constructed based on the model parameters, the parameters of the video classification model are adjusted to obtain a trained video classification model. In this way, during the classification of small sample videos, data augmentation is performed on the test video samples to obtain richer samples, i.e., virtual video samples. This improves the generalization ability and robustness of the trained video classification model despite the limited sample size.
[0019] Reference Figure 2 The training method for the video classification model provided in this embodiment specifically includes steps S202 to S208.
[0020] Step S202: Sample the video set to obtain test video samples and at least one training video sample, and generate paired labels for the test video samples and each training video sample.
[0021] In practical applications, during the classification of small sample videos, a large number of labeled base class samples and a small number of labeled new class samples are usually obtained. The goal of training is to enable the model trained on base class samples to obtain more accurate classification results when tested on new class samples. For a classification with a certain number of categories and a task of k support samples, the following steps are taken: construct a support video sample set, a query video sample set, and paired labels for query video samples in the query video sample set and each support video sample in the support video sample set. Based on these paired labels, data augmentation is performed on the query video samples to obtain more samples for model training.
[0022] In this embodiment, the video set includes a base class video sample set, which is the sample set accessible during training; the test video samples include query video samples; the training video samples include support video samples; and the paired labels are determined based on whether the paired query video samples and training video samples belong to the same category. Specifically, during model training, sampling is performed from the base class video sample set to obtain the test video sample set and the training video sample set. It should be noted that the base class video sample set in this embodiment includes a set containing a large number of labeled samples obtained from third-party channels; for example, the Kinetics dataset and the Something V2 dataset; the new class video sample set in this embodiment includes a set of a small number of labeled video samples; for example, in a user self-verification scenario, a video submitted by the user showing farmland; or in a resource lending scenario, a video taken by the user showing their farmland or house.
[0023] Optionally, the test video sample set includes test video samples of video categories for a given number of classifications, with the number of test video samples in each video category being a preset threshold; the training video sample set includes training video samples of the preset number of categories, with the number of training video samples in each video category exceeding the preset threshold. Optionally, the number of classifications corresponds to the number of classifications performed; for example, binary classification results in 2 classifications; 5-class classification results in 5 classifications. It should be noted that in this embodiment, the video categories are the categories of each sample in the few-shot learning dimension.
[0024] In the process of constructing the test video sample set and the training video sample set, the construction of the test video sample set and the training video sample set is based on the number of categories; for example, for a task with n categories and k support samples, the constructed test video sample set (query video sample set) contains n categories, each with one video sample, and the constructed training video sample set (support video sample set) contains n categories, each with k video samples.
[0025] In specific implementation, to improve the effectiveness of the sampled test video samples and training video samples, this embodiment provides an optional implementation method in which the following operations are performed during the process of sampling the video set to obtain test video samples and at least one training video sample, and generating paired labels for the test video samples and each training video sample: Based on the number of categories, the video set is sampled to obtain a test video sample set consisting of multiple test video samples and a training video sample set consisting of the at least one training video sample. Generate paired labels between each test video sample and each training video sample in the test video sample set.
[0026] Specifically, a test video sample set is formed by sampling a first number of test video samples from the video set, and a training video sample set is formed by sampling a second number of training video samples from the video set; optionally, the first number is equal to the number of categories, and the video categories of each test video sample in the test video sample set are different; the second number is equal to the product of the number of categories and the number of samples; the video categories of the training video samples included in the training video sample set are the same as the video categories of the test video samples included in the test video sample set, and the number of training videos under each video category in the training video sample set is equal to and equal to the number of samples.
[0027] After obtaining the test video sample set and the training video sample set, the test video samples in the test video sample set are paired with the test video samples in the training video sample set, and paired labels are generated. It should be noted that after obtaining the test video sample set and the training video sample set, the processing procedure for each test video sample in the test video sample set is similar. Therefore, this embodiment uses a single test video sample as an example to specifically explain the data augmentation and model training process.
[0028] In the specific implementation process, in order to improve the accuracy of the paired labels, this embodiment provides an optional implementation method in which the paired labels of the test video samples and each training video sample are generated in the following way: The paired labels of the test video sample and the training video sample belonging to the same video category as the test video sample are determined as the first label; The paired labels of the test video sample and the training video sample belonging to a different video category from the test video sample are determined as the second label.
[0029] Specifically, the paired labels for the test video samples and the training video samples are determined based on whether they belong to the same video category.
[0030] For example, the constructed query video sample set contains 5 query video samples, whose video categories are Category 1, Category 2 through Category 5, respectively. The constructed support video sample set contains 25 support video samples, where each of the 5 categories (Category 1 through Category 5) contains 5 support video samples. That is, among the support video samples, 5 support video samples belong to Category 1, 5 support video samples belong to Category 2, ..., and 5 support video samples belong to Category 5. In this case, during the generation of paired labels, each query video sample in the query video sample set has a paired label of 1 with a support video sample in the same video category as the support video sample set, and a paired label of 0 with other support video samples in the support video sample set. In other words, each query video sample has a paired label of 1 with the 5 support video samples and a paired label of 0 with the 20 support video samples.
[0031] In addition, step S202 can be replaced by sampling in the video set according to the number of categories to obtain a test video sample set and a training video sample set, and generating paired labels for each test video sample in the test video sample set and each training video sample in the training video sample set, and forming a new implementation method with other processing steps provided in this embodiment.
[0032] Step S204: Perform label conversion based on the paired labels of the test video samples and each training video sample to obtain model parameters and category labels of the test video samples and at least one training video sample.
[0033] To achieve data augmentation of test video samples based on training video samples, test video samples and training video samples are fused to obtain more virtual video samples. The model is then trained based on these virtual video samples to improve the generalization ability and robustness of the trained video classification model.
[0034] In the process of fusing test video samples and training video samples, because few-shot learning uses a meta-learning paradigm, while fusion is performed using a traditional classification paradigm, it's impossible to directly use paired labels for fusion in a few-shot meta-learning scenario. Therefore, label conversion is necessary, transforming paired labels into category labels, and then fusing the test video samples and training video samples based on these category labels. In other words, in this embodiment, to generate more virtual video samples during the meta-training stage, traditional classification fusion is introduced into meta-training. However, the prerequisite for fusion is the existence of a classifier structure in the network. Since meta-learning training maximizes the similarity between similar query video samples and supporting video samples, the model often does not have a classifier structure. Therefore, label conversion is required before fusion.
[0035] In this embodiment, the category labels include the category labels of the test video samples and training video samples in the traditional classification dimension.
[0036] In this embodiment, training a video classification model requires constructing a classifier; the model parameters include the classification weights of the classifier generated during the label conversion process. Once the classification weights are obtained, the classifier can be constructed.
[0037] In the specific process of label conversion, that is, converting paired labels into category labels, in order to improve the accuracy and effectiveness of the converted category labels, this embodiment provides an optional implementation method, which adopts the following approach to realize the label conversion based on the paired labels of the test video samples and each training video sample, and obtain the model parameters and the category labels of the test video samples and at least one training video sample: Read the first similarity algorithm for calculating positive sample similarity under the paired label dimension, and the second similarity algorithm for calculating negative sample similarity; The third similarity algorithm is used to calculate the similarity of video samples of the same category under the category label dimension, and the fourth similarity algorithm is used to calculate the similarity of video samples of different categories. Based on the test video samples and the training video sample set, a first classification weight is calculated according to the first similarity algorithm and the third similarity algorithm, and a second classification weight is calculated according to the second similarity algorithm and the fourth similarity algorithm, and the first classification weight and the second classification weight are determined as the model parameters.
[0038] Optionally, the first classification weight includes the classification weight corresponding to the category to which the test video sample belongs; The second classification weight includes the classification weights corresponding to categories other than the target category; the target number includes the number of categories of the classifier minus 1.
[0039] Specifically, for the test video samples, positive and negative samples were obtained. The positive samples of the test video samples are training video samples in the training video sample set that belong to the same video category as the test video samples; that is, training video samples whose paired label is 1. The negative samples of the test video samples are training video samples in the training video sample set that belong to different video categories than the test video samples; that is, training video samples whose paired label is 0.
[0040] For test video samples, the similarity between the test video samples and positive and negative samples can be calculated; For example, the similarity between a test video sample and a positive sample can be calculated as follows:
[0041] The similarity between the test video sample and the negative sample can be calculated as follows:
[0042] Where x represents the test video sample, x i x represents the positive samples of the test video samples. j This represents a negative sample from the test video sample. This indicates the similarity between the test video sample and the positive sample. This indicates the similarity between the test video and the negative sample.
[0043] The above describes the process of calculating the similarity between test video samples and positive and negative samples based on paired labels. The following describes the calculation of inter-class similarity and intra-class similarity based on category labels. Inter-class similarity refers to the similarity between test video samples and training video samples under different category labels; intra-class similarity refers to the similarity between test video samples and training video samples under the same category label. For example, the intra-class similarity of test video samples can be calculated as follows:
[0044] The intra-class similarity of test video samples can be calculated as follows:
[0045] Among them, w y This represents the classification weight corresponding to the category to which the test video sample belongs; w j This represents the classification weight for categories other than the category to which the test video sample belongs.
[0046] In other words, by constructing a classifier, paired labels are converted into category labels. Specifically, under the above constraints, the classification probabilities of positive and negative samples under the constructed classifier remain unchanged, and the label conversion can be completed. This can be achieved by calculating each weight of the classifier one by one, so that its inner product with its corresponding query video sample is the same as that calculated using paired labels.
[0047] In practice, the classifier is trained with the constraints that the intra-class similarity is equal to the similarity between the test video sample and the positive sample, and the inter-class similarity is equal to the similarity between the test video sample and the negative sample, to obtain the classification weights of the classifier. In other words, a classifier is built such that the similarity of a query video sample after passing through this classifier is consistent with the similarity of the query video sample with positive samples and with negative samples. In addition, in this embodiment, the model parameters can also be generated in the following way: Based on the paired labels of the test video samples and each training video, a first training video sample that is a positive sample to the test video sample and a second training video sample that is a negative sample to the test video sample are determined. The similarity between the test video samples and the first training video samples is calculated, and the similarity between the test video samples and the second training video samples is calculated. A classifier is constructed using the positive sample similarity and the negative sample similarity as constraints, and the classification weights of the classifier are obtained.
[0048] Based on the classification weights obtained from the classifier, each training video sample and test video sample is input into the classifier to classify the categories and obtain the category labels.
[0049] Step S206: According to the fusion parameters and the category labels, the test video sample and each training video sample are fused to obtain at least one virtual video sample.
[0050] The fusion parameters include the fusion ratio during the fusion process of the test video samples and training video samples; the virtual video samples include video samples obtained after fusion processing of the test video samples and each training video sample. Optionally, the fusion process includes: pixel-by-pixel interpolation of two images of the same size (test video samples and training video samples).
[0051] In the above steps, after label conversion based on the paired labels of the test video samples and each training video sample, category labels for the test video samples and training video samples are obtained. Optionally, the category labels include labels corresponding to the sample categories of the test video samples and each training video sample. It should be noted that the category labels in this embodiment include labels for the linear categories of samples in the traditional learning dimension.
[0052] In this step, in order to obtain more samples, after obtaining the category labels of the test video samples and each training video sample, the test video samples and each training video sample are fused based on the category labels to obtain at least one virtual video sample.
[0053] To improve the effectiveness of the virtual video samples obtained after fusion processing, the fusion parameters are changed sequentially without constraining the rate of change, thereby generating more diverse samples. In one optional implementation of this embodiment, the following operations are performed during the process of fusing test video samples with each training video sample according to the fusion parameters and category labels to obtain at least one virtual video sample: According to the fusion ratio, the test video sample and each image frame of each training video sample are fused to obtain at least one virtual video sample. According to the fusion ratio, the category labels of each image frame of the test video sample and the training video sample are fused to obtain the virtual category label of each virtual video sample.
[0054] Specifically, according to the fusion ratio corresponding to each image frame, the image frames of the test video sample and each training video sample are fused together. Furthermore, the category labels of the test video sample and each training video sample are fused together to obtain at least one virtual video sample and virtual category labels for each virtual video sample. Optionally, the virtual category label refers to the category label of the virtual video sample in the linear dimension.
[0055] The fusion process between test video samples and training video samples can be performed as follows:
[0056] When merging category tags, the following formula can be used:
[0057] in, The virtual video sample obtained by fusion is represented by x, where k represents the k-th image frame; i x represents the test video sample. j Indicates training video samples, This indicates the category label of the test video sample. λ represents the class label of the training video samples. k represents the mixing ratio of the k-th image frame, and T represents the total number of image frames in the test video sample.
[0058] It should be noted that, among them, λ k It can be calculated based on preset constants, the temporal parameters of the k-th frame of the test video samples and the training video samples. For example, the ratio of the temporal parameters of the k-th frame of the test video samples to the temporal parameters of the k-th frame of the training video samples can be calculated, and then the product of the preset constant and this ratio can be used as λ. k The above explanation of the calculation of the mixing ratio is merely illustrative; furthermore, λ k Other relevant parameters can also be used for calculation, but this embodiment does not limit the calculation.
[0059] In addition, step S206 can be replaced by fusing the test video sample with each training video sample according to the fusion parameters to obtain at least one virtual video sample, and fusing the category labels of the test video sample with each training video sample to obtain the category labels of each virtual video sample, and forming a new implementation method with other processing steps provided in this embodiment.
[0060] It should be noted that the above-described fusion process describes the fusion of one test video sample and one training video sample. In actual execution, the above method is used to fuse each test video and each training video. For example, one test video sample is fused with five training video samples to obtain five virtual video samples.
[0061] It should also be noted that this embodiment only provides one method for obtaining virtual video samples: fusing test video samples with each training video sample. In addition, virtual video samples can also be obtained by partially replacing test video samples and training video samples; for example, replacing frames 6-10 in a 10-frame test video sample with frames 6-10 in a training video sample to obtain virtual video samples; furthermore, discontinuous image frames can also be replaced; specifically, virtual video samples can also be obtained by using other methods to perform data augmentation on test video samples and training video samples.
[0062] In addition, step S206 can also be replaced by performing data augmentation on the test video samples based on each training video sample and the category label to obtain at least one virtual video sample, and forming a new implementation method with other processing steps provided in this embodiment.
[0063] Step S208: Input the at least one virtual video sample into the video classification model constructed based on the model parameters for video classification, and adjust the parameters of the video classification model based on the classification results of each virtual video sample to obtain a video classification model.
[0064] In the above steps, at least one virtual video sample is obtained. After obtaining at least one virtual video sample, the video classification model to be trained is trained based on the virtual video sample. That is, the at least one virtual video sample is input into the video classification model constructed based on the model parameters to classify the video, and the parameters of the video classification model are adjusted based on the classification results of each virtual video sample to obtain the trained video classification model.
[0065] Optionally, the video classification model includes a sampling layer, a feature extraction layer, and a classification sub-model constructed based on the model parameters. In this embodiment, to train the video classification model to be trained, at least one virtual video sample is used as the training sample. Specifically, each virtual video sample is input into the video classification model to be trained, that is, the video classification model constructed with the model parameters is used for video classification. In one optional implementation of this embodiment, the video classification of any virtual video sample is achieved in the following way: Image sampling is performed on any virtual video sample to obtain a number of sampled image frames, and the image frames are input into the feature extraction layer for feature extraction to obtain image frame features; The image frame features are input into the classification sub-model constructed by the model parameters to perform video classification on any virtual video sample, thereby obtaining the classification result of any virtual video sample.
[0066] Optionally, the classification sub-model performs video classification on any of the virtual video samples, including: The similarity between the image frame features and at least one training video feature under each category is calculated to obtain the similarity between the image frame features and each training video feature under each category. Based on the similarity, the probability of matching any virtual video sample with each of the categories is calculated, and the calculated category matching probability is used as the classification result.
[0067] Specifically, for any virtual video sample, the virtual video sample is input into the video classification model to be trained. The video classification model first performs image sampling on the virtual video sample based on the sampling layer to obtain the number of sampled image frames. Then, the sampled image frames are input into the feature extraction layer for feature extraction to obtain image frame features. Finally, the image frame features are input into the classification sub-model constructed based on the model parameters to perform video classification and obtain the classification result.
[0068] For example, a video classification model consists of three parts: a sampling layer called TSN (temporal segment network), a feature extraction layer called ResNet50 (residual neural network), and a classifier constructed from the aforementioned classification weights.
[0069] The classification sub-model, also known as the classifier, calculates the similarity between the image frame features of the virtual video sample and the features of each training video during the video classification process, thereby obtaining the similarity between the image frame features and the features of each training video; based on the similarity, it calculates the matching probability between the virtual video sample and each category.
[0070] In this embodiment, when there is only one training video feature under each category, the similarity between the image frame feature and the training video feature is calculated, and the matching probability between the virtual video sample and each category is calculated based on this similarity. If each category contains multiple training video features, the average similarity between the image frame and the multiple training video features is calculated, and the matching probability between the virtual video sample and each category is calculated based on the average similarity. In calculating the category matching probability, the similarity between the virtual video sample and each category can be proportionally divided to obtain the category matching probability for each category; other processing methods can also be used to calculate the category matching probability, which is not limited in this embodiment.
[0071] In specific implementation, to train the video classification model, after obtaining the classification results of each virtual video sample, the parameters of the video classification model are adjusted based on the classification results of each virtual video sample. To improve the robustness and generalization ability of the trained video classification model, in an optional implementation method provided in this embodiment, the following operations are performed during the parameter adjustment process of the video classification model based on the classification results of each virtual video sample: The training loss is calculated based on each virtual video sample and the classification result of each virtual video sample, and the parameters of the video classification model constructed based on the training loss are adjusted.
[0072] In one optional implementation of this embodiment, the process of calculating the training loss based on each virtual video sample and its classification result includes: The classification result parameters in the classification result of the first virtual video sample are corrected; the first virtual video sample is obtained by fusing test video samples and training video samples of different video categories; The training loss is calculated based on the classification results after parameter correction, the classification results of the second virtual sample, and the category labels of each virtual video sample; the second virtual video sample is obtained by fusing test video samples and support video samples of the same video category.
[0073] Optionally, the classification result parameters in the classification result of the first virtual video sample are corrected, including: Correction coefficients are calculated based on the fusion parameters for parameter adjustment; The classification prediction result is parameter-corrected based on the category label, number of categories, and correction coefficient corresponding to the first virtual video sample.
[0074] In practice, introducing virtual video samples aims to reduce the overconfidence of video classification models. By introducing uncertainty into the training process through category labels, the model is encouraged to achieve better generalization performance. The goal is to further enhance this uncertainty by mixing test and training video samples from different categories. For example, if the query video sample belongs to the ballet category, and the supporting video sample also belongs to this category, the video classification model should make a more confident judgment on the virtual video sample obtained by fusing the query and supporting video samples because they share many similar features. However, if the supporting video sample belongs to a different category, there are fewer similar features, making it more difficult for the model to distinguish the category of the virtual video sample obtained by fusing the query and supporting video samples from another category. Therefore, when the test and training video samples being fused belong to different categories, a correction coefficient is first calculated.
[0075] in, It belongs to the correction factor. It is a constant. This refers to the mixing ratio.
[0076] After calculating the correction coefficient, the predicted classification results are corrected in the following way:
[0077] Where n is the number of categories.
[0078] After calculating the classification result of the first virtual video sample after correction, the training loss is calculated based on the classification result of the first virtual video sample after parameter correction, the classification result of the second virtual video sample, and the category labels of the first and second virtual video samples. The parameters of the video classification model to be trained are then corrected based on the training loss to obtain the trained video classification model.
[0079] The above process illustrates the data augmentation of query video samples during meta-training. Meta-learning also includes a meta-testing phase; the following section provides a detailed explanation of the data augmentation process for query videos during the meta-testing phase.
[0080] In the meta-testing phase, the same backbone network as in the meta-training phase is retained. Therefore, this backbone network (video classification model) can correctly predict the fused virtual videos. Given a test video, L virtual videos can be generated using the training videos, where L is a hyperparameter in meta-learning. By analyzing the classification results of all virtual videos, a better classification prediction can be made for the test video.
[0081] To improve the accurate classification of test videos, this embodiment provides an optional implementation method that uses the following approach to perform the meta-testing process: Sample the target video set to obtain a test video and at least one training video; According to the fusion parameters, the test video and each training video are fused to obtain at least one virtual video corresponding to the test video; The video classification model obtained after training is used to input at least one virtual video to classify the video and obtain the classification results of each virtual video. The video classification result of the test video is calculated based on the classification results of each virtual video.
[0082] Optionally, the target video set may include a new type of video set.
[0083] In practical applications, the classification result of a single test sample is calculated in the following way:
[0084] in, This indicates the video classification result of the test video; it excludes the portion where the confidence level of the corresponding category increases due to the introduction of supporting video samples.
[0085] In practical applications, if fusion is performed randomly and without constraints, similar to meta-training, classification prediction will be more sensitive to noise when the proportion of query video samples is low. Therefore, a fixed fusion ratio greater than 0.5 is used, and the same number of support videos are selected for fusion for each category. Finally, the above formula can be simplified to the following form:
[0086] Where m represents the number of samples selected for fusion for each category.
[0087] Furthermore, in this embodiment, fusion processing can be performed in multiple dimensions or in multiple ways. The fusion processing is not limited to the data layer, but can also be applied to the feature map. For small sample video classification, since different frames of different videos are extracted through the same backbone network, the fusion processing can also be added to the backbone network. Alternatively, in addition to fusion processing of samples, more types of data augmentation methods can be introduced. For example, fusion processing can simply merge two samples pixel by pixel to generate virtual video samples without involving more complex operations such as cutting. Alternatively, a part of one sample can be swapped with another or more complex operations can be considered.
[0088] In this embodiment, a data augmentation method for meta-learning is proposed from the perspective of samples. By expanding the samples, the performance of the video classification model is improved. This method can be easily applied to any few-sample video classification method that uses meta-learning. By introducing very little computational overhead, the performance of the trained video classification model is steadily improved. This scheme achieves stable improvements across all evaluation metrics for two different few-sample video classification methods and three different few-sample video classification methods, as shown in the table below:
[0089] The values in parentheses represent the performance improvement of the video classification model trained in the above manner compared to the video classification model trained in the traditional manner.
[0090] In summary, the training method for the video classification model provided in this embodiment transforms pairwise learning into category learning, introducing the fusion method commonly used in traditional classification into meta-learning. For specific fusion methods, considering the characteristics of video samples and few-shot learning, temporal augmentation fusion and asymmetric fusion are proposed respectively. This preserves temporal diversity and performs asymmetric fusion processing on query and support video samples. Furthermore, a data augmentation method for the meta-testing stage is proposed, ultimately achieving significant performance improvements. It should be noted that this embodiment can be implemented using the PyTorch deep learning framework.
[0091] The following description uses the application of the video classification model training method provided in this embodiment in the training and testing scenarios of video classification models as an example to further illustrate the training method of the video classification model provided in this embodiment. (See also...) Figure 3 The training method for video classification models, applied to training and testing scenarios, includes the following steps.
[0092] Step S302: Sample the Kinetics dataset to obtain query video samples and at least one supporting video sample.
[0093] Step S304: Generate matching labels for the query video sample and each supporting video sample.
[0094] Step S306: Based on the similarity algorithm between the query video sample and each supporting video sample, and the similarity algorithm between the query video sample and each supporting video sample under the linear category label, calculate the classification parameters of the classifier and obtain the category labels of the query video sample and each supporting video sample.
[0095] Optionally, in calculating the classification parameters of the classifier, the classification weights of the classifier are calculated under the constraint that the similarity between the query video sample and the positive samples in the support video samples is equal to the similarity between the query video sample and the classification weight corresponding to the category to which the query video sample belongs in the classifier, and / or that the similarity between the query video sample and the negative samples in the support video samples is equal to the similarity between the query video sample and the classification weight corresponding to the category to which the query video sample belongs in the classifier. The classification parameters of the classifier include the classification weights corresponding to each category. The category label is the label of the category to which each video sample belongs under the classifier.
[0096] Step S308: According to the fusion ratio, the query video sample and each supporting video sample are fused to obtain at least one virtual video sample; and the category labels of the query video sample and each supporting video sample are fused to obtain the category labels of each virtual video sample.
[0097] Step S310: Input each virtual video sample into the video classification model to be trained, which contains a classifier, to classify the video. Based on the category classification results and category labels of each virtual video sample, adjust the parameters of the video classification model to be trained to obtain the trained video classification model.
[0098] Step S312: Sample the set of user self-certification videos of the user self-certification service to obtain the query video and at least one supporting video.
[0099] Step S314: According to the fusion ratio, the query video and each supporting video are fused to obtain at least one virtual video corresponding to the query video.
[0100] Step S316: Input each virtual video into the video classification model to classify the video and obtain the classification results for each virtual video.
[0101] Step S318: Calculate the video classification result of the query video based on the classification results of each virtual video.
[0102] This specification provides one or more embodiments of a video classification processing method as follows: The relevant content in the video classification processing method provided in this embodiment is similar to the relevant content in the training method of the video classification model provided in the above embodiments. When reading this embodiment, please refer to the relevant content of the above embodiments or make adaptive modifications to the relevant content of the above embodiments. Correspondingly, when reading the above embodiments, you can also refer to the relevant content of this embodiment. This embodiment will not be described in detail here.
[0103] Reference Figure 4 The video classification processing method provided in this embodiment specifically includes steps S402 to S408.
[0104] Step S402: Sample the target video set to obtain a test video and at least one training video.
[0105] In this embodiment, the target video set includes a small set of labeled videos; for example, in a user self-verification scenario, a video submitted by the user showing farmland being cultivated; or in a resource lending scenario, a video taken by the user showing their farmland or house. In other words, the target video set in this embodiment includes video data submitted by the user for participating in the service. That is, step S402 can also be replaced by sampling the service video set of the target service to obtain test videos and at least one training video.
[0106] The test video includes videos used for video classification. The at least one training video includes videos that assist in the video classification of the test video; optionally, this embodiment can also be applied to the process of testing the trained video classification model; in the process of testing the trained video classification model, the test video refers to a query video or a query video sample, and the training video refers to a support video or a support video sample.
[0107] Step S404: According to the fusion parameters, the test video and each training video are fused to obtain at least one virtual video corresponding to the test video.
[0108] The fusion parameter refers to the fusion ratio during the process of fusing the test video with each training video.
[0109] In this step, the test video and each training video are fused according to the fusion ratio to obtain at least one virtual video corresponding to the test video. The number of virtual videos obtained is equal to the number of training videos.
[0110] Specifically, according to the fusion ratio corresponding to each image frame, the image frames of the test video and each training video are fused. During the fusion process of the test video and training video, the fusion can be performed as follows:
[0111] in, The virtual video sample obtained by fusion is represented by x, where k represents the k-th image frame; i x represents the test video sample. j This represents the training video samples.
[0112] It should be noted that, among them, λ k It can be calculated based on preset constants, the temporal parameters of the k-th frame of the test video samples and the training video samples. For example, the ratio of the temporal parameters of the k-th frame of the test video samples to the temporal parameters of the k-th frame of the training video samples can be calculated, and then the product of the preset constant and this ratio can be used as λ. k The above explanation of the calculation of the mixing ratio is merely illustrative; furthermore, λ k Other relevant parameters can also be used for calculation, but this embodiment does not limit the calculation.
[0113] Besides fusing the test video with each training video according to the fusion parameters to obtain at least one virtual video corresponding to the test video, other methods can be used to obtain virtual videos. For example, a portion of image frames in the test video and any training video can be replaced to obtain a virtual video; for example, frames 6-10 of a 10-frame test video can be replaced with frames 6-10 of a training video to obtain a virtual video; furthermore, discontinuous image frames can be replaced; specifically, virtual videos can also be obtained by performing data augmentation on the test video and training videos in other ways. That is, step S404 can also be replaced by performing data augmentation on the test video based on each training video to obtain at least one virtual video, and forming a new implementation method with other processing steps provided in this embodiment. It should be noted that in addition to the methods described above, other methods can be used for data augmentation, such as splicing and replacement, which are not limited in this embodiment.
[0114] Step S406: Input the at least one virtual video into the video classification model to perform video classification and obtain the classification results for each virtual video.
[0115] Optionally, the video classification model is obtained by training the video classification model based on model parameters using at least one virtual video sample; the model parameters are obtained after label conversion based on the paired labels of the test video samples and each training video sample. Optionally, the video classification model includes a sampling layer, a feature extraction layer, and a classification sub-model constructed based on the model parameters. The classification sub-model can be a classifier; correspondingly, the classification weights of the classifier are the aforementioned model parameters.
[0116] The following section provides a detailed explanation of the training process for the video classification model.
[0117] In this embodiment, to train the video classification model to be trained, at least one virtual video sample is used as the training sample. Specifically, each virtual video sample is input into the video classification model to be trained, that is, the video classification model constructed with the model parameters is used for video classification. In an optional implementation method provided in this embodiment, the video classification of any virtual video sample is achieved in the following way: Image sampling is performed on any virtual video sample to obtain a number of sampled image frames, and the image frames are input into the feature extraction layer for feature extraction to obtain image frame features; The image frame features are input into the classification sub-model constructed by the model parameters to perform video classification on any virtual video sample, thereby obtaining the classification result of any virtual video sample.
[0118] Optionally, the classification sub-model performs video classification on any of the virtual video samples, including: The similarity between the image frame features and at least one training video feature under each category is calculated to obtain the similarity between the image frame features and each training video feature under each category. Based on the similarity, the probability of matching any virtual video sample with each of the categories is calculated, and the calculated category matching probability is used as the classification result.
[0119] Specifically, for any virtual video sample, the virtual video sample is input into the video classification model to be trained. The video classification model first performs image sampling on the virtual video sample based on the sampling layer to obtain the number of sampled image frames. Then, the sampled image frames are input into the feature extraction layer for feature extraction to obtain image frame features. Finally, the image frame features are input into the classification sub-model constructed based on the model parameters to perform video classification and obtain the classification result.
[0120] For example, a video classification model consists of three parts: a sampling layer called TSN (temporal segment network), a feature extraction layer called ResNet50 (residual neural network), and a classifier constructed from the aforementioned classification weights.
[0121] The classification sub-model, also known as the classifier, calculates the similarity between the image frame features of the virtual video sample and the features of each training video during the video classification process, thereby obtaining the similarity between the image frame features and the features of each training video; based on the similarity, it calculates the matching probability between the virtual video sample and each category.
[0122] In this embodiment, when there is only one training video feature under each category, the similarity between the image frame feature and the training video feature is calculated, and the matching probability between the virtual video sample and each category is calculated based on this similarity. If each category contains multiple training video features, the average similarity between the image frame and the multiple training video features is calculated, and the matching probability between the virtual video sample and each category is calculated based on the average similarity. In calculating the category matching probability, the similarity between the virtual video sample and each category can be proportionally divided to obtain the category matching probability for each category; other processing methods can also be used to calculate the category matching probability, which is not limited in this embodiment.
[0123] In specific implementation, to train the video classification model, after obtaining the classification results of each virtual video sample, the parameters of the video classification model are adjusted based on the classification results of each virtual video sample. To improve the robustness and generalization ability of the trained video classification model, in an optional implementation method provided in this embodiment, the following operations are performed during the parameter adjustment process of the video classification model based on the classification results of each virtual video sample: The training loss is calculated based on each virtual video sample and the classification result of each virtual video sample, and the parameters of the video classification model constructed based on the training loss are adjusted.
[0124] In one optional implementation of this embodiment, the process of calculating the training loss based on each virtual video sample and its classification result includes: The classification result parameters in the classification result of the first virtual video sample are corrected; the first virtual video sample is obtained by fusing test video samples and training video samples of different video categories; The training loss is calculated based on the classification results after parameter correction, the classification results of the second virtual sample, and the category labels of each virtual video sample; the second virtual video sample is obtained by fusing test video samples and support video samples of the same video category.
[0125] Optionally, the classification result parameters in the classification result of the first virtual video sample are corrected, including: Correction coefficients are calculated based on the fusion parameters for parameter adjustment; The classification prediction result is parameter-corrected based on the category label, number of categories, and correction coefficient corresponding to the first virtual video sample.
[0126] In practice, introducing virtual video samples aims to reduce the overconfidence of video classification models. By introducing uncertainty into the training process through category labels, the model is encouraged to achieve better generalization performance. The goal is to further enhance this uncertainty by mixing test and training video samples from different categories. For example, if the query video sample belongs to the ballet category, and the supporting video sample also belongs to this category, the video classification model should make a more confident judgment on the virtual video sample obtained by fusing the query and supporting video samples because they share many similar features. However, if the supporting video sample belongs to a different category, there are fewer similar features, making it more difficult for the model to distinguish the category of the virtual video sample obtained by fusing the query and supporting video samples from another category. Therefore, when the test and training video samples being fused belong to different categories, a correction coefficient is first calculated.
[0127] in, It belongs to the correction factor. It is a constant. This refers to the mixing ratio.
[0128] After calculating the correction coefficient, the predicted classification results are corrected in the following way:
[0129] Where n is the number of categories.
[0130] After calculating the classification result of the first virtual video sample after correction, the training loss is calculated based on the classification result of the first virtual video sample after parameter correction, the classification result of the second virtual video sample, and the category labels of the first and second virtual video samples. The parameters of the video classification model to be trained are then corrected based on the training loss to obtain the trained video classification model.
[0131] In the actual implementation process, the video classification model retains the same backbone network as in the meta-training stage. Therefore, this backbone network (video classification model) can make correct predictions on the fused virtual videos. Given a test video, L virtual videos can be generated using the training videos, where L is a hyperparameter in meta-learning. By analyzing the classification results of all virtual videos, a better classification prediction can be made for the test video.
[0132] Step S408: Calculate the video classification result of the test video based on the classification results of each virtual video.
[0133] In practical applications, the classification result of a single test sample is calculated in the following way:
[0134] in, This indicates the video classification result of the test video; it excludes the portion where the confidence level of the corresponding category increases due to the introduction of supporting video samples.
[0135] In practical applications, if fusion is performed randomly and without constraints, as in meta-training, classification prediction will be more sensitive to noise when the proportion of query videos is low. Therefore, a fixed fusion ratio greater than 0.5 is used, and the same number of support videos are selected for fusion for each category. Finally, the above formula can be simplified to the following form:
[0136] Where m represents the number of samples selected for fusion from each category. f() represents the features of the corresponding video.
[0137] Specifically, this formula is used to calculate the video classification results for the query video.
[0138] Furthermore, in this embodiment, fusion processing can be performed in multiple dimensions or in multiple ways. The fusion processing is not limited to the data layer, but can also be applied to the feature map. For small sample video classification, since different frames of different videos are extracted through the same backbone network, the fusion processing can also be added to the backbone network. Alternatively, in addition to fusion processing of samples, more types of data augmentation methods can be introduced. For example, fusion processing can simply merge two samples pixel by pixel to generate virtual video samples without involving more complex operations such as cutting. Alternatively, a part of one sample can be swapped with another or more complex operations can be considered.
[0139] In this embodiment, a data augmentation method for meta-learning is proposed from the perspective of samples. By expanding the samples, the performance of the video classification model is improved. It can be easily applied to any few-sample video classification method that uses meta-learning. By introducing very little computational overhead, the performance of the trained video classification model is steadily improved. In two different few-sample video classification methods and three different few-sample video classification methods, this solution achieves stable improvement under various evaluation metrics.
[0140] This specification provides one or more embodiments of a training device for a video classification model, as follows: In the above embodiments, a training method for a video classification model is provided, and correspondingly, a training device for a video classification model is also provided, which will be described below with reference to the accompanying drawings.
[0141] Reference Figure 5 This illustration shows a schematic diagram of an embodiment of a training device for a video classification model provided in this embodiment.
[0142] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.
[0143] This embodiment provides a training device for a video classification model, including: The pairing label generation module 502 is configured to sample a video set to obtain test video samples and at least one training video sample, and generate pairing labels for the test video samples and each training video sample. The label conversion module 504 is configured to perform label conversion based on the paired labels of the test video samples and each training video sample to obtain model parameters and category labels of the test video samples and the at least one training video sample; The fusion processing module 506 is configured to perform fusion processing on the test video sample and each training video sample according to the fusion parameters and the category label to obtain at least one virtual video sample; The parameter adjustment module 508 is configured to input the at least one virtual video sample into a video classification model constructed based on the model parameters for video classification, and to adjust the parameters of the video classification model based on the classification results of each virtual video sample to obtain a video classification model.
[0144] This specification provides one or more embodiments of a video classification and processing device as follows: In the above embodiments, a video classification processing method is provided, and correspondingly, a video classification processing apparatus is also provided, which will be described below with reference to the accompanying drawings.
[0145] Reference Figure 6 The diagram illustrates an embodiment of a video classification processing device provided in this embodiment.
[0146] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.
[0147] This embodiment provides a video classification processing device, including: The sampling module 602 is configured to sample the target video set to obtain a test video and at least one training video. The fusion processing module 604 is configured to perform fusion processing on the test video and each training video according to fusion parameters to obtain at least one virtual video corresponding to the test video; The video classification module 606 is configured to input the at least one virtual video into a video classification model for video classification and obtain the classification results of each virtual video. The result calculation module 608 is configured to calculate the video classification result of the test video based on the classification results of each virtual video; The video classification model is obtained by training the video classification model based on model parameters using at least one virtual video sample; the model parameters are obtained by label conversion based on the paired labels of the test video sample and each training video sample.
[0148] This specification provides one or more embodiments of a training device for a video classification model, as follows: Corresponding to the video classification model training method described above, based on the same technical concept, one or more embodiments of this specification also provide a video classification model training device, which is used to execute the video classification model training method provided above. Figure 7 This is a schematic diagram of the structure of a training device for a video classification model provided in one or more embodiments of this specification.
[0149] This embodiment provides a training device for a video classification model, comprising: like Figure 7 As shown, the training device for the video classification model can vary significantly due to differences in configuration or performance. It may include one or more processors 701 and a memory 702, where one or more application programs or data may be stored. The memory 702 can be temporary or persistent storage. The application programs stored in the memory 702 may include one or more modules (not shown), each module comprising a series of computer-executable instructions for the video classification model training device. Furthermore, the processor 701 may be configured to communicate with the memory 702, executing the series of computer-executable instructions in the memory 702 on the video classification model training device. The video classification model training device may also include one or more power supplies 703, one or more wired or wireless network interfaces 704, one or more input / output interfaces 705, one or more keyboards 706, etc.
[0150] In one specific embodiment, the training device for the video classification model includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the training device for the video classification model, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: The video set is sampled to obtain test video samples and at least one training video sample, and paired labels are generated for the test video samples and each training video sample. Label conversion is performed based on the paired labels of the test video samples and each training video sample to obtain model parameters and category labels of the test video samples and at least one training video sample; According to the fusion parameters and the category labels, the test video sample and each training video sample are fused to obtain at least one virtual video sample; The at least one virtual video sample is input into a video classification model constructed based on the model parameters for video classification, and the parameters of the video classification model are adjusted based on the classification results of each virtual video sample to obtain a video classification model.
[0151] This specification provides one or more embodiments of a video classification and processing device as follows: Corresponding to the video classification processing method described above, based on the same technical concept, one or more embodiments of this specification also provide a video classification processing device for performing the video classification processing method provided above. Figure 8 This is a schematic diagram of the structure of a video classification processing device provided for one or more embodiments of this specification.
[0152] This embodiment provides a video classification processing device, including: like Figure 8 As shown, video classification processing devices can vary significantly due to differences in configuration or performance. They may include one or more processors 801 and memory 802, with memory 802 storing one or more application programs or data. Memory 802 can be temporary or persistent storage. The application programs stored in memory 802 may include one or more modules (not shown), each module including a series of computer-executable instructions from the video classification processing device. Furthermore, processor 801 may be configured to communicate with memory 802, executing the series of computer-executable instructions in memory 802 on the video classification processing device. The video classification processing device may also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input / output interfaces 805, one or more keyboards 806, etc.
[0153] In one specific embodiment, the video classification processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the video classification processing device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Sample the target video set to obtain a test video and at least one training video; According to the fusion parameters, the test video and each training video are fused together to obtain at least one virtual video corresponding to the test video; The at least one virtual video is input into a video classification model for video classification to obtain the classification results for each virtual video; The video classification result of the test video is calculated based on the classification results of each virtual video; The video classification model is obtained by training the video classification model based on model parameters using at least one virtual video sample; the model parameters are obtained by label conversion based on the paired labels of the test video sample and each training video sample.
[0154] This specification provides one or more embodiments of a storage medium as follows: Corresponding to the training method of the video classification model described above, based on the same technical concept, one or more embodiments of this specification also provide a storage medium.
[0155] The storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed by a processor, implement the following process: The video set is sampled to obtain test video samples and at least one training video sample, and paired labels are generated for the test video samples and each training video sample. Label conversion is performed based on the paired labels of the test video samples and each training video sample to obtain model parameters and category labels of the test video samples and at least one training video sample; According to the fusion parameters and the category labels, the test video sample and each training video sample are fused to obtain at least one virtual video sample; The at least one virtual video sample is input into a video classification model constructed based on the model parameters for video classification, and the parameters of the video classification model are adjusted based on the classification results of each virtual video sample to obtain a video classification model.
[0156] It should be noted that the embodiments concerning storage media in this specification and the embodiments concerning training methods for video classification models in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding methods described above, and the repeated parts will not be described again.
[0157] One or more embodiments of another storage medium provided in this specification are as follows: Corresponding to the video classification processing method described above, based on the same technical concept, one or more embodiments of this specification also provide a storage medium.
[0158] The storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed by a processor, implement the following process: Sample the target video set to obtain a test video and at least one training video; According to the fusion parameters, the test video and each training video are fused together to obtain at least one virtual video corresponding to the test video; The at least one virtual video is input into a video classification model for video classification to obtain the classification results for each virtual video; The video classification result of the test video is calculated based on the classification results of each virtual video; The video classification model is obtained by training the video classification model based on model parameters using at least one virtual video sample; the model parameters are obtained by label conversion based on the paired labels of the test video sample and each training video sample.
[0159] It should be noted that the embodiments concerning storage media in this specification and the embodiments concerning video classification processing methods in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0160] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, please refer to each other. Each embodiment focuses on describing the differences from other embodiments. For example, the device embodiment, equipment embodiment, and storage medium embodiment are all similar to the method embodiment, so the description is relatively simple. For reading the relevant content of the device embodiment, equipment embodiment, and storage medium embodiment, please refer to the description of the method embodiment.
[0161] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0162] In the 1930s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many improvements to the methodology today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that an improvement to the methodology cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0163] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0164] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0165] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0166] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0167] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0168] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0169] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0170] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0171] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0172] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0173] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0174] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0175] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0176] The above description is merely an embodiment of this document and is not intended to limit the scope of this document. Various modifications and variations can be made to this document by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this document should be included within the scope of the claims of this document.
Claims
1. A training method for a video classification model, comprising: The video set is sampled to obtain test video samples and at least one training video sample, and paired labels are generated for the test video samples and each training video sample. Based on the paired labels, a first training video sample that is a positive sample to the test video sample and a second training video sample that is a negative sample to the test video sample are determined. Positive sample similarity is calculated based on the test video sample and the first training video sample, and negative sample similarity is calculated based on the test video sample and the second training video sample. Each training video sample and the test video sample are input into a classifier for category classification to obtain category labels. The classifier is constructed with positive sample similarity and negative sample similarity as constraints. According to the fusion parameters and the category labels, the test video sample and each training video sample are fused to obtain at least one virtual video sample; The at least one virtual video sample is input into the video classification model for video classification, and the parameters of the video classification model are adjusted based on the classification results of each virtual video sample to obtain the video classification model.
2. The training method for the video classification model according to claim 1, wherein fusing the test video sample with each training video sample according to the fusion parameters and the category label to obtain at least one virtual video sample includes: According to the fusion ratio, the test video sample and each image frame of each training video sample are fused to obtain at least one virtual video sample. According to the fusion ratio, the category labels of each image frame of the test video sample and each training video sample are fused to obtain the virtual category label of each virtual video sample.
3. The training method for the video classification model according to claim 1, wherein sampling the video set to obtain test video samples and at least one training video sample, and generating paired labels for the test video samples and each training video sample, comprises: Based on the number of categories, the video set is sampled to obtain a test video sample set consisting of multiple test video samples and a training video sample set consisting of the at least one training video sample. Generate paired labels between each test video sample and each training video sample in the test video sample set.
4. The training method for the video classification model according to claim 3, wherein the test video sample set contains test video samples of video categories of a certain number of categories, and the number of test video samples under each video category is a preset threshold; The training video sample set contains training video samples of the preset number of categories, and the number of training video samples under each video category is greater than the preset threshold.
5. The training method for the video classification model according to claim 1, wherein generating the paired labels of the test video samples and each training video sample comprises: The paired labels of the test video sample and the training video sample belonging to the same video category as the test video sample are determined as the first label; The paired labels of the test video sample and the training video sample belonging to a different video category from the test video sample are determined as the second label.
6. The training method for the video classification model according to claim 1, wherein video classification of any virtual video sample in the at least one virtual video sample includes: Image sampling is performed on any virtual video sample to obtain a number of sampled image frames, and the image frames are input into the feature extraction layer for feature extraction to obtain image frame features; The image frame features are input into the classification sub-model to perform video classification on any virtual video sample, thereby obtaining the classification result of any virtual video sample.
7. The training method for the video classification model according to claim 6, wherein the classification sub-model performs video classification on any virtual video sample, comprising: The similarity between the image frame features and at least one training video feature under each category is calculated to obtain the similarity between the image frame features and each training video feature under each category. Based on the similarity, the probability of matching any virtual video sample with each of the categories is calculated, and the calculated category matching probability is used as the classification result.
8. The training method for the video classification model according to claim 1, wherein adjusting the parameters of the video classification model based on the classification results of each virtual video sample includes: The training loss is calculated based on each virtual video sample and its classification result, and the parameters of the video classification model are adjusted based on the training loss.
9. The training method for the video classification model according to claim 8, wherein calculating the training loss based on each virtual video sample and the classification result of each virtual video sample includes: The classification result parameters in the classification result of the first virtual video sample are corrected. The first virtual video sample is obtained by fusing test video samples and training video samples from different video categories; Based on the classification results after parameter correction, the classification results of the second virtual video sample, and the category labels of each virtual video sample, the training loss is calculated; The second virtual video sample is obtained by fusing test video samples and supporting video samples of the same video category.
10. The training method for the video classification model according to claim 9, wherein the parameter correction of the classification result parameters in the classification result of the first virtual video sample includes: Correction coefficients are calculated based on the fusion parameters for parameter adjustment; The classification result is parameter-corrected based on the category label, number of categories, and correction coefficient corresponding to the first virtual video sample.
11. The training method for the video classification model according to claim 1, further comprising: Sample the target video set to obtain a test video and at least one training video; According to the fusion parameters, the test video and each training video are fused to obtain at least one virtual video corresponding to the test video; The at least one virtual video is input into the trained video classification model to classify the video and obtain the classification results for each virtual video. The video classification result of the test video is calculated based on the classification results of each virtual video.
12. A video classification processing method, comprising: Sample the target video set to obtain a test video and at least one training video; According to the fusion parameters, the test video and each training video are fused together to obtain at least one virtual video corresponding to the test video; The at least one virtual video is input into a video classification model for video classification to obtain the classification results for each virtual video; The video classification result of the test video is calculated based on the classification results of each virtual video; The video classification model is obtained in the following way: Based on the paired labels of the test video samples and each training video sample, a first training video sample that is a positive sample to the test video sample and a second training video sample that is a negative sample to the test video sample are determined. The positive sample similarity is calculated based on the test video sample and the first training video sample, and the negative sample similarity is calculated based on the test video sample and the second training video sample. The training video samples and the test video samples are input into a classifier for category classification to obtain category labels. The classifier is constructed with positive sample similarity and negative sample similarity as constraints. According to the fusion parameters and the category labels, the test video samples and each training video sample are fused to obtain at least one virtual video sample, and the at least one virtual video sample is input into the video classification model for training.
13. The video classification processing method according to claim 12, wherein the video classification of any virtual video in the at least one virtual video includes: Image sampling is performed on any of the virtual videos to obtain a number of sampled image frames, and the image frames are input into the feature extraction layer for feature extraction to obtain image frame features; The image frame features are input into the classification sub-model to perform video classification on any virtual video, thereby obtaining the classification result of any virtual video.
14. The video classification processing method according to claim 12, wherein fusing the test video with each training video according to fusion parameters to obtain at least one virtual video corresponding to the test video includes: According to the fusion ratio, each image frame of the test video and each training video is fused to obtain the at least one virtual video.
15. A training device for a video classification model, comprising: The matching tag generation module is configured to sample a video set to obtain test video samples and at least one training video sample, and generate matching tags for the test video samples and each training video sample. The label conversion module is configured to determine a first training video sample that is a positive sample to the test video sample and a second training video sample that is a negative sample to the test video sample based on the paired labels; calculate the positive sample similarity based on the test video sample and the first training video sample; calculate the negative sample similarity based on the test video sample and the second training video sample; and input each training video sample and the test video sample into a classifier for category classification to obtain category labels. The classifier is constructed with positive sample similarity and negative sample similarity as constraints. The fusion processing module is configured to perform fusion processing on the test video sample and each training video sample according to the fusion parameters and the category label to obtain at least one virtual video sample; The parameter adjustment module is configured to input the at least one virtual video sample into the video classification model for video classification, and to adjust the parameters of the video classification model based on the classification results of each virtual video sample to obtain a video classification model.
16. A video classification processing apparatus, comprising: The sampling module is configured to sample the target video set to obtain a test video and at least one training video; The fusion processing module is configured to perform fusion processing on the test video and each training video according to fusion parameters to obtain at least one virtual video corresponding to the test video; The video classification module is configured to input the at least one virtual video into a video classification model for video classification, and obtain the classification results for each virtual video; The result calculation module is configured to calculate the video classification result of the test video based on the classification results of each virtual video; The video classification model is obtained in the following way: Based on the paired labels of the test video samples and each training video sample, a first training video sample that is a positive sample to the test video sample and a second training video sample that is a negative sample to the test video sample are determined. The positive sample similarity is calculated based on the test video sample and the first training video sample, and the negative sample similarity is calculated based on the test video sample and the second training video sample. The training video samples and the test video samples are input into a classifier for category classification to obtain category labels. The classifier is constructed with positive sample similarity and negative sample similarity as constraints. According to the fusion parameters and the category labels, the test video samples and each training video sample are fused to obtain at least one virtual video sample, and the at least one virtual video sample is input into the video classification model for training.
17. A training device for a video classification model, comprising: processor; And, a memory configured to store computer-executable instructions, which, when executed, cause the processor to: The video set is sampled to obtain test video samples and at least one training video sample, and paired labels are generated for the test video samples and each training video sample. Based on the paired labels, a first training video sample that is a positive sample to the test video sample and a second training video sample that is a negative sample to the test video sample are determined. Positive sample similarity is calculated based on the test video sample and the first training video sample, and negative sample similarity is calculated based on the test video sample and the second training video sample. Each training video sample and the test video sample are input into a classifier for category classification to obtain category labels. The classifier is constructed with positive sample similarity and negative sample similarity as constraints. According to the fusion parameters and the category labels, the test video sample and each training video sample are fused to obtain at least one virtual video sample; The at least one virtual video sample is input into the video classification model for video classification, and the parameters of the video classification model are adjusted based on the classification results of each virtual video sample to obtain the video classification model.
18. A video classification processing device, comprising: processor; And, a memory configured to store computer-executable instructions, which, when executed, cause the processor to: Sample the target video set to obtain a test video and at least one training video; According to the fusion parameters, the test video and each training video are fused together to obtain at least one virtual video corresponding to the test video; The fusion parameters include the fusion ratio between the test video and each training video; The at least one virtual video is input into a video classification model for video classification to obtain the classification results for each virtual video; The video classification result of the test video is calculated based on the classification results of each virtual video; The video classification model is obtained in the following way: Based on the paired labels of the test video samples and each training video sample, a first training video sample that is a positive sample to the test video sample and a second training video sample that is a negative sample to the test video sample are determined. The positive sample similarity is calculated based on the test video sample and the first training video sample, and the negative sample similarity is calculated based on the test video sample and the second training video sample. The training video samples and the test video samples are input into a classifier for category classification to obtain category labels. The classifier is constructed with positive sample similarity and negative sample similarity as constraints. According to the fusion parameters and the category labels, the test video samples and each training video sample are fused to obtain at least one virtual video sample, and the at least one virtual video sample is input into the video classification model for training.
19. A storage medium for storing computer-executable instructions, which, when executed by a processor, perform the following process: The video set is sampled to obtain test video samples and at least one training video sample, and paired labels are generated for the test video samples and each training video sample. Based on the paired labels, a first training video sample that is a positive sample to the test video sample and a second training video sample that is a negative sample to the test video sample are determined. Positive sample similarity is calculated based on the test video sample and the first training video sample, and negative sample similarity is calculated based on the test video sample and the second training video sample. Each training video sample and the test video sample are input into a classifier for category classification to obtain category labels. The classifier is constructed with positive sample similarity and negative sample similarity as constraints. According to the fusion parameters and the category labels, the test video sample and each training video sample are fused to obtain at least one virtual video sample; The at least one virtual video sample is input into the video classification model for video classification, and the parameters of the video classification model are adjusted based on the classification results of each virtual video sample to obtain the video classification model.
20. A storage medium for storing computer-executable instructions, which, when executed by a processor, perform the following process: Sample the target video set to obtain a test video and at least one training video; According to the fusion parameters, the test video and each training video are fused together to obtain at least one virtual video corresponding to the test video; The at least one virtual video is input into a video classification model for video classification to obtain the classification results for each virtual video; The video classification result of the test video is calculated based on the classification results of each virtual video; The video classification model is obtained in the following way: Based on the paired labels of the test video samples and each training video sample, a first training video sample that is a positive sample to the test video sample and a second training video sample that is a negative sample to the test video sample are determined. The positive sample similarity is calculated based on the test video sample and the first training video sample, and the negative sample similarity is calculated based on the test video sample and the second training video sample. The training video samples and the test video samples are input into a classifier for category classification to obtain category labels. The classifier is constructed with positive sample similarity and negative sample similarity as constraints. According to the fusion parameters and the category labels, the test video samples and each training video sample are fused to obtain at least one virtual video sample, and the at least one virtual video sample is input into the video classification model for training.