Video Classification Method and Device, Electronic Device, and Computer Readable Storage Medium
By integrating and weighting video features with text features, the problem of low classification accuracy in massive videos is solved, and more efficient video classification is achieved, suitable for a variety of complex data distribution scenarios.
Patent Information
- Application Number
- CN202210749546.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-06-29
AI Technical Summary
With the explosive growth of the number of videos, it is difficult for users to find the videos of interest from massive videos. When the existing technology classifies video content based on video content, there is a problem of low classification accuracy.
By obtaining the characteristics of the pending video and text, feature extraction and fusion are performed, and the similarity between video features and text features is used for weighted fusion to improve the accuracy of video classification.
The accuracy of video classification is improved, especially in the case of closed sets, long-tail distribution, and few samples, and the recognition ability of video categories is enhanced.
Smart Images

Figure CN115063726B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular, to a video classification method and apparatus, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the explosive growth of the number of various videos, it has become increasingly difficult for users to find interesting videos from a vast amount of videos. Classifying videos based on video content helps users quickly find interesting videos according to the video classification results. Therefore, how to classify videos is of great significance. Summary of the Invention
[0003] This application provides a video classification method and apparatus, an electronic device, and a computer-readable storage medium.
[0004] In a first aspect, a video classification method is provided, and the method includes:
[0005] Obtain a video to be processed and at least one first text feature; the at least one first text feature carries semantic information for describing at least one first category;
[0006] Perform feature extraction processing on the video to be processed to obtain a first video feature;
[0007] Fuse the first video feature and the at least one first text feature to obtain a first fusion feature;
[0008] Classify the video to be processed according to the first fusion feature to obtain a second category of the video to be processed.
[0009] In this aspect, when the video classification apparatus obtains a first video feature by performing feature extraction processing on the video to be processed, a first fusion feature is obtained by fusing the first video feature and at least one first text feature, enriching the information for determining the category of the video to be processed. Then, the video classification apparatus classifies the video to be processed according to the first fusion feature to obtain a second category of the video to be processed, which can improve the accuracy of the second category.
[0010] Combined with any implementation manner of this application, when the number of the first text features is greater than 1, the at least one first text feature includes a second text feature and a third text feature;
[0011] After performing feature extraction processing on the video to be processed to obtain a first video feature, and before fusing the first video feature and the at least one first text feature to obtain a first fusion feature, the method further includes:
[0012] Obtain a first similarity between the second text feature and the first video feature, and a second similarity between the third text feature and the first video feature;
[0013] According to the first similarity and the second similarity, obtain a first weight of the second text feature and a second weight of the third text feature; when the first similarity is greater than the second similarity, the first weight is greater than the second weight; when the first similarity is equal to the second similarity, the first weight is equal to the second weight;
[0014] The fusing of the first video feature and the at least one first text feature to obtain a first fused feature includes:
[0015] According to the first weight and the second weight, perform weighted fusion on the second text feature and the third text feature to obtain a second fused feature;
[0016] Fuse the second fused feature and the first video feature to obtain the first fused feature.
[0017] In this embodiment, since both the second text feature and the third text feature carry semantic information describing the video category (i.e., the first category), and the semantic information describing the first category is beneficial to determining the category of the video to be processed. Therefore, using the information carried by the second text feature and the information carried by the third text feature to determine the second category of the video to be processed can improve the accuracy of the second category of the video to be processed.
[0018] Also, because the information carried by the second text feature is different from the information carried by the third text feature, the degree of improvement of the accuracy of the second category by the information carried by the second text feature is different from the degree of improvement of the accuracy of the second category by the information carried by the third text feature.
[0019] Specifically, the degree of correlation between the information carried by the second text feature and the information carried by the first video feature is called the first correlation degree, and the degree of correlation between the information carried by the third text feature and the information carried by the first video feature is called the second correlation degree. When the first similarity is greater than the second similarity, the first correlation degree is greater than the second correlation degree. That is to say, the information carried by the second text feature is closer to the information carried by the first video feature than the information carried by the third text feature.
[0020] At this time, the information carried by the second text feature can improve the accuracy of the second category to a greater extent compared to the information carried by the third text feature. Therefore, when the first similarity is greater than the second similarity, making more use of the information carried by the second text feature is more conducive to improving the accuracy of the classification result of the video to be processed, that is, more conducive to improving the accuracy of the second category.
[0021] Therefore, when the first similarity is greater than the second similarity, the video classification device makes the first weight greater than the second weight, uses the first weight as the weight of the second text feature, uses the second weight as the weight of the third text feature, and performs weighted fusion on the second text feature and the third text feature to obtain a second fusion feature. In this way, in the process of using the second fusion feature to classify the video to be processed, not only can the semantic information carried by the second text feature and the semantic information carried by the third text feature be used to determine the second category of the video to be processed, but also the semantic information carried by the second text feature and the semantic information carried by the third text feature can be better used to determine the second category of the video to be processed, thereby improving the accuracy of the second category.
[0022] Thus, by fusing the second fusion feature and the first video feature, the video classification device obtains a first fusion feature, which can make the information carried by the first fusion feature more conducive to improving the accuracy of the second category.
[0023] Combined with any implementation manner of the present application, classifying the video to be processed according to the first fusion feature to obtain the second category of the video to be processed includes:
[0024] Predicting a third category of the video to be processed and a first confidence level of the third category according to the first fusion feature; the third category belongs to the at least one first category;
[0025] When the first confidence level is less than or equal to a confidence level threshold, determining that the second category of the video to be processed is a category other than the at least one first category;
[0026] When the first confidence level is greater than the confidence level threshold, determining that the third category is the second category.
[0027] In this implementation manner, the video classification device predicts the third category of the video to be processed and the first confidence level of the third category according to the first fusion feature. When the first confidence level is less than or equal to the confidence level threshold, it determines that the second category of the video to be processed is a category other than the at least one first category. When the first confidence level is greater than the confidence level threshold, it determines that the third category is the second category, which can improve the accuracy of the second category.
[0028] Combined with any implementation manner of the present application, the obtaining of at least one first text feature includes:
[0029] Obtain at least one fourth text feature and a second confidence level of the at least one fourth text feature; the second confidence level characterizes the accuracy of the semantic information carried by the fourth text feature in describing the first category corresponding to the fourth text feature;
[0030] From the at least one fourth text feature, determine n fourth text features with the highest second confidence level corresponding to each of the first categories to obtain the at least one first text feature.
[0031] In this implementation manner, the video classification device screens at least one first text feature from at least one fourth text feature based on the second confidence level, which can improve the accuracy of the description of at least one first text feature for at least one first category.
[0032] Combined with any implementation manner of the present application, the performing of feature extraction processing on the video to be processed to obtain a first video feature includes:
[0033] Perform feature extraction processing on at least one frame of the image to be processed in the video to be processed to obtain the frame feature of the at least one frame of the image to be processed;
[0034] Fuse the frame features of the at least one frame of the image to be processed and the timestamp information of the at least one frame of the image to be processed to obtain the first video feature.
[0035] In this implementation manner, the video classification device can make the frame features carry timestamp information by fusing the frame features of at least one frame of the image to be processed and the timestamp information of at least one frame of the image to be processed. In this way, in the first video feature obtained by fusion, the frame features carry timestamp information, and according to this timestamp information, the sequence relationship of different frame features in the time dimension can be determined.
[0036] Combined with any implementation manner of the present application, the video classification method is implemented through a video classification network, and the video classification network includes a video encoding module;
[0037] The performing of feature extraction processing on the video to be processed to obtain a first video feature includes:
[0038] Perform feature extraction processing on the video to be processed through the video encoding module to obtain a first video feature;
[0039] The video classification method further includes the training process of the video classification network:
[0040] Obtain a first training video;
[0041] The first training video is subjected to feature extraction processing by the video encoding module to obtain second video features;
[0042] The second video features and the at least one first text feature are fused to obtain third fused features;
[0043] Based on the third fused features, a fourth category of the first training video is obtained;
[0044] Based on the first difference between the fourth category and the label of the first training video, a first loss of the video classification network is obtained;
[0045] Based on the first loss, the parameters of the video classification network are updated to obtain the video classification network.
[0046] Through this implementation manner, the video classification device completes the training of the video classification network, so that the video classification method described above can be implemented through the video classification network.
[0047] In addition, there are currently many methods for classifying videos through deep learning techniques. However, these methods have low classification accuracy when there is at least one of a closed set, long-tail distribution, and few-shot in the training data. In the training method of the video classification network provided in the embodiments of the present application, since the information carried by the second video features and the information carried by at least one text feature are both used for classification, the classification accuracy can be improved when there is at least one of a closed set, long-tail distribution, and few-shot.
[0048] Combined with any implementation manner of the present application, the updating the parameters of the video classification network based on the first loss includes:
[0049] Based on the first loss, the parameters of the video classification network other than the parameters of the video encoding module are updated;
[0050] The video encoding module is obtained by training a video classification training network, and the video classification training network includes the video encoding module;
[0051] The video classification method further includes a training process of the video classification training network:
[0052] A second training video and at least two first training texts are obtained; the labels of the at least two first training texts include at least two fifth categories, and the labels of the at least two first training texts include the at least one first category; the at least two first training texts include a second training text, and the fifth category of the second training text is the same as the sixth category of the second training video;
[0053] The second training video is subjected to feature extraction processing by the video encoding module to obtain third video features;
[0054] According to the third similarity between the third video features and the second training text, a second loss of the video classification training network is obtained; the second loss is negatively correlated with the third similarity;
[0055] According to the second loss, the parameters of the video classification training network are updated to obtain the video classification training network.
[0056] In this embodiment, the fifth category of the second training text is the same as the sixth category of the second training video. The video classification device determines the second loss according to the third similarity between the third video features of the second training video and the second training text, where the second loss and the third similarity are negatively correlated. In this way, updating the parameters of the video classification training network according to the second loss can make the third video features extracted by the video encoding module of the video classification training network have a high matching degree with the second training text. In other words, it makes the third video features extracted by the video encoding module have a high matching degree with the text of the sixth category.
[0057] Combined with any embodiment of the present application, the at least two first training texts further include a third training text, and the fifth category described by the third training text is different from the sixth category;
[0058] Before obtaining the second loss of the video classification training network according to the third similarity between the third video features and the second training text, the method further includes:
[0059] Determine a fourth similarity between the third video features and the third training text;
[0060] The obtaining of the second loss of the video classification training network according to the third similarity between the third video features and the second training text includes:
[0061] According to the third similarity and the fourth similarity, a second loss of the video classification training network is obtained; the second loss is positively correlated with the fourth similarity.
[0062] In this embodiment, the fifth category of the third training text is different from the sixth category of the second training video. The video classification device determines the second loss according to the third similarity and the fourth similarity in the case of determining the fourth similarity between the third video features of the second training video and the third training text, where the third similarity is negatively correlated with the second loss, and the fourth similarity is positively correlated with the second loss.
[0063] In this way, updating the parameters of the video classification training network according to the second loss can not only make the third video features extracted by the video encoding module of the video classification training network have a high degree of matching with the second training text, but also make the third video features extracted by the video encoding module of the video classification training network have a low degree of matching with the third training text. In other words, it makes the third video features extracted by the video encoding module have a high degree of matching with the text of the sixth category, and makes the extracted third video features have a low degree of matching with the text other than the sixth category.
[0064] Combined with any implementation manner of the present application, the video classification training network further includes a text encoding module;
[0065] Before obtaining the second loss of the video classification training network according to the third similarity between the third video feature and the second training text, the method further includes:
[0066] Performing feature extraction processing on the second training text through the text encoding module to obtain fifth text features of the second training text;
[0067] Determining the similarity between the fifth text feature and the third video feature as the third similarity between the third video feature and the second training text.
[0068] In this implementation manner, when the video classification device extracts the fifth text features of the second training text through the text encoding module, it determines the similarity between the fifth text feature and the third video feature as the third similarity.
[0069] Combined with any implementation manner of the present application, when obtaining at least one first text feature, including: obtaining at least one fourth text feature and a second confidence level of the at least one fourth text feature; the second confidence level represents the accuracy of the semantic information carried by the fourth text feature in describing the first category corresponding to the fourth text feature; and determining, from the at least one fourth text feature, n fourth text features with the highest second confidence level corresponding to each first category to obtain the at least one first text feature, obtaining the at least one fourth text feature and the second confidence level of the at least one fourth text feature includes:
[0070] Performing feature extraction processing on the at least one first training text through the text encoding module to obtain text features of the at least one training text as the at least one fourth text feature; the at least one fourth text feature includes the fifth text feature;
[0071] The second confidence level of the at least one fourth text feature includes the second confidence level of the fifth text feature. The obtaining of the second confidence level of the at least one fourth text feature includes:
[0072] Obtaining the second confidence level of the fifth text feature according to the third similarity; the second confidence level of the fifth text feature is positively correlated with the third similarity.
[0073] In this embodiment, the video classification device performs feature extraction processing on at least one first training text through a text encoding module to obtain at least one fourth text feature. When the third similarity is positively correlated with the second confidence level of the fifth text feature, obtaining the second confidence level of the fifth text feature according to the third similarity can improve the accuracy of the second confidence level of the fifth text feature.
[0074] Combined with any embodiment of the present application, before updating the parameters of the video classification training network according to the second loss, the method further includes:
[0075] Obtaining a fourth video feature of the second training video; the fourth video feature is extracted from the second training video through a trained video feature extraction model;
[0076] Obtaining a third loss according to the first difference between the third video feature and the fourth video feature; the third loss is positively correlated with the first difference;
[0077] The updating of the parameters of the video classification training network according to the second loss includes:
[0078] Updating the parameters of the video classification training network according to the second loss and the third loss.
[0079] In this embodiment, the fourth video feature is extracted from the second training video through a trained video feature extraction model. Therefore, the fourth video feature can be used as the ground truth (GT) of the video feature of the second training video. The video classification device obtains a third loss according to the first difference between the third video feature and the fourth video feature. In this way, during the training of the video classification network according to the second loss and the third loss, the accuracy of the third video feature extracted by the video encoding module from the second training video can be improved.
[0080] Combined with any embodiment of the present application, before updating the parameters of the video classification training network according to the second loss and the third loss, the method further includes:
[0081] Obtain the seventh text feature of the second training text; the seventh text feature is extracted from the second training text by a trained text feature extraction model;
[0082] Obtain a fourth loss according to the second difference between the fifth text feature and the seventh text feature; the fourth loss is positively correlated with the second difference;
[0083] The updating the parameters of the video classification training network according to the second loss and the third loss includes:
[0084] Update the parameters of the video classification training network according to the second loss, the third loss, and the fourth loss.
[0085] In this embodiment, the seventh text feature is extracted from the second training text by a trained text feature extraction model. Therefore, the seventh text feature can be used as the GT of the text feature of the second training text. The video classification device obtains a fourth loss according to the second difference between the fifth text feature and the seventh text feature. In this way, during the training of the video classification network according to the second loss, the third loss, and the fourth loss, the accuracy of the fifth text feature extracted by the text encoding module from the second training text can be improved.
[0086] Combined with any embodiment of the present application, the obtaining the second training video and at least two first training texts includes:
[0087] Obtain a training data set, the training data set is constructed according to at least two of the following training tasks: closed set, long-tail distribution, few-shot, open set;
[0088] Sample the second training video and the at least two first training texts from the training data set.
[0089] In this embodiment, the training data set is constructed according to at least two of the closed set, long-tail distribution, few-shot, and open set. Therefore, the training data set can be used to make the model perform at least two different training tasks simultaneously, which can improve the training efficiency.
[0090] In a second aspect, a video classification device is provided, the device includes:
[0091] An obtaining unit, configured to obtain a video to be processed and at least one first text feature; the at least one first text feature carries semantic information for describing at least one first category;
[0092] A first processing unit, configured to perform feature extraction processing on the video to be processed to obtain a first video feature;
[0093] A second processing unit, configured to fuse the first video feature and the at least one first text feature to obtain a first fused feature;
[0094] A third processing unit, configured to classify the video to be processed according to the first fused feature to obtain a second category of the video to be processed.
[0095] Combined with any embodiment of the present application, when the number of the first text features is greater than 1, the at least one first text feature includes a second text feature and a third text feature;
[0096] The obtaining unit is further configured to obtain a first similarity between the second text feature and the first video feature, and a second similarity between the third text feature and the first video feature;
[0097] The second processing unit is configured to:
[0098] Obtain a first weight of the second text feature and a second weight of the third text feature according to the first similarity and the second similarity; when the first similarity is greater than the second similarity, the first weight is greater than the second weight; when the first similarity is equal to the second similarity, the first weight is equal to the second weight;
[0099] Perform weighted fusion on the second text feature and the third text feature according to the first weight and the second weight to obtain a second fused feature;
[0100] Fuse the second fused feature and the first video feature to obtain the first fused feature.
[0101] Combined with any embodiment of the present application, the third processing unit is configured to:
[0102] Predict a third category of the video to be processed and a first confidence level of the third category according to the first fused feature; the third category belongs to the at least one first category;
[0103] When the first confidence level is less than or equal to a confidence level threshold, determine that the second category of the video to be processed is a category other than the at least one first category;
[0104] When the first confidence level is greater than the confidence level threshold, determine that the third category is the second category.
[0105] Combined with any embodiment of the present application, the obtaining unit is configured to:
[0106] Obtain at least one fourth text feature and a second confidence level of the at least one fourth text feature; the second confidence level characterizes the accuracy of the semantic information carried by the fourth text feature in describing the first category corresponding to the fourth text feature;
[0107] From the at least one fourth text feature, determine n fourth text features with the highest second confidence level corresponding to each of the first categories to obtain the at least one first text feature.
[0108] Combined with any embodiment of the present application, the first processing unit is configured to:
[0109] Perform feature extraction processing on at least one frame of to-be-processed image in the to-be-processed video to obtain the frame features of the at least one frame of to-be-processed image;
[0110] Fuse the frame features of the at least one frame of to-be-processed image and the timestamp information of the at least one frame of to-be-processed image to obtain the first video feature.
[0111] Combined with any embodiment of the present application, the video classification method is implemented through a video classification network, and the video classification network includes a video encoding module;
[0112] The first processing unit is configured to perform feature extraction processing on the to-be-processed video through the video encoding module to obtain a first video feature;
[0113] The video classification device further includes a training unit, and the training unit is configured to execute the training process of the video classification network:
[0114] Obtain a first training video;
[0115] Perform feature extraction processing on the first training video through the video encoding module to obtain a second video feature;
[0116] Fuse the second video feature and the at least one first text feature to obtain a third fusion feature;
[0117] Obtain a fourth category of the first training video according to the third fusion feature;
[0118] Obtain a first loss of the video classification network according to a first difference between the fourth category and a label of the first training video;
[0119] Update parameters of the video classification network according to the first loss to obtain the video classification network.
[0120] Combined with any embodiment of the present application, the training unit is configured to:
[0121] Update the parameters in the video classification network except for the parameters of the video encoding module according to the first loss;
[0122] The video encoding module is obtained by training a video classification training network, and the video classification training network includes the video encoding module;
[0123] The video classification method further includes the training process of the video classification training network:
[0124] Obtain a second training video and at least two first training texts; the labels of the at least two first training texts include at least two fifth categories, and the labels of the at least two first training texts include the at least one first category; the at least two first training texts include the second training text, and the fifth category of the second training text is the same as the sixth category of the second training video;
[0125] Perform feature extraction processing on the second training video through the video encoding module to obtain a third video feature;
[0126] Obtain the second loss of the video classification training network according to the third similarity between the third video feature and the second training text; the second loss is negatively correlated with the third similarity;
[0127] Update the parameters of the video classification training network according to the second loss to obtain the video classification training network.
[0128] Combined with any implementation manner of the present application, the at least two first training texts further include a third training text, and the fifth category described by the third training text is different from the sixth category;
[0129] The training unit is further configured to:
[0130] Determine the fourth similarity between the third video feature and the third training text;
[0131] Obtain the second loss of the video classification training network according to the third similarity and the fourth similarity; the second loss is positively correlated with the fourth similarity.
[0132] Combined with any implementation manner of the present application, the video classification training network further includes a text encoding module;
[0133] The training unit is further configured to:
[0134] Perform feature extraction processing on the second training text through the text encoding module to obtain the fifth text feature of the second training text;
[0135] Determine the similarity between the fifth text feature and the third video feature as the third similarity between the third video feature and the second training text.
[0136] Combined with any embodiment of the present application, in the obtaining unit, it is used to: obtain at least one fourth text feature and a second confidence level of the at least one fourth text feature; the second confidence level characterizes the accuracy of the semantic information carried by the fourth text feature in describing the first category corresponding to the fourth text feature; when determining n fourth text features with the highest second confidence levels corresponding to each of the first categories from the at least one fourth text feature to obtain the at least one first text feature, the obtaining unit is specifically used to:
[0137] Perform feature extraction processing on the at least one first training text through the text encoding module to obtain text features of the at least one training text as the at least one fourth text feature; the at least one fourth text feature includes the fifth text feature;
[0138] Obtain the second confidence level of the fifth text feature according to the third similarity; the second confidence level of the fifth text feature is positively correlated with the third similarity.
[0139] Combined with any embodiment of the present application, the training unit is further used to:
[0140] Obtain a fourth video feature of the second training video; the fourth video feature is extracted from the second training video through a trained video feature extraction model;
[0141] Obtain a third loss according to a first difference between the third video feature and the fourth video feature; the third loss is positively correlated with the first difference;
[0142] Update parameters of the video classification training network according to the second loss and the third loss.
[0143] Combined with any embodiment of the present application, the training unit is further used to:
[0144] Obtain a seventh text feature of the second training text; the seventh text feature is extracted from the second training text through a trained text feature extraction model;
[0145] Obtain a fourth loss according to a second difference between the fifth text feature and the seventh text feature; the fourth loss is positively correlated with the second difference;
[0146] The updating the parameters of the video classification training network according to the second loss and the third loss includes:
[0147] Update the parameters of the video classification training network according to the second loss, the third loss, and the fourth loss.
[0148] In combination with any implementation manner of the present application, the training unit is configured to:
[0149] Obtain a training data set, where the training data set is constructed based on at least two of the following training tasks: closed set, long-tail distribution, few-shot, and open set;
[0150] Sample the second training video and the at least two first training texts from the training data set.
[0151] In a third aspect, an electronic device is provided, which includes: a processor and a memory. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the method according to the first aspect and any possible implementation manner thereof as described above.
[0152] In a fourth aspect, another electronic device is provided, which includes: a processor, a sending device, an input device, an output device, and a memory. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the method according to the first aspect and any possible implementation manner thereof as described above.
[0153] In a fifth aspect, a computer-readable storage medium is provided. A computer program is stored in the computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the method according to the first aspect and any possible implementation manner thereof as described above.
[0154] In a sixth aspect, a computer program product is provided. The computer program product includes a computer program or instructions. When the computer program or instructions run on a computer, the computer is caused to execute the method according to the first aspect and any possible implementation manner thereof as described above.
[0155] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. Description of the Drawings
[0156] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the following will describe the drawings required to be used in the embodiments of the present application or the background art.
[0157] The accompanying drawings here are incorporated into and form a part of this specification. These drawings illustrate embodiments consistent with this application and, together with the specification, are used to explain the technical solutions of this application.
[0158] Figure 1 It is a schematic flowchart of a video classification method provided for an embodiment of this application;
[0159] Figure 2 It is a structural diagram of a video classification network and a video classification training network provided for an embodiment of this application;
[0160] Figure 3 It is a schematic structural diagram of a video classification device provided for an embodiment of this application;
[0161] Figure 4 It is a schematic hardware structure diagram of a video classification device provided for an embodiment of this application. Detailed implementation manners
[0162] In order to enable those skilled in the art to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of this application.
[0163] The terms "first", "second", etc. in the specification and claims of this application and the above accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.
[0164] It should be understood that in this application, "at least one (item)" means one or more, "a plurality" means two or more, "at least two (items)" means two or three and more than three, and "and / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " can represent an "or" relationship between the front and rear associated objects, which refers to any combination of these items, including any combination of single items or plural items. For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple. The character " / " can also represent the division sign in mathematical operations. For example, a / b = a divided by b; 6 / 3 = 2. "At least one (item) of the following" or its similar expressions.
[0165] Reference to "embodiments" in this context means that the specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of this application. The occurrence of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0166] The execution subject of the embodiments of this application is a video classification device. Among them, the video classification device can be any electronic device that can execute the technical solutions disclosed in the method embodiments of this application. Optionally, the video classification device can be one of the following: a mobile phone, a computer, a tablet computer, a wearable intelligent device.
[0167] It should be understood that the method embodiments of this application can also be implemented by a processor executing computer program code. The embodiments of this application will be described below in conjunction with the accompanying drawings in the embodiments of this application. Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a video classification method provided by the embodiments of this application.
[0168] 101. Obtain a video to be processed and at least one first text feature.
[0169] In the embodiments of this application, the video to be processed can be an offline video or an online video. Among them, the offline video can be a video obtained by collecting through a camera or a mobile intelligent device. The online video can be a video collected in real time by a camera.
[0170] In an implementation of obtaining a video to be processed, a video classification device receives the video to be processed input by a user through an input component. The above input component includes at least one of the following: keyboard, mouse, touch screen, touch pad, audio input device.
[0171] In another implementation of obtaining a video to be processed, a video classification device receives the video to be processed sent by a terminal. The above terminal can be any one of the following: mobile phone, computer, tablet computer, server.
[0172] In still another implementation of obtaining a video to be processed, there is a communication connection between a video classification device and a camera, and the camera obtains the video to be processed collected by the camera through this communication connection.
[0173] Before proceeding with the following elaboration, the video categories are defined first. In the embodiments of the present application, video categories are used to distinguish different videos, and the basis for distinction is the video content. That is, video categories represent video content, and videos with the same category have the same video category.
[0174] For example, if video a is a video of a basketball game, then the category of video a can be basketball. For another example, if video b is a video of a singing competition, then the category of video b can be singing. For still another example, if video c and video d are both videos obtained by shooting outdoors, then the categories of video c and video d both include outdoors. If the video content of video c includes a person riding a bicycle and the video content of video d includes construction workers constructing outdoors, then the category of video c can also include cycling, and the category of video d can also include building construction.
[0175] In the embodiments of the present application, the first category is a video category, that is, the categories in at least one first category are all video categories. The first text feature is a feature of text, and the first text feature carries semantic information for describing the first category, that is, the first text feature carries semantic information for describing the video category.
[0176] For example, the first text feature a carries semantic information for describing the video category of basketball. Among them, the semantic information includes: Basketball is a physical confrontation sport centered around the hands. At this time, Basketball is a physical confrontation sport centered around the hands is the semantic information for describing the video category of basketball.
[0177] For another example, the first text feature b is the text feature of text c, where text c is: Construction is the production activity in the implementation stage of project construction. At this time, the first text feature b carries the semantic information of text c, that is, the semantic information carried by the first text feature b includes: Construction is the production activity in the implementation stage of project construction.
[0178] If the first text feature b carries semantic information for describing the video category of building construction, then building construction refers to the production activities in the implementation stage of project construction, which is the semantic information for describing the video category of building construction.
[0179] Optionally, a first text feature corresponds to a first category, that is, the semantic information carried by a first text feature is used to describe a first category.
[0180] In one implementation manner of obtaining at least one first text feature, the video classification device receives at least one first text feature input by the user through the input component.
[0181] In another implementation manner of obtaining at least one first text feature, the video classification device receives at least one first text feature sent by the terminal.
[0182] It should be understood that in the embodiments of the present application, the step of obtaining the video to be processed and the step of obtaining at least one first text feature can be executed simultaneously or separately. Moreover, in the case of separate execution, the present application does not limit the execution order of the step of obtaining the video to be processed and the step of obtaining at least one first text feature.
[0183] 102. Perform feature extraction processing on the above-mentioned video to be processed to obtain a first video feature.
[0184] In the embodiments of the present application, the first video feature is the video feature of the video to be processed, and the first video feature carries the feature information of the images in the video to be processed and the timestamp information of the images in the video to be processed. The feature information of the images includes at least one of the following information: the texture information of the images, the color information of the images, the shape information of the objects in the images, the brightness information of the images, and the spatial relationship information of different objects in the images. According to the timestamp information of the images, the timestamp of the images can be determined, where the timestamp of the images represents the playing time of the images in the video to be processed.
[0185] For example, the video to be processed includes a first frame image, a second frame image, and a third frame image, where the timestamp of the first frame image is t1, the timestamp of the second frame image is t2, and the timestamp of the third frame image is t3. At this time, the first video feature carries the feature information of the first frame image, the feature information of the second frame image, the feature information of the third frame image, the timestamp of the feature information of the first frame image is t1, the timestamp of the feature information of the second frame image is t2, and the timestamp of the feature information of the third frame image is t3.
[0186] 103. Fuse the above-mentioned first video feature and the above-mentioned at least one first text feature to obtain a first fusion feature.
[0187] Since the first video feature carries the video features of the video to be processed, and at least one first text feature carries semantic information for describing at least one first category, the video classification device can enrich the information for determining the category of the video to be processed by fusing the first video feature and at least one first text feature. In this way, during the subsequent process of classifying the video to be processed using the information carried by the first fusion feature, the accuracy of the classification result can be improved.
[0188] In a possible implementation manner, the video classification device obtains the first fusion feature by performing a concatenation process (concatenate) on the first video feature and at least one first text feature.
[0189] In another possible implementation manner, the video classification device determines the most matching text feature with the highest similarity to the first video feature from at least one first text feature. The most matching text feature and the first video feature are fused to obtain the first fusion feature.
[0190] 104. Classify the video to be processed according to the above first fusion feature to obtain the second category of the video to be processed.
[0191] In the embodiments of the present application, the second category is the classification result obtained by classifying the video to be processed according to the first fusion feature. For example, the video classification device classifies the video to be processed according to the first fusion feature and determines that the category of the video to be processed is basketball. At this time, the second category is basketball.
[0192] In a possible implementation manner, the video classification device processes the first fusion feature through the softmax function to implement the classification of the video to be processed and obtain the second category of the video to be processed.
[0193] In another possible implementation manner, the video classification device processes the first fusion feature through a support vector machine (SVM) to implement the classification of the video to be processed and obtain the second category of the video to be processed.
[0194] In the embodiments of the present application, when the video classification device obtains the first video feature through feature extraction processing of the video to be processed, the first fusion feature is obtained by fusing the first video feature and at least one first text feature, enriching the information for determining the category of the video to be processed. Then, the video classification device classifies the video to be processed according to the first fusion feature to obtain the second category of the video to be processed, which can improve the accuracy of the second category.
[0195] As an alternative implementation, when the number of the first text features is greater than 1, at least one of the first text features includes a second text feature and a third text feature, that is, the second text feature and the third text feature are any two of the at least one first text feature.
[0196] In this implementation, the video classification device further performs the following steps:
[0197] 201. Obtain a first similarity between the second text feature and the first video feature, and a second similarity between the third text feature and the first video feature.
[0198] In the embodiments of the present application, the first similarity is the similarity between the second text feature and the first video feature. The first similarity characterizes the similarity between the semantic information carried by the second text feature and the information carried by the first video feature. The second similarity is the similarity between the third text feature and the first video feature. The second similarity characterizes the similarity between the semantic information carried by the third text feature and the information carried by the first video feature.
[0199] In a possible implementation manner, the video classification device receives the first similarity and the second similarity input by the user through the input component.
[0200] In another possible implementation manner, the video classification device receives the first similarity and the second similarity sent by the terminal.
[0201] In still another possible implementation manner, the video classification device determines the cosine similarity between the second text feature and the first video feature to obtain the first similarity, and the video classification device determines the cosine similarity between the third text feature and the first video feature to obtain the second similarity.
[0202] 202. Obtain a first weight of the second text feature and a second weight of the third text feature according to the first similarity and the second similarity.
[0203] In the embodiments of the present application, when the first similarity is greater than the second similarity, the first weight is greater than the second weight. When the first similarity is equal to the second similarity, the first weight is equal to the second weight. When the first similarity is less than the second similarity, the first weight is less than the second weight.
[0204] Optionally, the first weight is positively correlated with the first similarity, and the second weight is positively correlated with the second similarity. In this way, the first weight can better characterize the first similarity, and the second weight can better characterize the second similarity.
[0205] After obtaining the first weight and the second weight, the video classification device performs the following steps during the execution of step 103:
[0206] 203. According to the above first weight and the above second weight, perform weighted fusion on the above second text feature and the above third text feature to obtain a second fusion feature.
[0207] Specifically, the video classification device uses the first weight as the weight of the second text feature and the second weight as the weight of the third text feature, and performs weighted fusion on the second text feature and the third text feature to obtain a second fusion feature.
[0208] 204. Perform fusion on the above second fusion feature and the above first video feature to obtain the above first fusion feature.
[0209] In this implementation manner, since both the second text feature and the third text feature carry semantic information describing the video category (i.e., the first category), and the semantic information describing the first category is beneficial to determining the category of the video to be processed. Therefore, using the information carried by the second text feature and the information carried by the third text feature to determine the second category of the video to be processed can improve the accuracy of the second category of the video to be processed.
[0210] Moreover, because the information carried by the second text feature is different from the information carried by the third text feature, the degree of improvement of the accuracy of the second category by the information carried by the second text feature is different from the degree of improvement of the accuracy of the second category by the information carried by the third text feature.
[0211] Specifically, the correlation between the information carried by the second text feature and the information carried by the first video feature is called the first correlation, and the correlation between the information carried by the third text feature and the information carried by the first video feature is called the second correlation. In the case where the first similarity is greater than the second similarity, the first correlation is greater than the second correlation. That is to say, the information carried by the second text feature is closer to the information carried by the first video feature than the information carried by the third text feature.
[0212] At this time, the information carried by the second text feature improves the accuracy of the second category to a greater extent than the information carried by the third text feature. Therefore, in the case where the first similarity is greater than the second similarity, making more use of the information carried by the second text feature is more conducive to improving the accuracy of the classification result of the video to be processed, that is, more conducive to improving the accuracy of the second category.
[0213] Therefore, when the first similarity is greater than the second similarity, the video classification device makes the first weight greater than the second weight, uses the first weight as the weight of the second text feature, uses the second weight as the weight of the third text feature, and performs weighted fusion on the second text feature and the third text feature to obtain a second fusion feature. In this way, when using the second fusion feature to classify the video to be processed, not only can the semantic information carried by the second text feature and the semantic information carried by the third text feature be used to determine the second category of the video to be processed, but also the semantic information carried by the second text feature and the semantic information carried by the third text feature can be better used to determine the second category of the video to be processed, thereby improving the accuracy of the second category.
[0214] Thereby, the video classification device fuses the second fusion feature and the first video feature to obtain a first fusion feature, which can make the information carried by the first fusion feature more conducive to improving the accuracy of the second category.
[0215] As an optional implementation manner, when the video classification device executes step 104, the following steps are performed:
[0216] 301. Predict a third category of the video to be processed and a first confidence level of the third category according to the first fusion feature.
[0217] In this implementation manner, the video classification device first predicts the category of the video to be processed according to the first fusion feature to obtain a third category and a first confidence level. Among them, the third category belongs to at least one first category, and the first confidence level is the confidence level of the third category, that is, the first confidence level represents the accuracy of the category of the video to be processed being the third category.
[0218] 302. When the first confidence level is less than or equal to the confidence level threshold, determine that the second category of the video to be processed is a category other than the at least one first category.
[0219] 303. When the first confidence level is greater than the confidence level threshold, determine that the third category is the second category.
[0220] Because the video classification device uses the semantic information carried by at least one first text feature during the process of determining the category of the video to be processed, and the semantic information carried by at least one first text feature is used to describe at least one first category, the third category of the video to be processed predicted according to the first fusion feature belongs to at least one first category. And when the category of the video to be processed belongs to at least one first category, the video classification device has a high classification accuracy for the video to be processed.
[0221] That is to say, when the category of the video to be processed belongs to at least one first category, the confidence level of the third category predicted by the video classification device is high, and when the confidence level of the third category is low, the probability that the third category does not belong to at least one first category is high.
[0222] To improve the classification accuracy of the video to be processed, when the confidence level of the third category is high, the video classification device determines that the third category is the category of the video to be processed. When the confidence level of the third category is low, the video classification device determines that the category of the video to be processed is a category other than at least one first category.
[0223] For example, if at least one first category includes basketball, then the category other than at least one first category is the category other than basketball. For another example, if at least one first category includes basketball, cycling, and construction, then the category other than at least one first category is the category other than basketball, cycling, and construction.
[0224] In the embodiments of the present application, based on the confidence threshold, it is determined whether the confidence level of the third category is high or low. Specifically, if the confidence level of the third category is greater than or equal to the confidence threshold, it indicates that the confidence level of the third category is high; if the confidence level of the third category is less than the confidence threshold, it indicates that the confidence level of the third category is low.
[0225] Therefore, when the first confidence level is less than or equal to the confidence threshold, the video classification device determines that the second category of the video to be processed is a category other than at least one first category; when the first confidence level is greater than the confidence threshold, the video classification device determines that the third category is the second category.
[0226] In this implementation manner, the video classification device predicts the third category of the video to be processed and the first confidence level of the third category according to the first fusion feature. When the first confidence level is less than or equal to the confidence threshold, the video classification device determines that the second category of the video to be processed is a category other than at least one first category. When the first confidence level is greater than the confidence threshold, the video classification device determines that the third category is the second category, which can improve the accuracy of the second category.
[0227] As an alternative implementation manner, the video classification device obtains at least one first text feature by performing the following steps:
[0228] 401. Obtain at least one fourth text feature and the second confidence level of the at least one fourth text feature.
[0229] In the embodiments of the present application, the fourth text feature is the feature of the text, and the fourth text feature carries semantic information for describing the first category, that is, the fourth text feature carries semantic information for describing the video category.
[0230] Optionally, a fourth text feature corresponds to a first category, that is, the semantic information carried by a fourth text feature is used to describe a first category.
[0231] In the embodiments of the present application, the second confidence level represents the accuracy of the semantic information carried by the fourth text feature in describing the first category corresponding to the fourth text feature.
[0232] For example, the first category described by the semantic information carried by the fourth text feature is basketball, and the semantic information carried by the fourth text feature is that one person passes the ball to another person. Since for ball games other than basketball, it may also involve one person passing the ball to another person (such as football, volleyball, hockey), the accuracy of the semantic information carried by the fourth text feature in describing basketball is low.
[0233] Another example is that the first category described by the semantic information carried by the fourth text feature is building construction, and the semantic information carried by the fourth text feature is that one person passes the ball to another person. Obviously, the semantic information carried by the fourth text feature does not match building construction, so the accuracy of the semantic information carried by the fourth text feature in describing basketball is low.
[0234] In an implementation manner of obtaining at least one fourth text feature, the video classification device receives at least one fourth text feature input by the user through the input component.
[0235] In another implementation manner of obtaining at least one fourth text feature, the video classification device receives at least one fourth text feature sent by the terminal.
[0236] 402. Determine n fourth text features with the highest second confidence level corresponding to each of the above first categories from the above at least one fourth text feature to obtain the above at least one first text feature.
[0237] In the embodiments of the present application, n is a positive integer. Since a fourth text feature corresponds to a first category, at least one fourth text feature can be divided into at least one category set based on the first category.
[0238] For example, the at least one fourth text feature includes the fourth text feature a, where the first category described by the semantic information carried by the fourth text feature a is basketball, that is, the fourth text feature a corresponds to basketball. At this time, the category set obtained by dividing the at least one fourth text feature based on the first category is the basketball set, and the basketball set includes the fourth text feature a.
[0239] For another example, at least one fourth text feature includes a fourth text feature b, a fourth text feature c, and a fourth text feature d. Among them, the semantic information carried by the fourth text feature b and the information carried by the fourth text feature c are both used to describe the first category of basketball, and the semantic information carried by the fourth text feature d is used to describe the first category of cycling. That is, both the fourth text feature b and the fourth text feature c correspond to the first category of basketball, and the fourth text feature d corresponds to the first category of cycling.
[0240] At this time, the category set obtained by dividing at least one fourth text feature based on the first category includes a basketball set and a cycling set. The basketball set includes the fourth text feature b and the fourth text feature c, and the cycling set includes the fourth text feature d.
[0241] By executing step 402, the video classification device can respectively determine the n fourth text features with the second highest confidence in each category set, that is, determine the n fourth text features with the second highest confidence corresponding to each first category, and then obtain at least one first text feature.
[0242] For example, at least one fourth text feature includes a fourth text feature A, a fourth text feature B, a fourth text feature C, a fourth text feature D, and a fourth text feature E. Among them, both the fourth text feature A and the fourth text feature B correspond to the first category of basketball, and the fourth text feature C, the fourth text feature D, and the fourth text feature E all correspond to the first category of cycling.
[0243] At this time, the fourth text features corresponding to the first category of basketball include the fourth text feature A and the fourth text feature B, and the fourth text features corresponding to the first category of cycling include the fourth text feature C, the fourth text feature D, and the fourth text feature E.
[0244] If n = 1, the second confidence of the fourth text feature A is z1, the second confidence of the fourth text feature B is z2, the second confidence of the fourth text feature C is z3, the second confidence of the fourth text feature D is z4, and the second confidence of the fourth text feature E is z5, where z1 is greater than z2, z3 is greater than z4, and z4 is greater than z5.
[0245] Since z1 is greater than z2, the fourth text feature with the second highest confidence corresponding to the first category of basketball is the fourth text feature A. Since z3 is greater than z4 and z4 is greater than z5, the fourth text feature with the second highest confidence corresponding to the first category of cycling is the fourth text feature C.
[0246] In this embodiment, the video classification device screens out at least one first text feature from at least one fourth text feature based on the second confidence level, which can improve the accuracy of the description of at least one first category by at least one first text feature.
[0247] As an alternative embodiment, the video classification device performs the following steps during the execution of step 102:
[0248] 501. Perform feature extraction processing on at least one frame of the to-be-processed images in the to-be-processed video to obtain the frame features of at least one frame of the to-be-processed images.
[0249] In the embodiments of the present application, the to-be-processed images and the frame features are in one-to-one correspondence, that is, the video classification device obtains the frame features of one frame of the to-be-processed images by performing feature extraction processing on one frame of the to-be-processed images. When the number of to-be-processed images is greater than 1, the video classification device obtains the frame features of each frame of the to-be-processed images by performing feature extraction processing on each frame of the to-be-processed images respectively. At this time, the frame features of at least one frame of the to-be-processed images include the frame features of each frame of the to-be-processed images.
[0250] In the embodiments of the present application, the frame features include the feature information of the to-be-processed images. For example, at least one frame of the to-be-processed images in the to-be-processed video includes image a. The video classification device obtains the frame features of image a by performing feature extraction processing on image a. At this time, the frame features of at least one frame of the to-be-processed images include the frame features of image a, where the frame features of image a include the feature information of the image.
[0251] 502. Fuse the frame features of at least one frame of the to-be-processed images and the timestamp information of at least one frame of the to-be-processed images to obtain the first video features.
[0252] In the embodiments of the present application, according to the timestamp information of at least one frame of the to-be-processed images, the timestamps of at least one frame of the to-be-processed images can be determined, and further, according to the timestamps of at least one frame of the to-be-processed images in the to-be-processed video, the timestamps of the frame features of at least one frame of the to-be-processed images can be determined.
[0253] By fusing the frame features of at least one frame of the to-be-processed images and the timestamp information of at least one frame of the to-be-processed images, the video classification device can make the frame features carry timestamp information. In this way, in the first video features obtained by fusion, the frame features carry timestamp information, and according to this timestamp information, the chronological relationship of different frame features in the time dimension can be determined.
[0254] As an alternative implementation, the video classification method provided by the embodiments of the present application is implemented through a video classification network, where the video classification network includes a video encoding module. By performing feature extraction processing on the video to be processed through the video encoding module, a first video feature can be obtained, that is, step 102 can be executed through the video encoding module.
[0255] In this implementation, the video classification device also executes the training process of the video classification network:
[0256] 601. Obtain a first training video.
[0257] In the embodiments of the present application, the first training video may be an offline video. Optionally, the first training video includes a label, where the label includes the category of the first training video.
[0258] In one implementation manner of obtaining the first training video, the video classification device receives the first training video input by the user through an input component.
[0259] In another implementation manner of obtaining the first training video, the video classification device receives the first training video sent by the terminal.
[0260] 602. Perform feature extraction processing on the first training video through the above video encoding module to obtain a second video feature.
[0261] The video classification device extracts the video feature of the first training video through step 602 to obtain a second video feature. The implementation manner of this step can refer to the implementation manner of step 102. Specifically, the first training video in this step corresponds to the video to be processed in step 102, and the second video feature in this step corresponds to the first video feature in step 102.
[0262] 603. Fuse the second video feature and the at least one first text feature to obtain a third fusion feature.
[0263] The video classification device extracts the video feature of the first training video through step 602 to obtain a second video feature. The implementation manner of this step can refer to the implementation manner of step 103. Specifically, the second video feature in this step corresponds to the first video feature in step 102, and the third fusion feature in this step corresponds to the first fusion feature in step 103.
[0264] 604. Obtain a fourth category of the first training video according to the third fusion feature.
[0265] The video classification device can classify the first training video according to the third fusion feature to obtain the fourth category of the first training video. The implementation method of this step can refer to the implementation method of step 104. Specifically, the third fusion feature in this step corresponds to the first fusion feature in step 103, and the fourth category in this step corresponds to the second category in step 104.
[0266] 605. Obtain the first loss of the video classification network according to the first difference between the above-mentioned fourth category and the label of the above-mentioned first training video.
[0267] In the embodiments of the present application, the first difference is positively correlated with the first loss. In a possible implementation manner, the video classification device processes the fourth category and the label of the first training video through a classification loss (cls loss) to obtain the first loss.
[0268] 606. Update the parameters of the above-mentioned video classification network according to the above-mentioned first loss to obtain a video classification network.
[0269] In a possible implementation manner, the video classification device updates the parameters of the video classification network according to the first loss until the first loss converges, and completes the training of the video classification network.
[0270] Through this implementation manner, the video classification device completes the training of the video classification network, so that the video classification method described above can be implemented through the video classification network.
[0271] In addition, there are currently many methods for classifying videos through deep learning techniques. However, these methods have low classification accuracy when there is at least one of a closed set, long-tail distribution, and few-shot in the training data. In the training method of the video classification network provided in the embodiments of the present application, since the information carried by the second video feature is used for classification and the information carried by at least one text feature is also used for classification, the classification accuracy can be improved when there is at least one of a closed set, long-tail distribution, and few-shot.
[0272] As an optional implementation manner, the video classification device executes the following steps during the execution of step 606:
[0273] 701. Update the parameters of the video classification network except for the parameters of the video encoding module according to the first loss.
[0274] In this implementation manner, during the training process of the video classification network, the video classification device does not update the parameters of the video encoding module, but updates the parameters except for the parameters of the video encoding module according to the first loss.
[0275] The video encoding module is obtained by training a video classification training network. The video classification training network includes the video encoding module. That is, during the training process of the video classification training network, the parameters of the video encoding module are also updated. When the training of the video classification training network is completed, the training of the video encoding module is also completed. And the video encoding module in the video classification network is the video encoding module that has completed training.
[0276] In this embodiment, the video classification device also performs the training process of the video classification training network:
[0277] 702. Obtain a second training video and at least two first training texts.
[0278] In the embodiments of the present application, the second training video may be an offline video. Optionally, the second training video includes a label, where the label includes the category of the second training video. It should be understood that the second training video and the first training video may be the same or different, and the present application does not make any limitations in this regard.
[0279] In one implementation manner of obtaining the second training video, the video classification device receives the second training video input by the user through the input component.
[0280] In another implementation manner of obtaining the second training video, the video classification device receives the second training video sent by the terminal.
[0281] In the embodiments of the present application, the label of the first training text includes the category described by the first training text (i.e., the fifth category). For example, if the label of the first training text is basketball, then the first training text is a text for describing basketball.
[0282] The labels of at least two first training texts include at least two fifth categories, and the labels of at least two first training texts include at least one first category. For example, the labels of at least two first training texts include basketball and cycling. At this time, at least two fifth categories include basketball and cycling, and at least one first category is at least one of basketball and cycling. For example, at least one first category is basketball, or at least one first category is cycling, or at least one first category includes both basketball and cycling.
[0283] In one implementation manner of obtaining at least two first training texts, the video classification device receives at least two first training texts input by the user through the input component.
[0284] In another implementation manner of obtaining at least two first training texts, the video classification device receives at least two first training texts sent by the terminal.
[0285] In the embodiments of the present application, at least two first training texts include a second training text, where the fifth category of the second training text is the same as the sixth category of the second training video, that is, the fifth category indicated by the label of the second training text is the same as the sixth category. For example, if both the fifth category and the sixth category are basketball, then the second training text is a text describing basketball.
[0286] 703. Perform feature extraction processing on the above-mentioned second training video through the above-mentioned video encoding module to obtain the third video feature of the above-mentioned second training video.
[0287] The video classification device extracts the video feature of the second training video by executing step 703 to obtain the third video feature. The implementation manner of this step can refer to the implementation manner of step 102. Specifically, the second training video in this step corresponds to the video to be processed in step 102, and the third video feature in this step corresponds to the first video feature in step 102.
[0288] 704. Obtain the second loss of the above-mentioned video classification training network according to the third similarity between the above-mentioned third video feature and the above-mentioned second training text.
[0289] In the embodiments of the present application, the similarity between the third video feature and the second training text is the third similarity. The second loss is negatively correlated with the third similarity, that is, the smaller the third similarity, the larger the second loss. In other words, the smaller the similarity between the third video feature and the second training text, the larger the second loss.
[0290] 705. Update the parameters of the above-mentioned video classification training network according to the above-mentioned second loss to obtain a video classification training network.
[0291] In a possible implementation manner, the video classification device updates the parameters of the video classification training network according to the second loss until the second loss converges, completing the training of the video classification training network. At this time, the training of the video encoding module is also completed.
[0292] In this implementation manner, the fifth category of the second training text is the same as the sixth category of the second training video. The video classification device determines the second loss according to the third similarity between the third video feature of the second training video and the second training text, where the second loss and the third similarity are negatively correlated. In this way, updating the parameters of the video classification training network according to the second loss can make the third video feature extracted by the video encoding module of the video classification training network have a high matching degree with the second training text. In other words, make the third video feature extracted by the video encoding module have a high matching degree with the text of the sixth category.
[0293] It should be understood that in the embodiments of the present application, the second training text is a description object determined for concisely describing the implementation process of the technical solution, and it should not be understood that in the training process of the video classification training network, the second loss is obtained only by processing the second training text, and the parameters of the video classification training network are updated according to the second loss. In practical applications, the video classification device can perform the same processing as that of the second training text on the training text in which any one of the fifth categories in at least two first training texts is the same as the sixth category of the second training video, that is, for the training text in which the category in at least two first training texts is the same as the sixth category of the second training video, the same processing as that of the second training text can be performed.
[0294] As an optional implementation manner, at least two first training texts further include a third training text, where the fifth category described in the third training text is different from the above-mentioned sixth category, that is, the fifth category indicated by the label of the third training text is different from the sixth category. For example, the fifth category of the third training text is basketball and the sixth category is football.
[0295] In this implementation manner, the video classification device further performs the following steps:
[0296] 801. Determine the fourth similarity between the above-mentioned third video feature and the above-mentioned third training text.
[0297] When the fourth similarity is obtained, the video classification device performs the following steps in the process of executing step 704:
[0298] 802. Obtain the second loss of the above-mentioned video classification training network according to the above-mentioned third similarity and the above-mentioned fourth similarity.
[0299] In the embodiments of the present application, the second loss is positively correlated with the fourth similarity, that is, the smaller the fourth similarity, the smaller the second loss. In other words, the smaller the similarity between the third video feature and the third training text, the smaller the second loss.
[0300] In this implementation manner, the fifth category of the third training text is different from the sixth category of the second training video. When the video classification device determines the fourth similarity between the third video feature of the second training video and the third training text, the second loss is obtained according to the third similarity and the fourth similarity, where the third similarity is negatively correlated with the second loss and the fourth similarity is positively correlated with the second loss.
[0301] In this way, by updating the parameters of the video classification training network according to the second loss, the third video features extracted by the video encoding module of the video classification training network can have a high degree of matching with the second training text, and at the same time, the third video features extracted by the video encoding module of the video classification training network can have a low degree of matching with the third training text. In other words, the third video features extracted by the video encoding module have a high degree of matching with the text of the sixth category, and a low degree of matching with the text of categories other than the sixth category.
[0302] It should be understood that in the embodiments of the present application, the third training text is a description object determined for concisely describing the implementation process of the technical solution. In actual applications, the video classification device can perform the same processing as that of the third training text on any training text in which the fifth category of at least two first training texts is different from the sixth category of the second training video, that is, for the training text in which the category of at least two first training texts is different from the sixth category of the second training video, the same processing as that of the third training text can be performed.
[0303] As an optional implementation manner, the video classification training network further includes a text encoding module. The text encoding module is used to perform feature extraction processing on the text to obtain text features. The video classification device further performs the following steps:
[0304] 901. Perform feature extraction processing on the second training text through the above text encoding module to obtain the fifth text features of the second training text.
[0305] 902. Determine the similarity between the fifth text features and the third video features as the third similarity between the third video features and the second training text.
[0306] In this implementation manner, when the video classification device extracts the fifth text features of the second training text through the text encoding module, it determines the similarity between the fifth text features and the third video features as the third similarity.
[0307] As an optional implementation manner, the video classification device further performs the following steps: 1001. Perform feature extraction processing on the third training text through the above text encoding module to obtain the seventh text features of the third training text.
[0308] 1002. Determine the similarity between the seventh text features and the third video features as the fourth similarity between the third video features and the second training text.
[0309] In this implementation, when the video classification device extracts the seventh text feature of the third training text through the text encoding module, it determines that the similarity between the seventh text feature and the third video feature is the fourth similarity.
[0310] As an optional implementation, during the execution of step 401, the video classification device performs the following steps:
[0311] 1101. Perform feature extraction processing on the at least one first training text through the above text encoding module to obtain the text features of the at least one training text as the at least one fourth text feature.
[0312] In step 1101, the video classification device extracts the text features of each first training text in the at least one first training text through the text encoding module to obtain at least one fourth text feature. At this time, the at least one fourth text feature includes the text feature of the second training text (i.e., the fifth text feature).
[0313] 1102. Obtain the second confidence level of the fifth text feature according to the above third similarity.
[0314] Since the category of the second training video (i.e., the sixth category) is the same as the fifth category of the second training text, and the third video feature is the video feature of the second training video, and the fifth text feature is the text feature of the second training text, the higher the similarity between the third video feature and the fifth text feature, the more accurate the information carried by the fifth text feature for describing the sixth category, and thus the higher the second confidence level of the fifth text feature can be determined. Therefore, the second confidence level of the fifth text feature is positively correlated with the third similarity, that is, the greater the similarity between the third video feature and the second training text, the greater the second confidence level of the fifth text feature.
[0315] In this implementation, the video classification device performs feature extraction processing on the at least one first training text through the text encoding module to obtain at least one fourth text feature. In the case where the third similarity is positively correlated with the second confidence level of the fifth text feature, obtaining the second confidence level of the fifth text feature according to the third similarity can improve the accuracy of the second confidence level of the fifth text feature.
[0316] It should be understood that the third similarity and the second confidence level of the fifth text feature in this implementation are description objects determined for concisely describing the implementation process of the technical solution. In actual applications, for the fourth text feature in the at least one fourth text feature whose category is the same as the sixth category, the second confidence level can be determined in the same way as determining the second confidence level of the fifth text feature above.
[0317] As an alternative implementation, when the video classification device obtains at least one fourth text feature by executing step 1101, it further determines the second confidence level of the seventh text feature by executing the following steps:
[0318] 1103. Obtain the second confidence level of the seventh text feature according to the above fourth similarity.
[0319] Since the category of the second training video (i.e., the sixth category) is different from the fifth category of the third training text, and the third video feature is the video feature of the second training video, and the seventh text feature is the text feature of the third training text, the higher the similarity between the third video feature and the seventh text feature, the closer the information carried by the seventh text feature is to the sixth category, and further the lower the accuracy of the description of the information carried by the seventh text feature for the fifth category of the third training text. Therefore, it can be determined that the second confidence level of the seventh text feature is lower. Therefore, the second confidence level of the seventh text feature is negatively correlated with the fourth similarity, that is, the smaller the similarity between the third video feature and the third training text, the greater the second confidence level of the seventh text feature.
[0320] In this implementation, when the fourth similarity and the second confidence level of the seventh text feature are negatively correlated, the video classification device obtains the second confidence level of the seventh text feature according to the fourth similarity, which can improve the accuracy of the second confidence level of the seventh text feature.
[0321] It should be understood that the fourth similarity and the second confidence level of the seventh text feature in this implementation are description objects determined for concisely describing the implementation process of the technical solution. In actual applications, for the fourth text features in at least one fourth text feature whose categories are different from the sixth category, the second confidence level can be determined in the same way as determining the second confidence level of the seventh text feature above.
[0322] As an alternative implementation, the video classification device further executes the following steps:
[0323] 1201. Obtain the fourth video feature of the second training video.
[0324] In the embodiments of the present application, the fourth video feature is extracted from the second training video by a trained video feature extraction model. Optionally, the trained video feature extraction module is a CLIP model.
[0325] In one implementation of obtaining the fourth video feature, the video classification device receives the fourth video feature input by the user through the input component.
[0326] In another implementation of obtaining the fourth video feature, the video classification device receives the fourth video feature sent by the terminal.
[0327] In yet another implementation for obtaining the fourth video feature, the video classification device performs feature extraction processing on the second training video through a trained video feature extraction model to obtain the fourth video feature.
[0328] 1202. Obtain a third loss based on the first difference between the above-mentioned third video feature and the above-mentioned fourth video feature.
[0329] In the embodiments of the present application, the third loss is positively correlated with the first difference. In a possible implementation, the video classification device uses the first difference as the third loss.
[0330] In another possible implementation, the video classification device uses the product of the first difference and the first hyperparameter as the third loss.
[0331] In the case of obtaining the third loss, the video classification device performs the following steps during the execution of step 705:
[0332] 1203. Update the parameters of the above-mentioned video classification training network according to the above-mentioned second loss and the above-mentioned third loss.
[0333] In a possible implementation, the video classification device calculates the sum of the second loss and the third loss to obtain the total loss of the video classification training network. Update the parameters of the video classification training network according to the total loss until the total loss converges, and complete the training of the video classification training.
[0334] In another possible implementation, the video classification device performs weighted summation on the second loss and the third loss to obtain the total loss of the video classification training network. Update the parameters of the video classification training network according to the total loss until the total loss converges, and complete the training of the video classification training network.
[0335] In yet another possible implementation, the video classification device calculates the sum of the second loss and the third loss to obtain the total loss of the video classification training network. Update the parameters of the video classification training network according to the total loss until the total loss, the second loss, and the third loss all converge, and complete the training of the video classification training network.
[0336] In this implementation, the fourth video feature is extracted from the second training video through a trained video feature extraction model. Therefore, the fourth video feature can be used as the ground truth (GT) of the video feature of the second training video. The video classification device obtains the third loss based on the first difference between the third video feature and the fourth video feature. In this way, during the training of the video classification network according to the second loss and the third loss, the accuracy of the third video feature extracted by the video encoding module from the second training video can be improved.
[0337] As an alternative implementation, the video classification device further performs the following steps:
[0338] 1301. Obtain the seventh text feature of the second training text described above.
[0339] In the embodiments of the present application, the seventh text feature is extracted from the second training text by a trained text feature extraction model. Optionally, the trained text feature extraction module is a CLIP model.
[0340] In one implementation manner of obtaining the seventh text feature, the video classification device receives the seventh text feature input by the user through the input component.
[0341] In another implementation manner of obtaining the seventh text feature, the video classification device receives the seventh text feature sent by the terminal.
[0342] In yet another implementation manner of obtaining the seventh text feature, the video classification device performs feature extraction processing on the second training text through a trained text feature extraction model to obtain the seventh text feature.
[0343] 1302. Obtain a fourth loss according to the second difference between the fifth text feature and the seventh text feature described above.
[0344] In the embodiments of the present application, the fourth loss is positively correlated with the second difference. In one possible implementation manner, the video classification device uses the second difference as the fourth loss.
[0345] In another possible implementation manner, the video classification device uses the product of the second difference and the second hyperparameter as the fourth loss.
[0346] In the case of obtaining the fourth loss, the video classification device performs the following steps during the execution of step 1203:
[0347] 1303. Update the parameters of the video classification training network according to the second loss, the third loss, and the fourth loss described above.
[0348] In one possible implementation manner, the video classification device calculates the sum of the second loss, the third loss, and the fourth loss to obtain the total loss of the video classification training network. Update the parameters of the video classification training network according to the total loss until the total loss converges, and complete the training of the video classification training.
[0349] In another possible implementation manner, the video classification device performs weighted summation on the second loss, the third loss, and the fourth loss to obtain the total loss of the video classification training network. Update the parameters of the video classification training network according to the total loss until the total loss converges, and complete the training of the video classification training network.
[0350] In yet another possible implementation, the video classification device calculates the sum of the second loss, the third loss, and the fourth loss to obtain the total loss of the video classification training network. The parameters of the video classification training network are updated according to the total loss until the total loss, the second loss, the third loss, and the fourth loss all converge, completing the training of the video classification training network.
[0351] In this implementation, the seventh text feature is extracted from the second training text by the trained text feature extraction model. Therefore, the seventh text feature can be used as the GT of the text feature of the second training text. The video classification device obtains the fourth loss according to the second difference between the fifth text feature and the seventh text feature. In this way, during the training of the video classification network according to the second loss, the third loss, and the fourth loss, the accuracy of the fifth text feature extracted by the text encoding module from the second training text can be improved.
[0352] As an optional implementation, the video classification device performs the following steps during the execution of step 601:
[0353] 1401. Obtain the training data set.
[0354] In the embodiments of the present application, the training data set is constructed based on at least two of the following training tasks: closed set, long-tail distribution, few-shot, open set.
[0355] When constructing the training data set according to the closed set, it is necessary to consider enriching the information carried by the training data set as much as possible. Therefore, in a possible implementation, when constructing the training data set according to the closed set, while constructing the target video with a label, a target text for describing the category indicated by the label can be constructed. The combination of the target video and the target text is used as the training data set.
[0356] In this way, when using the training data set to train the model, the semantic information provided by the target text can be utilized to improve the classification accuracy of the model.
[0357] When constructing the training data set according to the long-tail distribution, it is necessary to make the training data set simulate the long-tail distribution. For example, in the long-tail distribution training task, the number of samples of the category of basketball is large, and the number of samples of the category of building construction is small. At this time, the training data set constructed according to the long-tail distribution includes samples of the category of basketball and samples of the category of building construction, and the number of samples of the category of building construction is less than the number of samples of the category of basketball.
[0358] Moreover, when constructing the training data set according to the long-tail distribution, it is also necessary to consider that using the training data set to train the model can improve the classification accuracy of the tail class.
[0359] In a possible implementation manner, a training dataset is constructed according to the long-tail distribution. While constructing a target video that conforms to the long-tail distribution and has labels, a target text for describing the category indicated by the labels is constructed. The combination of the target video and the target text is used as the training dataset.
[0360] In this way, when using the training dataset to train the model, the semantic information provided by the target text can be utilized to improve the classification accuracy of the model for the tail classes.
[0361] Constructing a training dataset according to few-shot samples requires the training dataset to be available for performing few-shot training tasks. For example, the few-shot training task is to improve the classification accuracy of videos of the category of building construction when the number of samples of the category of building construction is small. At this time, the training dataset constructed according to few-shot samples includes samples of the category of building construction, and the number of samples of the category of building construction is small.
[0362] Moreover, when constructing a training dataset according to few-shot samples, it is also necessary to consider that using the training dataset to train the model can improve the classification accuracy of the categories with few samples.
[0363] In a possible implementation manner, when constructing a training dataset according to few-shot samples, while constructing a target video that can be used to perform few-shot training tasks, a target text for describing the category indicated by the labels of the target video is constructed. The combination of the target video and the target text is used as the training dataset.
[0364] In this way, when using the training dataset to train the model, the semantic information provided by the target text can be utilized to improve the classification accuracy of the model for the categories with few samples.
[0365] In a possible implementation manner, constructing a training dataset according to the open set requires the training dataset to be available for performing open set training tasks. Moreover, when constructing a training dataset according to the open set, it is also necessary to consider that using the training dataset to train the model can improve the classification accuracy of the categories that do not appear in the training dataset.
[0366] For example, the open set training task is to use samples of the category of basketball for training to improve the classification accuracy of videos of categories other than basketball. At this time, the training dataset constructed according to the open set includes samples of the category of basketball, and the test dataset includes samples of categories other than basketball.
[0367] In a possible implementation manner, when constructing a training dataset according to the open set, while constructing a target video that can be used to perform open set training tasks, a target text for describing the category indicated by the labels of the target video is constructed. The combination of the target video and the target text is used as the training dataset.
[0368] In this way, when training the model using the training dataset, the semantic information provided by the target text can be utilized to improve the classification accuracy of the model for the category corresponding to the target video (i.e., the categories that have appeared in the training dataset), and further improve the classification accuracy for the categories that have not appeared in the training dataset.
[0369] In the embodiments of the present application, the training dataset is constructed based on at least two of the closed set, long-tailed distribution, few-shot, and open set. Therefore, the training dataset can be used to enable the model to perform at least two different training tasks simultaneously, which can improve the training efficiency.
[0370] Optionally, the above training dataset is constructed according to the Kinetics400 dataset and at least two training datasets of the closed set, long-tailed distribution, few-shot, and open set.
[0371] 1402. Sample the above second training video and the above at least two first training texts from the above training dataset.
[0372] In this implementation manner, when both the second training video and the at least two first training texts are sampled from the training dataset, the process of training the video classification training network using the second training video and the at least two first training texts described above is equivalent to training the video classification training network using the second training video and the at least two first training texts during the process of training the video classification training network using the training dataset.
[0373] The video classification device can train the video classification training network through the training dataset, enabling the video classification network to perform at least two different training tasks simultaneously, thereby improving the training efficiency.
[0374] As an optional implementation manner, Figure 2 shows a structural diagram of a video classification network and a video classification training network. As Figure 2 shown, the video classification training network includes a first stage and a second stage. Among them, the first stage is the training of the video classification training network, and the second stage is the training of the video classification network.
[0375] During the training process of the first stage, the training data includes the second training video with the category of category 1, the second training video with the category of category 2, the second training video with the category of category 3, the first training text with the described category of category 1, the first training text with the described category of category 2, and the first training text with the described category of category 2.
[0376] The frame encoding module performs feature extraction processing on each frame of training images in each second training video to obtain the frame features of at least one frame of training images. The temporal aggregation module fuses the frame features of at least one frame of training images and the timestamp information of at least one frame of training images to obtain the video features (i.e., the third video features) of each second training video.
[0377] The text encoding module performs feature extraction processing on each first training text to obtain the text features of at least one first training text. Among them, at least one first training text includes the above-mentioned second training text and the above-mentioned third training text, and the text features of at least one first training text include the above-mentioned fifth text feature and the above-mentioned seventh text feature.
[0378] For the video features of each second training video and the text features in at least one first training text, the noise contrastive estimation (NCE) function is calculated respectively to obtain the NCE loss, that is, the above-mentioned second loss, which is denoted by hereinafter.
[0379] Optionally, satisfies the following formula:
[0380]
[0381] wherein, represents the NCE loss obtained based on the text features in at least one first training text, represents the NCE loss obtained based on the video features of each second training video, sim(·,·) represents the cosine similarity, and δ is a hyperparameter.
[0382] represents the set of first training videos of the i-th category in all second training videos, and any video in has the same category as T i , where T i represents the text used to describe the i-th category in at least one first training text. represents the set of texts describing the i-th category in at least one first training text, and any text in describes the same category as v i , where v i represents the second training video of the i-th category. v k represents the second training video whose category is not the i-th category, and T k represents the text in at least one first training text that describes a category other than the i-th category.
[0383] In a possible implementation, is used as the loss of the video classification training network (i.e., ), and by updating the parameters of the video classification network, the effect of reducing the distance between video features and text features of the same category and increasing the distance between video features and text features of different categories can be achieved.
[0384] Optionally, in order to reduce the negative impact of overfitting caused by at least one first training text, the CLIP model is used as the teacher model to perform knowledge distillation on the video features of the second training video output by the video classification training network and the text features of at least one first training text output by the text encoding module, and a distillation loss is obtained, where the distillation loss is the sum of the above third loss and fourth loss. Then, according to the distillation loss and obtain
[0385] Optionally, the distillation loss satisfies the following formula:
[0386]
[0387] where S represents the cosine similarity score based on the output of the video classification network, S′ represents the cosine similarity score based on the output of the CLIP model, and the calculation method of S′ v is the same as that of S v , and the calculation method of S′ τ is the same as that of S τ .
[0388] After obtaining and , and satisfy the following formula:
[0389]
[0390] where α is a non-negative number less than or equal to 1.
[0391] When the training of the first stage is completed, the training of the second stage is executed. First, the temporal aggregation module of the video classification network is initialized by the temporal aggregation module of the video classification training network, and the frame encoding module of the video classification network is initialized by the frame encoding module of the video classification training network. Moreover, during the training process of the second stage, the parameters of the frame encoding module and the temporal aggregation module are not updated.
[0392] During the training process of the second stage, the first training video is processed by the frame encoding module and the temporal aggregation module in sequence to obtain the fifth video feature of the first training video. At least one text feature of the at least one first training text extracted by the text encoding module during the training process of the first stage is screened through a text selection rule to obtain at least one screened text feature.
[0393] Specifically, the text selection rule is that for text features with the same category described in the text features of at least one first training text, n text features less than or equal to the screening threshold are selected, where is obtained through formula (1).
[0394] The fifth video feature is linearized and normalized to obtain a query. Each screened text feature is linearly and normalized in sequence to obtain a key. The transpose of each screened text feature is calculated to obtain a value.
[0395] The query, key, and value are processed through an attention mechanism to obtain a fourth fusion feature. Optionally, the attention mechanism is a bi-modal attention head. Optionally, the query, key, value, and fourth fusion feature satisfy the following formula:
[0396]
[0397] where Q represents the query, K represents the key, V represents the value, D represents the dimension of the screened text feature, and Softmax(·) represents the logistic regression function.
[0398] The fifth video feature is linearized to obtain the above-mentioned second video feature. The fourth fusion feature and the second video feature are element-wise added to obtain the above-mentioned third fusion feature.
[0399] According to the third fusion feature, the first training video is classified to obtain the fourth category of the first training video. According to the first difference between the fourth category and the label of the first training video, the first loss of the video classification network is obtained, hereinafter denoted by represented.
[0400] Optionally, satisfies the following formula:
[0401]
[0402] where represents the cross entropy loss function, y represents the label of the first training video, and MLP(EV ) represents the second video feature, and G represents the fourth fusion feature.
[0403] According to the first loss, update the parameters of the video classification network, and the training of the video classification network can be completed.
[0404] Optionally, when the training task of the video classification network includes an open set, according to the third fusion feature, predict the seventh category of the first training video and the third confidence level of the seventh category. When the third confidence level is less than or equal to the confidence level threshold, the post-processing module determines that the fourth category of the first training video is a category other than at least one first category; when the third confidence level is greater than the confidence level threshold, the post-processing module determines that the fourth category of the first training video is the seventh category.
[0405] Based on the technical solutions provided in the embodiments of the present application, several possible application scenarios are also provided in the embodiments of the present application.
[0406] Scenario A: As the number of network videos increases, classifying videos helps users find corresponding videos according to their own needs. Based on the technical solutions provided in the embodiments of the present application, the video classification device can classify network videos and add corresponding labels to the network videos based on the categories of the network videos. In this way, when a user needs to view a video, videos with labels that match the user's needs can be pushed to the user.
[0407] Scenario B: In recent years, the application of driverless has been increasing. Currently, driverless technology is usually implemented through a deep learning model. Specifically, the deep learning model has an environmental perception function. Through the environmental perception function, the environment where the vehicle is located can be perceived, and the vehicle can then determine the driving route according to the perception result. Among them, the deep learning model has an environmental perception function through training the deep learning model.
[0408] Due to the limited driverless training data used for training the deep learning model, the driving scenarios learned by deep learning through training are also limited, which results in low perception accuracy of the deep learning model for scenarios not covered by the driverless training data, and thus leads to the existence of potential safety hazards.
[0409] For example, the driverless training data used for training the deep learning model includes videos collected in an urban scenario. Then, when the scenario where the vehicle is located is a highway scenario or a rural scenario, the perception accuracy of the deep learning model is low.
[0410] Based on the technical solution provided by the implementation of this application, during the driving process of a vehicle, a video of the scene where the vehicle is located can be taken to obtain a video to be processed, and then the scene where the vehicle is located can be determined by classifying the video to be processed. When it is determined that the scene where the vehicle is located is a scene not involved in the driverless training data, an alarm prompt is output, thereby reducing potential safety hazards.
[0411] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order that constitutes any limitation on the implementation process, and the specific execution order of each step should be determined according to its function and possible internal logic.
[0412] If the technical solution of this application involves personal information, before the product applying the technical solution of this application processes personal information, the rules for processing personal information have been clearly informed, and the individual's independent consent has been obtained. If the technical solution of this application involves sensitive personal information, before the product applying the technical solution of this application processes sensitive personal information, the individual's separate consent has been obtained, and at the same time, the requirements of "express consent" are met. For example, at a personal information collection device such as a camera, a clear and prominent sign is set to inform that the personal information collection range has been entered and personal information will be collected. If an individual voluntarily enters the collection range, it is regarded as consenting to the collection of their personal information; or on the device for processing personal information, when the rules for processing personal information are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information by themselves; among them, personal information processing may include information such as personal information processors, personal information processing purposes, processing methods, and types of personal information processed.
[0413] The method of the embodiment of this application is elaborated in detail above, and the device of the embodiment of this application is provided below.
[0414] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a video classification device provided by an embodiment of this application. The video classification device 1 includes an acquisition unit 11, a first processing unit 12, a second processing unit 13, and a third processing unit 14. Optionally, the video classification device 1 further includes a training unit 15. Specifically:
[0415] The acquisition unit 11 is configured to acquire a video to be processed and at least one first text feature; the at least one first text feature carries semantic information for describing at least one first category;
[0416] The first processing unit 12 is configured to perform feature extraction processing on the video to be processed to obtain a first video feature;
[0417] A second processing unit 13, configured to fuse the first video feature and the at least one first text feature to obtain a first fused feature;
[0418] A third processing unit 14, configured to classify the video to be processed according to the first fused feature to obtain a second category of the video to be processed.
[0419] Combined with any implementation manner of the present application, when the number of the first text features is greater than 1, the at least one first text feature includes a second text feature and a third text feature;
[0420] The obtaining unit 11 is further configured to obtain a first similarity between the second text feature and the first video feature, and a second similarity between the third text feature and the first video feature;
[0421] The second processing unit 13 is configured to:
[0422] Obtain a first weight of the second text feature and a second weight of the third text feature according to the first similarity and the second similarity; when the first similarity is greater than the second similarity, the first weight is greater than the second weight; when the first similarity is equal to the second similarity, the first weight is equal to the second weight;
[0423] Perform weighted fusion on the second text feature and the third text feature according to the first weight and the second weight to obtain a second fused feature;
[0424] Fuse the second fused feature and the first video feature to obtain the first fused feature.
[0425] Combined with any implementation manner of the present application, the third processing unit 14 is configured to:
[0426] Predict a third category of the video to be processed and a first confidence level of the third category according to the first fused feature; the third category belongs to the at least one first category;
[0427] When the first confidence level is less than or equal to a confidence level threshold, determine that the second category of the video to be processed is a category other than the at least one first category;
[0428] When the first confidence level is greater than the confidence level threshold, determine that the third category is the second category.
[0429] Combined with any implementation manner of the present application, the obtaining unit 11 is configured to:
[0430] Obtain at least one fourth text feature and a second confidence level of the at least one fourth text feature; the second confidence level characterizes the accuracy of the semantic information carried by the fourth text feature in describing the first category corresponding to the fourth text feature.
[0431] From the at least one fourth text feature, determine n fourth text features with the highest second confidence level corresponding to each of the first categories, to obtain the at least one first text feature.
[0432] In combination with any embodiment of the present application, the first processing unit 12 is configured to:
[0433] Perform feature extraction processing on at least one frame of to-be-processed image in the to-be-processed video, to obtain the frame features of the at least one frame of to-be-processed image.
[0434] Fuse the frame features of the at least one frame of to-be-processed image and the timestamp information of the at least one frame of to-be-processed image, to obtain the first video feature.
[0435] In combination with any embodiment of the present application, the video classification method is implemented through a video classification network, and the video classification network includes a video encoding module.
[0436] The first processing unit 12 is configured to perform feature extraction processing on the to-be-processed video through the video encoding module, to obtain a first video feature.
[0437] The video classification device further includes a training unit 15, and the training unit 15 is configured to execute the training process of the video classification network:
[0438] Obtain a first training video.
[0439] Perform feature extraction processing on the first training video through the video encoding module, to obtain a second video feature.
[0440] Fuse the second video feature and the at least one first text feature, to obtain a third fusion feature.
[0441] Obtain a fourth category of the first training video according to the third fusion feature.
[0442] Obtain a first loss of the video classification network according to a first difference between the fourth category and the label of the first training video.
[0443] Update parameters of the video classification network according to the first loss, to obtain the video classification network.
[0444] In combination with any embodiment of the present application, the training unit 15 is configured to:
[0445] Update the parameters in the video classification network except for the parameters of the video encoding module according to the first loss;
[0446] The video encoding module is obtained by training a video classification training network, and the video classification training network includes the video encoding module;
[0447] The video classification method further includes a training process of the video classification training network:
[0448] Obtain a second training video and at least two first training texts; the labels of the at least two first training texts include at least two fifth categories, and the labels of the at least two first training texts include the at least one first category; the at least two first training texts include the second training text, and the fifth category of the second training text is the same as the sixth category of the second training video;
[0449] Perform feature extraction processing on the second training video through the video encoding module to obtain a third video feature;
[0450] Obtain a second loss of the video classification training network according to the third similarity between the third video feature and the second training text; the second loss is negatively correlated with the third similarity;
[0451] Update the parameters of the video classification training network according to the second loss to obtain the video classification training network.
[0452] Combined with any implementation manner of the present application, the at least two first training texts further include a third training text, and the fifth category described by the third training text is different from the sixth category;
[0453] The training unit 15 is further configured to:
[0454] Determine a fourth similarity between the third video feature and the third training text;
[0455] Obtain a second loss of the video classification training network according to the third similarity and the fourth similarity; the second loss is positively correlated with the fourth similarity.
[0456] Combined with any implementation manner of the present application, the video classification training network further includes a text encoding module;
[0457] The training unit 15 is further configured to:
[0458] Perform feature extraction processing on the second training text through the text encoding module to obtain a fifth text feature of the second training text;
[0459] Determine the similarity between the fifth text feature and the third video feature as the third similarity between the third video feature and the second training text.
[0460] Combined with any implementation manner of the present application, in the obtaining unit 11, it is used for: obtaining at least one fourth text feature and the second confidence level of the at least one fourth text feature; the second confidence level characterizes the accuracy of the semantic information carried by the fourth text feature in describing the first category corresponding to the fourth text feature; when determining n fourth text features with the highest second confidence level corresponding to each of the first categories from the at least one fourth text feature to obtain the at least one first text feature, the obtaining unit 11 is specifically used for:
[0461] Performing feature extraction processing on the at least one first training text through the text encoding module to obtain the text features of the at least one training text as the at least one fourth text feature; the at least one fourth text feature includes the fifth text feature;
[0462] Obtain the second confidence level of the fifth text feature according to the third similarity; the second confidence level of the fifth text feature is positively correlated with the third similarity.
[0463] Combined with any implementation manner of the present application, the training unit 15 is further used for:
[0464] Obtain the fourth video feature of the second training video; the fourth video feature is extracted from the second training video through a trained video feature extraction model;
[0465] Obtain a third loss according to the first difference between the third video feature and the fourth video feature; the third loss is positively correlated with the first difference;
[0466] Update the parameters of the video classification training network according to the second loss and the third loss.
[0467] Combined with any implementation manner of the present application, the training unit 15 is further used for:
[0468] Obtain the seventh text feature of the second training text; the seventh text feature is extracted from the second training text through a trained text feature extraction model;
[0469] Obtain a fourth loss according to the second difference between the fifth text feature and the seventh text feature; the fourth loss is positively correlated with the second difference;
[0470] Updating the parameters of the video classification training network according to the second loss and the third loss includes:
[0471] Updating the parameters of the video classification training network according to the second loss, the third loss, and the fourth loss.
[0472] Combined with any implementation manner of the present application, the training unit 15 is configured to:
[0473] Obtain a training data set, where the training data set is constructed based on at least two of the following training tasks: closed set, long-tailed distribution, few-shot, open set;
[0474] Sample the second training video and the at least two first training texts from the training data set.
[0475] In the embodiments of the present application, when the video classification device obtains a first video feature by performing feature extraction processing on the video to be processed, a first fusion feature is obtained by fusing the first video feature and at least one first text feature, enriching the information for determining the category of the video to be processed. Then, the video classification device classifies the video to be processed according to the first fusion feature to obtain a second category of the video to be processed, which can improve the accuracy of the second category.
[0476] Optionally, the acquisition unit 11 is a data interface, the first processing unit 12, the second processing unit 13, and the third processing unit 14 are all graphics processors, and the training unit 15 is a processor.
[0477] In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0478] Figure 4 This is a schematic hardware structure diagram of an electronic device provided in an embodiment of the present application. The electronic device 2 includes a processor 21 and a memory 22. Optionally, the electronic device 2 further includes an input device 23 and an output device 24. The processor 21, the memory 22, the input device 23, and the output device 24 are coupled through a connector, and the connector includes various interfaces, transmission lines, or buses, etc. The embodiments of the present application do not limit this. It should be understood that in various embodiments of the present application, coupling refers to a specific way of mutual connection, including direct connection or indirect connection through other devices. For example, it can be connected through various interfaces, transmission lines, buses, etc.
[0479] The processor 21 may be one or more graphics processing units (GPUs). When the processor 21 is a GPU, the GPU may be a single-core GPU or a multi-core GPU. Optionally, the processor 21 may be a processor group composed of multiple GPUs, and multiple processors are coupled to each other through one or more buses. Optionally, the processor may also be other types of processors, etc., which are not limited in the embodiments of the present application.
[0480] The memory 22 can be used to store computer program instructions and various computer program codes including the program codes for executing the solutions of the present application. Optionally, the memory includes but is not limited to random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), and this memory is used for relevant instructions and data.
[0481] The input device 23 is used to input data and / or signals, and the output device 24 is used to output data and / or signals. The input device 23 and the output device 24 may be independent devices or an integrated device.
[0482] It can be understood that in the embodiments of the present application, the memory 22 can not only be used to store relevant instructions, but also be used to store relevant data. For example, the memory 22 can be used to store the video to be processed and at least one first text feature obtained through the input device 23, or the memory 22 can also be used to store the second category of the video to be processed obtained through the processor 21, etc. The embodiments of the present application do not limit the specific data stored in this memory.
[0483] It can be understood that Figure 4 Only a simplified design of an electronic device is shown. In practical applications, the electronic device may also respectively include necessary other components, including but not limited to any number of input / output devices, processors, memories, etc., and all electronic devices that can implement the embodiments of the present application are within the protection scope of the present application.
[0484] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0485] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. Those skilled in the art can also clearly understand that each embodiment of this application has its own emphasis. For the convenience and brevity of description, the same or similar parts may not be elaborated in different embodiments. Therefore, the parts not described or not detailedly described in a certain embodiment can be referred to the descriptions of other embodiments.
[0486] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0487] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0488] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0489] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital versatile disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0490] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware with a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
Claims
1. A video classification method, characterized in that, the method includes: obtaining a video to be processed and at least one first text feature; the at least one first text feature carries semantic information for describing at least one first category; the number of the first text features is greater than 1, and the at least one first text feature includes a second text feature and a third text feature; performing feature extraction processing on the video to be processed to obtain a first video feature; obtaining a first similarity between the second text feature and the first video feature, and a second similarity between the third text feature and the first video feature; obtaining a first weight of the second text feature and a second weight of the third text feature according to the first similarity and the second similarity; when the first similarity is greater than the second similarity, the first weight is greater than the second weight; when the first similarity is equal to the second similarity, the first weight is equal to the second weight; fusing the first video feature and the at least one first text feature to obtain a first fusion feature; the fusing the first video feature and the at least one first text feature to obtain a first fusion feature includes: performing weighted fusion on the second text feature and the third text feature according to the first weight and the second weight to obtain a second fusion feature; fusing the second fusion feature and the first video feature to obtain the first fusion feature; classifying the video to be processed according to the first fusion feature to obtain a second category of the video to be processed.
2. The method according to claim 1, characterized in that, the classifying the video to be processed according to the first fusion feature to obtain a second category of the video to be processed includes: predicting a third category of the video to be processed and a first confidence level of the third category according to the first fusion feature; the third category belongs to the at least one first category; when the first confidence level is less than or equal to a confidence level threshold, determining that the second category of the video to be processed is a category other than the at least one first category; when the first confidence level is greater than the confidence level threshold, determining that the third category is the second category.
3. The method according to claim 1 or 2, characterized in that, the obtaining at least one first text feature includes: obtaining at least one fourth text feature and a second confidence level of the at least one fourth text feature; the second confidence level represents the accuracy of the semantic information carried by the fourth text feature in describing the first category corresponding to the fourth text feature; determining n fourth text features with the highest second confidence levels corresponding to each of the first categories from the at least one fourth text feature to obtain the at least one first text feature.
4. The method according to claim 1 or 2, characterized in that, the performing feature extraction processing on the video to be processed to obtain a first video feature includes: Performing feature extraction processing on at least one frame of the to-be-processed image in the to-be-processed video to obtain the frame features of the at least one frame of the to-be-processed image; Fusing the frame features of the at least one frame of the to-be-processed image and the timestamp information of the at least one frame of the to-be-processed image to obtain the first video feature.
5. The method according to claim 1 or 2, characterized in that, the video classification method is implemented through a video classification network, and the video classification network includes a video encoding module; the performing feature extraction processing on the to-be-processed video to obtain the first video feature includes: performing feature extraction processing on the to-be-processed video through the video encoding module to obtain the first video feature; the video classification method further includes the training process of the video classification network: obtaining a first training video; performing feature extraction processing on the first training video through the video encoding module to obtain a second video feature; fusing the second video feature and the at least one first text feature to obtain a third fusion feature; obtaining a fourth category of the first training video according to the third fusion feature; obtaining a first loss of the video classification network according to the first difference between the fourth category and the label of the first training video; updating the parameters of the video classification network according to the first loss to obtain the video classification network.
6. The method according to claim 5, characterized in that, the updating the parameters of the video classification network according to the first loss includes: updating the parameters of the video classification network other than the parameters of the video encoding module according to the first loss; the video encoding module is obtained by training a video classification training network, and the video classification training network includes the video encoding module; the video classification method further includes the training process of the video classification training network: obtaining a second training video and at least two first training texts; the labels of the at least two first training texts include at least two fifth categories, and the labels of the at least two first training texts include the at least one first category; the at least two first training texts include a second training text, and the fifth category of the second training text is the same as the sixth category of the second training video; performing feature extraction processing on the second training video through the video encoding module to obtain a third video feature; obtaining a second loss of the video classification training network according to the third similarity between the third video feature and the second training text; the second loss is negatively correlated with the third similarity; updating the parameters of the video classification training network according to the second loss to obtain the video classification training network.
7. The method according to claim 6, characterized in that, the at least two first training texts further include a third training text, and the fifth category described by the third training text is different from the sixth category; before the obtaining the second loss of the video classification training network according to the third similarity between the third video feature and the second training text, the method further includes: Determine the fourth similarity between the third video feature and the third training text; The obtaining the second loss of the video classification training network according to the third similarity between the third video feature and the second training text includes: Obtaining the second loss of the video classification training network according to the third similarity and the fourth similarity; the second loss is positively correlated with the fourth similarity.
8. The method according to claim 6, wherein, the video classification training network further includes a text encoding module; Before obtaining the second loss of the video classification training network according to the third similarity between the third video feature and the second training text, the method further includes: Performing feature extraction processing on the second training text through the text encoding module to obtain a fifth text feature of the second training text; Determining the similarity between the fifth text feature and the third video feature as the third similarity between the third video feature and the second training text.
9. The method according to claim 8, wherein, the obtaining at least one first text feature includes: Obtaining at least one fourth text feature and a second confidence level of the at least one fourth text feature; the second confidence level represents the accuracy of the semantic information carried by the fourth text feature in describing the first category corresponding to the fourth text feature; Determining, from the at least one fourth text feature, n fourth text features with the highest second confidence level corresponding to each of the first categories to obtain the at least one first text feature; wherein, the obtaining at least one fourth text feature includes: Performing feature extraction processing on the at least one first training text through the text encoding module to obtain text features of the at least one training text as the at least one fourth text feature; the at least one fourth text feature includes the fifth text feature; the second confidence level of the at least one fourth text feature includes the second confidence level of the fifth text feature, and the obtaining the second confidence level of the at least one fourth text feature includes: Obtaining the second confidence level of the fifth text feature according to the third similarity; the second confidence level of the fifth text feature is positively correlated with the third similarity.
10. The method according to claim 9, wherein, Before updating the parameters of the video classification training network according to the second loss, the method further includes: Obtaining a fourth video feature of the second training video; the fourth video feature is extracted from the second training video through a trained video feature extraction model; Obtaining a third loss according to a first difference between the third video feature and the fourth video feature; the third loss is positively correlated with the first difference; The updating the parameters of the video classification training network according to the second loss includes: Updating the parameters of the video classification training network according to the second loss and the third loss.
11. The method according to claim 10, wherein, Before updating the parameters of the video classification training network according to the second loss and the third loss, the method further includes: Obtaining a seventh text feature of the second training text; the seventh text feature is extracted from the second training text by a trained text feature extraction model; Obtaining a fourth loss according to a second difference between the fifth text feature and the seventh text feature; the fourth loss is positively correlated with the second difference; The updating the parameters of the video classification training network according to the second loss and the third loss includes: Updating the parameters of the video classification training network according to the second loss, the third loss and the fourth loss.
12. The method according to claim 6, wherein, the obtaining the second training video and at least two first training texts includes: Obtaining a training data set, the training data set being constructed according to at least two of the following training tasks: closed set, long-tailed distribution, few-shot, open set; Sampling the second training video and the at least two first training texts from the training data set.
13. A video classification device, wherein, the device includes: An obtaining unit, configured to obtain a video to be processed and at least one first text feature; the at least one first text feature carries semantic information for describing at least one first category; the number of the first text features is greater than 1, and the at least one first text feature includes a second text feature and a third text feature; A first processing unit, configured to perform feature extraction processing on the video to be processed to obtain a first video feature; The obtaining unit is further configured to obtain a first similarity between the second text feature and the first video feature, and a second similarity between the third text feature and the first video feature; A second processing unit, configured to obtain a first weight of the second text feature and a second weight of the third text feature according to the first similarity and the second similarity; in a case where the first similarity is greater than the second similarity, the first weight is greater than the second weight; in a case where the first similarity is equal to the second similarity, the first weight is equal to the second weight; The second processing unit is configured to fuse the first video feature and the at least one first text feature to obtain a first fusion feature; the fusing the first video feature and the at least one first text feature to obtain a first fusion feature includes: performing weighted fusion on the second text feature and the third text feature according to the first weight and the second weight to obtain a second fusion feature; fusing the second fusion feature and the first video feature to obtain the first fusion feature; A third processing unit, configured to classify the video to be processed according to the first fusion feature to obtain a second category of the video to be processed.
14. An electronic device, wherein, includes: A processor and a memory, the memory being configured to store computer program code, the computer program code including computer instructions, and in the case where the processor executes the computer instructions, the electronic device executes the method according to any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, the computer program including program instructions, and in the case where the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Video dynamic thumbnail generation method and model training method and device
CN109885723A
Video processing method, machine learning model training method and related device and equipment
CN114419515A