Video Tag Recognition Method, Device, Electronic Device and Computer Readable Medium

By matching the similarity between the target feature vectors of the videos to be identified with multiple cluster centers, the video tags are automatically determined, which solves the problems of low manual review efficiency and insufficient accuracy of machine learning algorithms in the prior art, and achieves efficient and accurate video tag recognition.

CN114637889BActive Publication Date: 2025-06-27BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210270414.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-18
Publication Date
2025-06-27
Estimated Expiration
2042-03-18

AI Technical Summary

Technical Problem

In the video tag identification, the existing technology has problems such as low manual review efficiency, the accuracy of machine learning algorithms is affected by the training data scale and sample distribution, and the generalization ability is insufficient.

Method used

By acquiring the target feature vector of the video to be identified and a preset multiple first clustering centers, the similarity between the video to be identified and each clustering center is determined, and the tag of the clustering center with the greatest similarity is used as the tag of the video to be identified.

Benefits of technology

It realizes accurate identification of the tags of the video to be identified, and obtains label predictions consistent with the distribution of the target video tags, saving manpower and reducing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114637889B_ABST
    Figure CN114637889B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a video tag recognition method, apparatus, electronic device, and computer-readable medium, which relate to the field of big data technology. The method includes: obtaining a video to be recognized and a plurality of preset first clustering centers, where the plurality of first clustering centers are obtained by clustering the target feature vectors of a plurality of target videos; obtaining the target feature vector of the video to be recognized, and determining the similarity between the video to be recognized and each of the first clustering centers according to the target feature vector of the video to be recognized; determining, from the plurality of first clustering centers, a target clustering center with the maximum similarity to the video to be recognized, and using the tag corresponding to the target clustering center as the tag of the video to be recognized. This method can accurately recognize the tags of videos with different durations, automatically mark the videos, save manpower, and effectively reduce costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and in particular, to a method, device, electronic device and computer-readable medium for video tag recognition. Background Art

[0002] With the development of mobile Internet and streaming media technology, the production of video products has become more and more convenient, and the number of videos has increased rapidly. In order to improve the user experience, when a video content provider launches a new video, it is necessary to mark the tags of the video, and the tags are used to reflect the classification, theme, flavor or picture style of the video. For example, the classification of videos can include TV dramas, movies, variety shows, documentaries, news, etc. The themes of videos can include: emotional, popular science, entertainment, etc. Flavor generally refers to the personality and individuality of things, which is a word borrowed from music terms to express different contents of multimedia. For example, brand flavor refers to the unique style of a brand in the market. The flavor of a video refers to the performance presented by the video and the perceived image reflected by the video.

[0003] However, the traditional method is to mark new videos manually, or to automatically mark new videos through machine learning algorithms. For example, given a training data set and some labeled class set samples for classification learning, a classifier is learned, and the classifier is used to mark new videos. However, in the face of a large number of newly added videos, manual review requires a large amount of manpower and is inefficient. In machine learning algorithms, the accuracy of the classifier is affected by many factors. For example, the scale of the training data set and the distribution difference of video samples of different classes in the training data set (i.e., sample imbalance) will affect the accuracy of the classifier; due to the development of the business, the distribution of video tags will also change, resulting in insufficient generalization ability of the classifier, and the tags predicted by the classifier cannot meet the real distribution requirements. Summary of the Invention

[0004] To solve the above technical problems or at least partially solve the above technical problems, embodiments of the present invention provide a method, device, electronic device and computer-readable storage medium for video tag recognition.

[0005] In the first aspect of the implementation of the present invention, a video tag recognition method is provided, including: obtaining a video to be recognized and the target feature vector of the video to be recognized, and obtaining a plurality of preset first clustering centers, where the plurality of first clustering centers are obtained by clustering the target feature vectors of a plurality of target videos, and each of the first clustering centers has a corresponding tag; determining the similarity between the video to be recognized and each of the first clustering centers according to the target feature vector of the video to be recognized; determining the target clustering center with the largest similarity to the video to be recognized from the plurality of first clustering centers, and taking the tag corresponding to the target clustering center as the tag of the video to be recognized.

[0006] Optionally, the process of obtaining the plurality of first clustering centers by clustering the target feature vectors of a plurality of target videos includes: sampling the plurality of target videos to obtain a plurality of sampled videos; performing a clustering operation on the target feature vectors of the plurality of sampled videos according to a preset clustering rule to obtain a plurality of second clustering centers; sampling the remaining target videos in the plurality of target videos except the sampled videos to obtain a plurality of new sampled videos, and iteratively updating the plurality of second clustering centers according to the target feature vectors of the plurality of new sampled videos until a preset stop condition is reached, and taking the second clustering centers after the last update as the first clustering centers.

[0007] Optionally, the method further includes: after obtaining the plurality of first clustering centers, for each first clustering center, determining the tag of the first clustering center according to the tags of the target videos within the cluster where the first clustering center is located.

[0008] Optionally, the process of obtaining the target feature vector of each target video includes: using a pre-constructed 3D convolutional neural network model to obtain the initial feature vectors of all target videos, where the dimensions of the initial feature vectors of target videos with different durations are different; performing weighted aggregation processing on the initial feature vector of each target video according to a preset weighted aggregation rule to obtain the target feature vector of each target video; the dimensions of the target feature vectors of each target video are the same.

[0009] Obtaining the target feature vector of the video to be recognized includes: using the pre-constructed 3D convolutional neural network model to obtain the initial feature vector of the video to be recognized, and performing weighted aggregation processing on the initial feature vector of the video to be recognized according to the weighted aggregation rule to obtain the target feature vector of the video to be recognized, and the target feature vector of the video to be recognized has the same dimension as the target feature vector of each target video.

[0010] Optionally, extracting the initial feature vector of each target video by using the pre-built 3D convolutional neural network model includes: for each target video, segmenting the target video according to a preset segmentation rule to obtain a plurality of sub-samples; the durations of the sub-samples obtained after segmenting all target videos are the same; for each sub-sample, inputting the sub-sample into the pre-built 3D convolutional neural network model, and taking the output of the pre-built 3D convolutional neural network model as the initial feature vector of the sub-sample, the initial feature vector of the sub-sample is a feature vector of W*C dimensions, and W and C are integers greater than 1 respectively; concatenating the initial feature vectors W*C of the plurality of sub-samples to obtain the initial feature vector of the target video, the initial feature vector of the target video is a feature vector of H*W*C dimensions, where H represents the number of sub-samples obtained after segmenting the target video;

[0011] Performing weighted aggregation processing on the initial feature vector of each target video according to a preset weighted aggregation rule to obtain the target feature vector of each target video includes: calculating the variance of the aggregation value of the channel feature maps of each channel according to the initial feature vectors H*W*C of all target videos, where the channel is the C dimensions of the initial feature vector, and the channel feature map is a two-dimensional matrix composed of the H dimensions and the W dimensions in the initial feature vector; sorting the variances of the aggregation values of the channel feature maps in descending order, and selecting the first N channels corresponding to the variances as the target channels, where N is an integer greater than or equal to 1; determining the normalized weight of the channel feature map of each target channel according to the feature map activation value in the channel feature map of the target channel and the sum of the feature map activation values of the channel feature maps of all target channels; determining the weighted sum of the channel feature map of the target channel according to the normalized weight of the channel feature map of the target channel and the initial feature vector of the target video; concatenating the weighted sums of the channel feature maps of N target channels to obtain the target feature vector of the target video;

[0012] Obtaining the initial feature vector of the video to be recognized by using the pre-built 3D convolutional neural network model includes: segmenting the video to be recognized according to a preset segmentation rule to obtain a plurality of sub-samples; the durations of the sub-samples obtained after segmenting the video to be recognized are the same; for each sub-sample, inputting the sub-sample into the pre-built 3D convolutional neural network model, and taking the output of the pre-built 3D convolutional neural network model as the initial feature vector of the sub-sample, the initial feature vector of the sub-sample is a feature vector of W*C dimensions, and W and C are integers greater than 1 respectively; concatenating the initial feature vectors W*C of the plurality of sub-samples to obtain the initial feature vector of the target video, the initial feature vector of the target video is a feature vector of K*W*C dimensions, where K represents the number of sub-samples corresponding to the video to be recognized;

[0013] Performing weighted aggregation processing on the initial feature vector of the video to be recognized according to the weighted aggregation rule to obtain the target feature vector of the video to be recognized includes: determining the weighted sum of the channel feature maps of the N target channels according to the initial feature vector K*W*C and the normalized weights of the channel feature maps of the pre-determined N target channels; splicing the weighted sums of the channel feature maps of the N target channels to obtain the target feature vector of the video to be recognized.

[0014] Optionally, the initial feature vector includes an initial video frame feature vector and an initial audio feature vector;

[0015] The process of obtaining the initial feature vector of each target video or video to be recognized includes: for each target video or video to be recognized, obtaining the initial video frame feature vector and the initial audio feature vector of each sub-target video or sub-video to be recognized; splicing the initial video frame feature vectors and the initial audio feature vectors of the multiple sub-target videos to obtain the initial video frame feature vector and the initial audio feature vector of the target video, or splicing the initial video frame feature vectors and the initial audio feature vectors of the multiple sub-videos to be recognized to obtain the initial video frame feature vector and the initial audio feature vector of the video to be recognized;

[0016] The process of obtaining the target feature vector of each target video or video to be recognized includes: performing weighted aggregation processing on the initial video frame feature vector of each target video or video to be recognized according to a preset weighted aggregation rule to obtain the target video frame feature vector of each target video or video to be recognized; performing weighted aggregation processing on the initial audio feature vector of each target video or video to be recognized according to a preset weighted aggregation rule to obtain the target audio feature vector of each target video or video to be recognized, where the target video frame feature vector of the video to be recognized has the same dimension as the target video frame feature vector of each target video, and the target audio feature vector of the video to be recognized has the same dimension as the target audio feature vector of each target video; fusing the target video frame feature vector and the target audio feature vector of each target video or video to be recognized to obtain the target audio-visual feature vector of each target video or video to be recognized;

[0017] The process of clustering the target feature vectors of multiple target videos to obtain the first cluster center includes: performing a clustering operation on the target audio-visual feature vectors of the multiple target videos to obtain the multiple first cluster centers;

[0018] Determining the similarity between the video to be recognized and each of the first cluster centers according to the target feature vector of the video to be recognized includes: determining the similarity between the video to be recognized and each of the first cluster centers according to the target audio-visual feature vector of the video to be recognized.

[0019] In a second aspect of the embodiments of the present invention, a video tag recognition device is provided, including: an acquisition module, configured to acquire a video to be recognized and a target feature vector of the video to be recognized, and acquire a plurality of preset first clustering centers, where the plurality of first clustering centers are obtained by clustering the target feature vectors of a plurality of target videos, and each of the first clustering centers has a corresponding tag; a calculation module, configured to determine a similarity between the video to be recognized and each of the first clustering centers according to the target feature vector of the video to be recognized; and a recognition module, configured to determine, from the plurality of first clustering centers, a target clustering center with the maximum similarity to the video to be recognized, and use the tag corresponding to the target clustering center as the tag of the video to be recognized.

[0020] In a third aspect of the embodiments of the present invention, a computer-readable storage medium is further provided. Instructions are stored in the computer-readable storage medium, and when the instructions run on a computer, the computer is enabled to execute any one of the above-mentioned video tag recognition methods.

[0021] In a fourth aspect of the embodiments of the present invention, a computer program product including instructions is further provided. When the computer program product runs on a computer, the computer is enabled to execute any one of the above-mentioned video tag recognition methods.

[0022] The video tag recognition method provided by the embodiments of the present invention determines the similarity between the video to be recognized and a plurality of preset first clustering centers through the target feature vector of the video to be recognized. The plurality of first clustering centers are obtained by clustering the target feature vectors of a plurality of target videos, and each first clustering center has a corresponding tag. Then, the tag of the target clustering center with the highest similarity to the video to be recognized is used as the tag of the video to be recognized, which can accurately recognize the tag of the video to be recognized, obtain a tag prediction consistent with the tag distribution of the target video, can automatically label the video to be recognized, save manpower, and effectively reduce costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.

[0024] Figure 1 Schematically shows a flowchart of the video tag recognition method according to the embodiments of the present invention;

[0025] Figure 2 Schematically shows a sub-flowchart of the video tag recognition method according to the embodiments of the present invention;

[0026] Figure 3 Schematically shows a structural diagram of the video tag recognition device according to the embodiments of the present invention;

[0027] Figure 4 Schematically shows the result schematic diagram of the electronic device according to an embodiment of the present invention. Detailed implementation manners

[0028] Next, the technical solutions in the embodiments of the present invention will be described with reference to the accompanying drawings in the embodiments of the present invention.

[0029] Figure 1 Schematically shows the flowchart of the video tag recognition method according to an embodiment of the present invention, as Figure 1 shown, the method includes:

[0030] Step 101: Obtain the video to be recognized and the target feature vector of the video to be recognized, and obtain a plurality of preset first clustering centers, where the plurality of first clustering centers are obtained by clustering the target feature vectors of a plurality of target videos, and each first clustering center has a corresponding label.

[0031] The video to be recognized may be a TV drama, a movie, a variety show, a short video, etc. In this step, the target feature vector of the video to be recognized can be obtained through a pre-trained feature acquisition model, and the pre-trained feature acquisition model may be a convolutional neural network model, which is not specifically limited.

[0032] Clustering is a process of dividing a set of physical or abstract objects into multiple classes composed of similar objects. The clusters generated by clustering are a set of data objects, and these objects are similar to each other within the same cluster and different from the objects in other clusters. The center within the cluster generated by clustering is the clustering center, and it is determined whether the object belongs to the cluster by calculating the distance between the object and it. In this step, the clustering center is obtained by clustering the target feature vectors of a plurality of target videos, so the clustering center is also a vector. Exemplarily, the target feature vectors of a plurality of target videos can be clustered by the K-Means (K-means) clustering algorithm, the mean shift clustering or the density-based clustering algorithm.

[0033] The target video can refer to a video available on the client or website of the video provider. Each target video has a corresponding tag, which can describe the type, theme, flavor, or style of the video, etc. For example, the types of videos can include TV dramas, movies, variety shows, documentaries, news, short videos, etc. The themes of the videos can include emotional, popular science, entertainment, etc. Flavor generally refers to the personality and individuality of things, which is a term borrowed from music to express different contents of multimedia. For example, brand flavor refers to the unique style of a brand in the market. The flavor of a video refers to what kind of performance the video presents, the perceptual image reflected by the video, and the flavor of the video can include nostalgic and retro, literary, etc. The styles of the videos can include positive energy, comedy, funny, etc. This tag can be determined by manual marking. The target feature vector can be obtained through the above-mentioned pre-trained feature acquisition model, for example, it can be obtained through a pre-trained neural network. This target feature vector can characterize the category, theme, flavor, or style of the target video.

[0034] In this step, the target videos within a period of time (such as within a month) can be obtained, feature extraction can be performed to obtain the target feature vectors of the target videos within this time period, and then the target feature vectors can be clustered to achieve clustering of the target videos, so as to cluster the videos with similar types, themes, flavors, or styles into one cluster, and divide the videos with dissimilar types, themes, flavors, or styles into different clusters.

[0035] Step 102: Determine the similarity between the video to be recognized and each of the first clustering centers according to the target feature vector of the video to be recognized.

[0036] The similarity between the video to be recognized and the first clustering center can be determined by the distance between the video to be recognized and this first clustering center. The smaller the distance, the greater the similarity. The distance between the video to be recognized and the first clustering center can be the Euclidean distance, Hamming distance, cosine similarity, etc. The Euclidean distance is a commonly used distance definition, which refers to the actual distance between two points in an m-dimensional space, or the natural length of a vector (i.e., the distance from this point to the origin). The Euclidean distance in two-dimensional and three-dimensional spaces is the actual distance between two points. The Hamming distance is the number of different values between two vectors. The cosine similarity refers to the cosine of the angle between two vectors.

[0037] Step 103: Determine the target clustering center with the greatest similarity to the video to be recognized from the multiple first clustering centers, and use the tag corresponding to the target clustering center as the tag of the video to be recognized.

[0038] In this embodiment, each first clustering center has a corresponding label, which is used to characterize the type, theme tone or style of the target video within the cluster where the first clustering center is located. The label of the first clustering center can be determined by the labels of the target videos within the cluster where the first clustering center is located.

[0039] The greater the similarity between the video to be recognized and the first clustering center, the more similar the video to be recognized is to the target videos within the cluster where the first clustering center is located. Therefore, the target clustering center with the greatest similarity is selected from multiple first clustering centers, and the label corresponding to the target clustering center is used as the label of the video to be recognized.

[0040] The video label recognition method according to the embodiment of the present invention determines the similarity between the video to be recognized and a plurality of preset first clustering centers through the target feature vector of the video to be recognized. The plurality of first clustering centers are obtained by clustering the target feature vectors of a plurality of target videos, and each first clustering center has a corresponding label. Then, the label of the target clustering center with the highest similarity to the video to be recognized is used as the label of the video to be recognized, which can accurately recognize the label of the video to be recognized, obtain a label prediction consistent with the label distribution of the target video, can automatically label the video to be recognized, save manpower, and effectively reduce costs.

[0041] In an alternative embodiment, since the volume of the target feature vectors of all target videos is too large to perform clustering operations at one time, a batch iterative calculation method can be adopted to determine the clustering centers. Instead of using all target videos in each iteration, equal sampling is performed on all target videos, and the sampled target videos are used to iteratively update the clustering centers. Based on this technical concept, the process of clustering the target feature vectors of a plurality of target videos to obtain a plurality of first clustering centers may include:

[0042] Sample the plurality of target videos to obtain a plurality of sampled videos. For example, sample Z sampled videos from M target videos, where M is much larger than Z.

[0043] According to the preset clustering rules, perform clustering operations on the target feature vectors of the plurality of sampled videos to obtain a plurality of second clustering centers. For example, use the K-Means algorithm to construct a model with K centroids, and the K centroids are the second clustering centers.

[0044] Sampling is performed on the remaining target videos among the multiple target videos except the sampled video to obtain multiple new sampled videos. According to the target feature vectors of the multiple new sampled videos, the multiple second clustering centers are updated until a preset stop condition is reached, and the second clustering center after the last update is used as the first clustering center. For example, continue to sample Z new sampled videos from the remaining target videos, and assign the new sampled videos to the nearest clustering center. Use the target feature vectors within the cluster where the clustering center is located to update the clustering center. Repeat the above steps until the clustering center is stable (the new clustering center is equal to or less than the specified threshold) or the number of iterations is reached, and the algorithm ends. The final clustering center is used as the final first clustering center. In each iteration of this algorithm, for each remaining target video, according to the distance between the remaining target video and each clustering center, the remaining target video is assigned to the nearest clustering center. After all target videos are examined, the iterative operation is completed, and new clustering centers are calculated. If the values of the clustering centers do not change before and after an iteration, it means that the algorithm has converged. After the algorithm converges, K first clustering centers are obtained.

[0045] In an alternative embodiment, the method further includes: after obtaining the multiple first clustering centers, for each first clustering center, determining the label of the first clustering center according to the labels of the target videos within the cluster where the first clustering center is located. For example, the labels of all target videos within the cluster where the first clustering center is located can be counted, and the label with the largest number is used as the label of the first clustering center, or the labels are sorted in descending order according to the number of labels, and the top P labels are used as the labels of the clustering center. For example, on a video playback platform, manually mark whether a video is suitable for online playback. Videos suitable for playback are marked as good, and videos not suitable for playback are marked as bad. After obtaining multiple first clustering centers, count the proportion of the labels good / bad of the target videos within the cluster where each first clustering center is located, and determine the label of each first clustering center according to the proportion of the labels good / bad. Sort in descending order according to the proportion of good / bad, set the labels of the first N first clustering centers to good, and set the tonal labels of the remaining first clustering centers to bad. For the video to be recognized, determine the label of the video to be recognized according to the label of the first clustering center most similar to the video to be recognized. If the tonal label of the first clustering center most similar to the video to be recognized is good, determine that the tonal label of the video to be recognized is good. If the tonal label of the first clustering center most similar to the video to be recognized is bad, determine that the tonal label of the video to be recognized is bad. If the tonal label of the video to be recognized is good, push the video to be recognized to the fast channel to facilitate the rapid listing of the video to be recognized.

[0046] In an alternative embodiment, the process of obtaining the target feature vector of each target video includes:

[0047] First, obtain the initial feature vectors of all target videos: Use a pre-constructed 3D convolutional neural network model to obtain the initial feature vectors of all target videos. Among them, the dimensions of the initial feature vectors of target videos with different durations are different.

[0048] Among them, the initial feature vector can be obtained through a pre-trained deep learning backbone network. The initial feature vector can be an initial video frame feature vector, an initial audio feature vector, or can also include both an initial video frame feature vector and an initial audio feature vector. The initial video frame feature vector can represent the video frame features of the video, and the initial audio feature vector can represent the audio features of the video. In this step, the initial feature vector of each target video can be extracted through a pre-constructed 3D convolutional neural network model. For example, if the initial feature vector is a video frame feature vector, it can be extracted through the pre-constructed 3D convolutional neural network model I3D. If the initial feature vector is an audio feature vector, it is extracted through another pre-constructed 3D convolutional neural network model VGG, without specific limitation. Among them, the convolutional kernel of the 3D convolutional neural network model is 3D. Using the 3D convolutional neural network model can better capture the temporal and spatial feature information in the video, that is, extract the feature information of the target video in the temporal dimension and the spatial dimension. The 3D convolutional neural network model stacks multiple consecutive frames to form a cube, and then applies a 3D convolutional kernel in the cube. In this structure, each feature map in the convolutional layer is connected to multiple adjacent consecutive frames in the previous layer, so the feature information of the target video in the temporal dimension can be captured. For target videos with different durations, the number of video frames they contain is different. Therefore, for target videos with different durations, the initial feature vectors extracted using the 3D convolutional neural network model are different in the temporal dimension, that is, the dimensions of the initial feature vectors of target videos with different durations are different.

[0049] Second, according to a preset weighted aggregation rule, perform weighted aggregation processing on the initial feature vector of each target video to obtain the target feature vector of each target video; the dimensions of the target feature vectors of each target video are the same. The preset weighted aggregation rule can be the PWA algorithm (Part-based Weighting Aggregation, partial weighted aggregation operation, also known as part-based weighted aggregation). Through the PWA algorithm, weighted aggregation processing is performed on the initial feature vector to obtain the target feature vector, and the dimensions of the target feature vectors of all target videos are the same. Through this step, the initial feature vectors of target videos with different durations can be unified to the same dimension, which is convenient for subsequent clustering operations.

[0050] In order to facilitate the recognition of the video to be recognized and improve the recognition accuracy, in this embodiment, the process of obtaining the target feature vector of the video to be recognized is the same as the process of obtaining the target feature vector of the target video, that is, the process of obtaining the target feature vector of the video to be recognized includes: using the pre-constructed 3D convolutional neural network model to obtain the initial feature vector of the video to be recognized, and performing weighted aggregation processing on the initial feature vector of the video to be recognized according to the weighted aggregation rule to obtain the target feature vector of the video to be recognized. The target feature vector of the video to be recognized has the same dimension as the target feature vector of each target video.

[0051] In an alternative embodiment, the process of using the pre-constructed 3D convolutional neural network model to obtain the initial feature vectors of all target videos includes:

[0052] For each target video, segment the target video according to a preset segmentation rule to obtain a plurality of sub-samples; the durations of the sub-samples obtained after segmentation of all target videos are the same;

[0053] For each sub-sample, input the sub-sample into the pre-constructed 3D convolutional neural network model, and use the output of the pre-constructed 3D convolutional neural network model as the initial feature vector of the sub-sample. The initial feature vector of the sub-sample is a feature vector of W*C dimensions, where W and C are integers greater than 1;

[0054] Concatenate the initial feature vectors W*C of the plurality of sub-samples to obtain the initial feature vector of the target video. The initial feature vector of the target video is a feature vector of H*W*C dimensions, where H represents the number of sub-samples obtained after segmentation of the target video.

[0055] In this embodiment, the preset segmentation rule is used to segment the target video, and a target video is segmented into multiple sub-samples. The durations of the sub-samples obtained after segmenting all target videos are the same, that is, the number of video frames included in the sub-samples obtained after segmenting all target videos is the same, and the durations of the audio included in the sub-samples obtained after segmenting all target videos are the same. Since the durations of the sub-samples are the same, the number of sub-samples obtained after segmenting target videos with different durations is different. For example, the duration of target video A is 1920 s, and the duration of target video B is 3600 s. Assuming that the duration of the sub-sample is 12.8 s, the number of sub-samples obtained after segmenting target video A is 150, and the number of sub-samples of target video B is 281. In this embodiment, the sub-samples are obtained by segmenting the target video in the time dimension, and the number of sub-samples is a feature of the target video in the time dimension. The sub-samples are input into a pre-constructed 3D convolutional neural network model to obtain the initial feature vector W*C of the sub-samples. The features output by the 3D convolutional neural network model are two-dimensional features. The initial feature vectors of multiple sub-samples are concatenated to obtain the initial feature vector H*W*C of the target video. As an example, a target video with a duration of t can be divided into multiple sub-samples, and each sub-sample has 64 frame images and 12.8 s of audio. If the initial feature vector is the picture feature vector of the target video, the 64-frame image sequence is input into the pre-constructed deep learning backbone network 3D-CNN to extract a 6*1024-dimensional picture feature vector. If the initial feature vector is the audio feature vector of the target video, the 12.8 s of audio is passed through another deep learning backbone network VGG to extract an 8*128-dimensional audio feature vector. Finally, each target video obtains a H*6*1024-dimensional picture feature vector and / or a H*8*128-dimensional audio feature vector, where H represents the number of sub-samples.

[0056] After obtaining the initial feature vector H*W*C of all target videos, the initial feature vector can be weighted and aggregated through the PWA algorithm (Part-based Weighting Aggregation, also known as part-based weighted aggregation) to obtain the target feature vector of the target video. The PWA algorithm selects representative and discriminative channels from all channels, and measures whether a channel is discriminative by the variance of the aggregation value of the feature map of each channel. The larger the variance, the greater the discriminability. Then, based on the channel feature maps of the N channels with larger discriminability, the target feature vector of the target video is determined. As Figure 2 shown, according to the preset weighted aggregation rule, the process of weighted aggregation processing of the initial feature vector of each target video to obtain the target feature vector of each target video includes:

[0057] Step 201: Calculate the variance of the aggregation values of the channel feature maps for each channel based on the initial feature vectors of all target videos. The channels are the C dimensions of the initial feature vectors, and the channel feature maps are two-dimensional matrices composed of the H and W dimensions in the initial feature vectors. Here, one channel is for detecting a certain feature, and the strength of a value at a certain position in the channel reflects the strength of the current feature. For example, an image may only contain three channels, which contain information about how much red, green, or blue each pixel in the image has. Mapping this concept to convolution, RGB data with three channels is obtained. In a convolutional neural network model, different filters are used for each channel, and the filters obtain different information from each channel. The convolutional layer in a 3D convolutional neural network model can interact between channels and then generate new channels in the next convolutional layer.

[0058] In this step, the variance of the aggregation values of the channel feature maps of all target videos is statistically calculated through the initial feature vectors of all target videos, so as to select discriminative channel feature maps. The larger the variance, the greater the discriminability of the channel feature map. The method for calculating the variance is as follows:

[0059] Perform sum pooling operation on D initial feature vectors of H*W*C along the C dimension, which is equivalent to adding up the feature values of the H and W dimensions. Here, D represents the number of initial feature vectors, that is, the number of training samples in the training sample set. For the i-th channel feature map of the m-th initial feature vector, calculate the aggregation value of this channel feature map according to the following formula:

[0060]

[0061] where, g m,k represents the aggregation value of the i-th channel feature map of the m-th image, and f i (x, y) represents the element in the channel feature map of the i-th channel.

[0062] For each channel feature map, calculate the aggregation value according to the above formula to obtain the aggregation value sequence G = {g1, g2... g D} of all channel feature maps, and then calculate the variance according to the following formula:

[0063]

[0064] where, represents the average value, and V i represents the variance of the i-th channel feature map.

[0065] Finally, obtain the variance sequence V = {V1, V2... V c}.

[0066] Step 202: Sort the variances of the aggregation values of the channel feature maps in descending order, and select the channels corresponding to the top N variances as the target channels, where N is an integer greater than or equal to 1.

[0067] Sort all the channel feature maps in descending order of variance, and select the top N, for example, the channels corresponding to the top 10 variances as the target channels. For example, the default channel indices are [0, 1, 2... 1023], and the channel indices after sorting in descending order of variance are [5, 109, 233, 17,......, 89, 10, 602, 45], and then select the top 10 channels as the target channels.

[0068] Step 203: Determine the normalized weight of the channel feature map of each target channel according to the feature map activation value in the channel feature map of the target channel and the sum of the feature map activation values of the channel feature maps of all target channels. In this step, the normalized weight of each H*W feature in the C dimension is calculated, and this normalized weight is calculated from the feature response of the target channel, as shown in the following formula:

[0069]

[0070] where, v n (x, y) represents the feature map activation value of the nth target channel among the top N target channels, and α and β are power transformation exponents respectively. As an example, the value of α is 2 and the value of β is also 2.

[0071] Step 204: Determine the weighted sum of the channel feature map of the target channel according to the normalized weight of the channel feature map of the target channel and the initial feature vector of the target video;

[0072] Step 205: Concatenate the weighted sums of the channel feature maps of N target channels to obtain the target feature vector of the target video.

[0073] where, the weighted sum of the channel feature maps of each target channel is calculated by using the normalized weights of each target channel, and its calculation formula is shown in the following formula:

[0074]

[0075] where, ψ n (I) represents the weighted sum of the channel feature map of the nth target channel.

[0076] In this embodiment, the channel feature maps of the selected discriminative and representative target channels are weighted and aggregated to obtain the target feature vector of the target video. The dimensions of the target feature vectors of each target video are the same, which not only facilitates subsequent clustering operations, but also improves accuracy and robustness, and can identify the labels of videos of any duration.

[0077] In the embodiment of the present invention, the process of obtaining the target feature vector of the video to be recognized is the same as the process of obtaining the target feature vector of the target video, that is, the process of obtaining the target feature vector of the video to be recognized includes: determining the weighted sum of the channel feature maps of the N target channels according to the initial feature vector K*W*C and the normalized weights of the channel feature maps of the pre-determined N target channels; concatenating the weighted sums of the channel feature maps of the N target channels to obtain the target feature vector of the video to be recognized. For the specific details not described here, please refer to the description of obtaining the target feature vector of each target video by weighted aggregation of the initial feature vector of each target video.

[0078] In an alternative embodiment, the initial feature vector of the target video includes the initial picture feature vector and the initial audio feature vector of the target video. The initial feature vector of the video to be recognized includes the initial picture feature vector and the initial audio feature vector of the video to be recognized.

[0079] The process of obtaining the initial feature vector of each target video or video to be recognized includes:

[0080] For each target video or video to be recognized, obtain the initial picture feature vector and the initial audio feature vector of each sub-target video or sub-video to be recognized;

[0081] Concatenate the initial picture feature vectors and the initial audio feature vectors of the multiple sub-target videos to obtain the initial picture feature vector and the initial audio feature vector of the target video, or concatenate the initial picture feature vectors and the initial audio feature vectors of the multiple sub-videos to be recognized to obtain the initial picture feature vector and the initial audio feature vector of the video to be recognized.

[0082] In the implementation of the present invention, the preset segmentation rule is used to segment the target video or the video to be recognized, segmenting each target video into multiple sub-target videos or segmenting the video to be recognized into multiple sub-videos to be recognized. The durations of the sub-target videos obtained after segmenting all target videos are the same, and the durations of the sub-samples obtained after segmenting the video to be recognized are the same as the durations of the sub-target videos. That is, a target video or a video to be recognized is segmented into multiple video segments, and each video segment contains the same number of video frames and audio of the same duration. Taking the target video as an example, a target video with a total duration of t can be divided into T sub-target videos, and each sub-target video has 64 frame images and 12.8 s of audio. Then, the 64-frame image sequence is input into the pre-constructed deep learning backbone network I3D to extract an initial frame feature vector of 6 * 1024 dimensions, and the 12.8 s of audio is passed through another deep learning backbone network VGG to extract an initial audio feature vector of 8 * 128 dimensions. Then, the initial frame feature vectors and the initial audio feature vectors of the T sub-target videos are fused, and finally, an initial frame feature vector of T * 6 * 1024 dimensions and an initial audio feature vector of T * 8 * 128 dimensions are obtained.

[0083] In the embodiment of the present invention, after obtaining the initial frame feature vector and the initial audio feature vector of each target video, it further includes: according to the preset weighted aggregation rule, performing weighted aggregation processing on the initial frame feature vector of each target video to obtain the target frame feature vector of each target video; according to the preset weighted aggregation rule, performing weighted aggregation processing on the initial audio feature vector of each target video to obtain the target audio feature vector of each target video.

[0084] The dimensions of the target frame feature vector and the target audio feature vector may be the same or different, and the present invention does not limit this here. Taking the case where the dimensions of the target frame feature vector and the target audio feature vector are the same as an example, after obtaining the frame feature vector of T * 6 * 1024 dimensions and the audio feature vector of T * 8 * 128 dimensions of each target video, according to the preset weighted aggregation rule, the frame feature vector of T * 6 * 1024 dimensions can be weighted and aggregated into a target frame feature vector of 10 * 1024 dimensions, and the audio feature vector of T * 8 * 128 dimensions can be weighted and aggregated into a target audio feature vector of 10 * 1024 dimensions.

[0085] Similarly, after obtaining the initial video frame feature vector and the initial audio feature vector of the video to be recognized, it further includes performing weighted aggregation processing on the initial video frame feature vector of the video to be recognized according to a preset weighted aggregation rule to obtain the target video frame feature vector of the video to be recognized; performing weighted aggregation processing on the initial audio feature vector of the video to be recognized according to a preset weighted aggregation rule to obtain the target audio feature vector of the video to be recognized, wherein the target video frame feature vector of the video to be recognized has the same dimension as the target video frame feature vector of each target video, and the target audio feature vector of the video to be recognized has the same dimension as the target audio feature vector of each target video.

[0086] In an alternative embodiment, after obtaining the target video frame feature vector and the target audio feature vector of each target video, the method further includes: for each target video, fusing the target video frame feature vector and the target audio feature vector of the target video to obtain the target audio-visual feature vector of the target video. This step realizes the fusion of two modal features by fusing the target video frame feature vector and the target audio feature vector, and obtains a multi-modal target audio-visual feature vector. As an example, the target video frame feature vector and the target audio feature vector can be fused by means of bilinear pooling.

[0087] After obtaining the target audio-visual feature vectors, clustering operations are performed on the target audio-visual feature vectors of multiple target videos to obtain the multiple first clustering centers. This step obtains the first clustering centers by clustering the multi-modal target audio-visual feature vectors, and its accuracy is higher than that of clustering using single-modal features.

[0088] Similarly, after obtaining the target video frame feature vector and the target audio feature vector of the video to be recognized, the method further includes: fusing the target video frame feature vector and the target audio feature vector of the video to be recognized to obtain the target audio-visual feature vector of the video to be recognized; determining the similarity between the video to be recognized and each of the first clustering centers according to the target audio-visual feature vector of the video to be recognized.

[0089] Figure 3 Schematically shows a structural diagram of a video label recognition device 300 according to an embodiment of the present invention, as Figure 3 shown, the video label recognition device 300 includes:

[0090] An acquisition module 301, configured to acquire a video to be recognized and the target feature vector of the video to be recognized, and acquire a preset plurality of first clustering centers, the plurality of first clustering centers being obtained by clustering the target feature vectors of a plurality of target videos, and each of the first clustering centers having a corresponding label;

[0091] A calculation module 302, configured to determine the similarity between the video to be recognized and each of the first cluster centers according to the target feature vector of the video to be recognized;

[0092] An identification module 303, configured to determine, from the multiple first cluster centers, a target cluster center with the maximum similarity to the video to be recognized, and use the label corresponding to the target cluster center as the label of the video to be recognized.

[0093] Optionally, the acquisition module is further configured to: sample the multiple target videos to obtain multiple sampled videos; perform a clustering operation on the target feature vectors of the multiple sampled videos according to a preset clustering rule to obtain multiple second cluster centers; in each round of iteratively updating the multiple second cluster centers, sample the remaining target videos in the multiple target videos except the sampled videos to obtain multiple new sampled videos, and iteratively update the multiple second cluster centers according to the target feature vectors of the multiple new sampled videos until a preset stop condition is reached, and use the second cluster centers after the last update as the first cluster centers.

[0094] Optionally, the device further includes a setting module, configured to: after obtaining the multiple first cluster centers, for each first cluster center, determine the label of the first cluster center according to the labels of the target videos within the cluster where the first cluster center is located.

[0095] Optionally, the acquisition module is further configured to: use a pre-constructed 3D convolutional neural network model to obtain the initial feature vectors of all target videos, where the dimensions of the initial feature vectors of target videos with different durations are different; perform a weighted aggregation process on the initial feature vectors of each target video according to a preset weighted aggregation rule to obtain the target feature vector of each target video; the dimensions of the target feature vectors of each target video are the same;

[0096] And use the pre-constructed 3D convolutional neural network model to obtain the initial feature vector of the video to be recognized, and perform a weighted aggregation process on the initial feature vector of the video to be recognized according to the weighted aggregation rule to obtain the target feature vector of the video to be recognized, and the target feature vector of the video to be recognized has the same dimension as the target feature vector of each target video.

[0097] Optionally, the obtaining module is further configured to: for each target video, segment the target video according to a preset segmentation rule to obtain a plurality of sub-samples; the durations of the sub-samples obtained after segmenting all target videos are the same; for each sub-sample, input the sub-sample into a pre-constructed 3D convolutional neural network model, and use the output of the pre-constructed 3D convolutional neural network model as the initial feature vector of the sub-sample, where the initial feature vector of the sub-sample is a feature vector of W*C dimensions, and W and C are integers greater than 1 respectively; concatenate the initial feature vectors of W*C of the plurality of sub-samples to obtain the initial feature vector of the target video, where the initial feature vector of the target video is a feature vector of H*W*C dimensions, and H represents the number of sub-samples obtained after segmenting the target video; according to the initial feature vectors of H*W*C of all target videos, calculate the variance of the aggregation value of the channel feature maps of each channel, where the channel is the C dimensions of the initial feature vector, and the channel feature map is a two-dimensional matrix formed by the H dimensions and the W dimensions in the initial feature vector; sort the variances of the aggregation values of the channel feature maps in descending order, and select the first N channels corresponding to the variances as the target channels, where N is an integer greater than or equal to 1; determine the normalized weight of the channel feature map of each target channel; according to the feature map activation value in the channel feature map of the target channel and the sum of the feature map activation values of the channel feature maps of all target channels, and according to the normalized weight of the channel feature map of the target channel and the initial feature vector of the target video, determine the weighted sum of the channel feature map of the target channel; concatenate the weighted sums of the channel feature maps of the N target channels to obtain the target feature vector of each target video.

[0098] The obtaining module is further configured to: use the pre-constructed 3D convolutional neural network model to obtain the initial feature vector of the video to be recognized, including: segment the video to be recognized according to a preset segmentation rule to obtain a plurality of sub-samples; the durations of the sub-samples obtained after segmenting the video to be recognized are the same; for each sub-sample, input the sub-sample into a pre-constructed 3D convolutional neural network model, and use the output of the pre-constructed 3D convolutional neural network model as the initial feature vector of the sub-sample, where the initial feature vector of the sub-sample is a feature vector of W*C dimensions, and W and C are integers greater than 1 respectively; concatenate the initial feature vectors of W*C of the plurality of sub-samples to obtain the initial feature vector of the video to be recognized, where the initial feature vector of the video to be recognized is a feature vector of K*W*C dimensions, and K represents the number of sub-samples corresponding to the video to be recognized; according to the initial feature vector of K*W*C and the normalized weights of the channel feature maps of the pre-determined N target channels, determine the weighted sums of the channel feature maps of the N target channels; concatenate the weighted sums of the channel feature maps of the N target channels to obtain the target feature vector of the video to be recognized.

[0099] Optionally, the initial feature vector includes an initial video frame feature vector and an initial audio feature vector;

[0100] The obtaining module is further configured to: for each target video or video to be recognized, obtain the initial video frame feature vector and the initial audio feature vector of each sub-target video or sub-video to be recognized; splice the initial video frame feature vectors and the initial audio feature vectors of the multiple sub-target videos to obtain the initial video frame feature vector and the initial audio feature vector of the target video, or splice the initial video frame feature vectors and the initial audio feature vectors of the multiple sub-videos to be recognized to obtain the initial video frame feature vector and the initial audio feature vector of the video to be recognized; perform weighted aggregation processing on the initial video frame feature vector of each target video or video to be recognized according to a preset weighted aggregation rule to obtain the target video frame feature vector of each target video or video to be recognized, where the target video frame feature vector of the video to be recognized has the same dimension as the target video frame feature vector of each target video, and the target audio feature vector of the video to be recognized has the same dimension as the target audio feature vector of each target video; perform weighted aggregation processing on the initial audio feature vector of each target video or video to be recognized according to a preset weighted aggregation rule to obtain the target audio feature vector of each target video or video to be recognized; fuse the target video frame feature vector and the target audio feature vector of each target video or video to be recognized to obtain the target audio-visual feature vector of each target video or video to be recognized;

[0101] The obtaining module is further configured to: perform a clustering operation on the target audio-visual feature vectors of multiple target videos to obtain the multiple first clustering centers;

[0102] The calculating module is further configured to: determine the similarity between the video to be recognized and each of the first clustering centers according to the target audio-visual feature vector of the video to be recognized.

[0103] An embodiment of the present invention further provides an electronic device, as Figure 4 shown, including a processor 401, a communication interface 402, a memory 403, and a communication bus 404. Among them, the processor 401, the communication interface 402, and the memory 403 complete communication with each other through the communication bus 404. The memory 403 is used to store a computer program; when the processor 401 executes the program stored in the memory 403, the following steps are implemented:

[0104] Obtain a video to be recognized and the target feature vector of the video to be recognized, and obtain a preset multiple first clustering centers, where the multiple first clustering centers are obtained by clustering the target feature vectors of multiple target videos, and each of the first clustering centers has a corresponding label;

[0105] Determine the similarity between the video to be recognized and each of the first clustering centers according to the target feature vector of the video to be recognized;

[0106] Determine, from the multiple first clustering centers, a target clustering center with the highest similarity to the video to be recognized, and use the label corresponding to the target clustering center as the label of the video to be recognized.

[0107] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0108] The communication interface is used for communication between the above terminal and other devices.

[0109] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0110] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0111] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium, in which instructions are stored, and when they run on a computer, the computer is made to execute any one of the video label recognition methods in the above embodiments.

[0112] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which when running on a computer, causes the computer to execute the video tag recognition method described in any one of the above embodiments.

[0113] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server, data center, etc. that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).

[0114] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.

[0115] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the related parts, reference can be made to the partial description of the method embodiment.

[0116] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A video tag recognition method, characterized in that, Including: Obtaining the video to be recognized and the target feature vector of the video to be recognized, and obtaining a plurality of preset first clustering centers, where the plurality of first clustering centers are obtained by clustering the target feature vectors of a plurality of target videos, and each of the first clustering centers has a corresponding label; Determining the similarity between the video to be recognized and each of the first clustering centers according to the target feature vector of the video to be recognized; Determining, from the plurality of first clustering centers, a target clustering center with the largest similarity to the video to be recognized, and using the label corresponding to the target clustering center as the label of the video to be recognized; Wherein, the process of obtaining the target feature vector of each target video includes: For each target video, segmenting the target video according to a preset segmentation rule to obtain a plurality of sub-samples; the durations of the sub-samples obtained after segmenting all target videos are the same; for each sub-sample, inputting the sub-sample into a pre-constructed 3D convolutional neural network model, and using the output of the pre-constructed 3D convolutional neural network model as the initial feature vector of the sub-sample; splicing the initial feature vectors of the plurality of sub-samples to obtain the initial feature vector of the target video; calculating the variance of the aggregation value of the channel feature maps of each channel according to the initial feature vectors of all target videos; sorting the variances of the aggregation values of the channel feature maps in descending order, and selecting the first N channels corresponding to the variances as the target channels, where N is an integer greater than or equal to 1; determining the normalized weight of the channel feature map of each target channel according to the feature map activation value in the channel feature map of the target channel and the sum of the feature map activation values of the channel feature maps of all target channels; determining the weighted sum of the channel feature map of the target channel according to the normalized weight of the channel feature map of the target channel and the initial feature vector of the target video; splicing the weighted sums of the channel feature maps of the N target channels to obtain the target feature vector of the target video.

2. The method according to claim 1, wherein The process of obtaining the plurality of first clustering centers by clustering the target feature vectors of a plurality of target videos includes: Sampling the plurality of target videos to obtain a plurality of sampled videos; Performing a clustering operation on the target feature vectors of the plurality of sampled videos according to a preset clustering rule to obtain a plurality of second clustering centers; Sampling the remaining target videos in the plurality of target videos except the sampled videos to obtain a plurality of new sampled videos, and iteratively updating the plurality of second clustering centers according to the target feature vectors of the plurality of new sampled videos until a preset stop condition is reached, and using the second clustering center after the last update as the first clustering center.

3. The method according to claim 2, characterized in that The method further includes: After obtaining the plurality of first clustering centers, for each first clustering center, determining the label of the first clustering center according to the labels of the target videos within the cluster where the first clustering center is located.

4. The method according to claim 1, characterized in that, The dimensions of the initial feature vectors of target videos with different durations are different; the dimensions of the target feature vectors of each target video are the same; obtaining the target feature vector of the video to be recognized includes: using the pre-constructed 3D convolutional neural network model to obtain the initial feature vector of the video to be recognized, and performing weighted aggregation processing on the initial feature vector of the video to be recognized according to the weighted aggregation rule to obtain the target feature vector of the video to be recognized, and the target feature vector of the video to be recognized has the same dimension as the target feature vectors of each target video.

5. The method according to claim 4, characterized in that, The initial feature vector of the subsample is a feature vector of W*C dimensions, where W and C are integers greater than 1; splicing the initial feature vectors of the multiple subsamples to obtain the initial feature vector of the target video includes: splicing the initial feature vectors of the multiple subsamples W*C to obtain the initial feature vector of the target video, and the initial feature vector of the target video is a feature vector of H*W*C dimensions, where H represents the number of subsamples obtained after the target video is segmented; Calculating the variance of the aggregation values of the channel feature maps of each channel according to the initial feature vectors of all target videos includes: calculating the variance of the aggregation values of the channel feature maps of each channel according to the initial feature vectors of all target videos H*W*C, where the channel is the C dimensions of the initial feature vector, and the channel feature map is a two-dimensional matrix composed of the H dimensions and the W dimensions in the initial feature vector; Using the pre-constructed 3D convolutional neural network model to obtain the initial feature vector of the video to be recognized includes: segmenting the video to be recognized according to a preset segmentation rule to obtain multiple subsamples; the durations of the multiple subsamples obtained after the video to be recognized is segmented are the same; for each subsample, inputting the subsample into the pre-constructed 3D convolutional neural network model, and taking the output of the pre-constructed 3D convolutional neural network model as the initial feature vector of the subsample, and the initial feature vector of the subsample is a feature vector of W*C dimensions, where W and C are integers greater than 1; splicing the initial feature vectors of the multiple subsamples W*C to obtain the initial feature vector of the target video, and the initial feature vector of the target video is a feature vector of K*W*C dimensions, where K represents the number of subsamples corresponding to the video to be recognized; Performing weighted aggregation processing on the initial feature vector of the video to be recognized according to the weighted aggregation rule to obtain the target feature vector of the video to be recognized includes: determining the weighted sum of the channel feature maps of the N target channels according to the initial feature vector K*W*C and the normalized weights of the channel feature maps of the pre-determined N target channels; splicing the weighted sums of the channel feature maps of the N target channels to obtain the target feature vector of the video to be recognized.

6. The method according to claim 5, wherein The initial feature vector includes an initial video frame feature vector and an initial audio feature vector; The process of obtaining the initial feature vectors of each target video or video to be recognized includes: for each target video or video to be recognized, obtaining the initial frame feature vectors and initial audio feature vectors of each sub-target video or sub-video to be recognized; concatenating the initial frame feature vectors and initial audio feature vectors of multiple sub-target videos to obtain the initial frame feature vector and initial audio feature vector of the target video, or concatenating the initial frame feature vectors and initial audio feature vectors of multiple sub-videos to be recognized to obtain the initial frame feature vector and initial audio feature vector of the video to be recognized. The target feature vector of each target video includes the target frame feature vector of each target video and the target audio feature vector of each target video. The target feature vector of the video to be recognized includes the target frame feature vector of the video to be recognized and the target audio feature vector of the video to be recognized. Wherein, the dimension of the target frame feature vector of the video to be recognized is the same as that of the target frame feature vector of each target video, and the dimension of the target audio feature vector of the video to be recognized is the same as that of the target audio feature vector of each target video; fusing the target frame feature vector and target audio feature vector of each target video or video to be recognized to obtain the target audio-visual feature vector of each target video or video to be recognized. The process of clustering the target feature vectors of multiple target videos to obtain the first cluster centers includes: performing a clustering operation on the target audio-visual feature vectors of multiple target videos to obtain the multiple first cluster centers. Determining the similarity between the video to be recognized and each of the first cluster centers according to the target feature vector of the video to be recognized includes: determining the similarity between the video to be recognized and each of the first cluster centers according to the target audio-visual feature vector of the video to be recognized.

7. A video tag recognition device, characterized in that Includes: An acquisition module, configured to acquire the video to be recognized and the target feature vector of the video to be recognized, and acquire a preset plurality of first cluster centers, the plurality of first cluster centers being obtained by clustering the target feature vectors of multiple target videos, and each of the first cluster centers has a corresponding label. A calculation module, configured to determine the similarity between the video to be recognized and each of the first cluster centers according to the target feature vector of the video to be recognized. An identification module, configured to determine, from the plurality of first cluster centers, a target cluster center with the maximum similarity to the video to be recognized, and use the label corresponding to the target cluster center as the label of the video to be recognized. Among them, the obtaining module is further configured to, for each target video, perform segmentation processing on the target video according to a preset segmentation rule to obtain a plurality of sub-samples; the durations of the sub-samples obtained after segmenting all target videos are the same; for each sub-sample, input the sub-sample into a pre-constructed 3D convolutional neural network model, and use the output of the pre-constructed 3D convolutional neural network model as the initial feature vector of the sub-sample; splice the initial feature vectors of the plurality of sub-samples to obtain the initial feature vector of the target video; calculate the variance of the aggregation value of the channel feature maps of each channel according to the initial feature vectors of all target videos; sort the variances of the aggregation values of the channel feature maps in descending order, and select the first N channels corresponding to the variances as target channels, where N is an integer greater than or equal to 1; determine the normalized weight of the channel feature map of each target channel according to the feature map activation value in the channel feature map of the target channel and the sum of the feature map activation values of the channel feature maps of all target channels; determine the weighted sum of the channel feature map of the target channel according to the normalized weight of the channel feature map of the target channel and the initial feature vector of the target video; splice the weighted sums of the channel feature maps of the N target channels to obtain the target feature vector of the target video.

8. The device according to claim 7, characterized in that, The obtaining module is further configured to: The dimensions of the initial feature vectors of target videos with different durations are different; the dimensions of the target feature vectors of each target video are the same; use the pre-constructed 3D convolutional neural network model to obtain the initial feature vector of the video to be recognized, and perform weighted aggregation processing on the initial feature vector of the video to be recognized according to the weighted aggregation rule to obtain the target feature vector of the video to be recognized, and the target feature vector of the video to be recognized has the same dimension as the target feature vector of each target video.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store a computer program; The processor is configured to, when executing the program stored on the memory, implement the method steps described in any one of claims 1-6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Video clustering method and device thereof

    CN113515668A