Video clustering method and device

By using feature extraction and clustering videos using feature extraction models and clustering modules, the problem of large workload of video annotation in the prior art is solved, and efficient video clustering and accurate feature extraction are achieved.

CN113515668BActive Publication Date: 2025-05-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110025310.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-08
Publication Date
2025-05-16
Estimated Expiration
2041-01-08

AI Technical Summary

Technical Problem

In the prior art, video labeling is required based on the content of the sample video, resulting in a large amount of work on labeling the sample video.

Method used

By obtaining the video set, using the feature extraction model to extract the video frame sequence to obtain the video semantic feature vector. The feature extraction model is trained based on image information and tag information of multiple sample videos, and the tag information includes video number and clustering category. The clustering module assists in learning cluster categories and reduces the workload of annotating sample videos.

Benefits of technology

It realizes the reduction of video annotation workload, improves the accuracy of video clustering, and reduces the training complexity of feature extraction models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113515668B_ABST
    Figure CN113515668B_ABST
Patent Text Reader

Abstract

The present application relates to the field of artificial intelligence technology, and specifically provides a video clustering method and device, the method comprising: obtaining a video set, the video set comprising a plurality of videos to be processed; extracting features from a video frame sequence of each video in the video set by a feature extraction model, and obtaining a video semantic feature vector of each video; the feature extraction model is obtained by training an original model using image information of a plurality of sample videos and label information corresponding to the sample videos, the label information comprising a first label for describing the number of the sample video in the plurality of sample videos, and a second label for describing the cluster category to which the sample video belongs: different sample videos have different first labels; clustering the videos in the video set according to the video semantic feature vectors of each video, and dividing each video in the video set into at least one cluster category; this scheme reduces the workload of labeling the sample videos and ensures the accuracy of clustering the videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a video clustering method and device. Background Art

[0002] In the field of video search, in order to facilitate video search, it is generally necessary to classify the video. In the existing video classification method, a feature extraction model based on deep learning is generally used to extract the feature vector of the video, and then the video is classified according to the feature vector of the video. In order to ensure the accuracy of the feature vector extracted by the feature extraction model, it is necessary to train the feature extraction model with sample data. In practice, relevant personnel are required to first watch the sample video in the sample data, and then label the sample video with a label that can represent the content of the sample video according to the content of the sample video, so as to supervise the feature extraction model based on the labeled label. In this process, since relevant personnel are required to label the sample video according to the video content, the workload of labeling the sample video is large. Summary of the invention

[0003] The embodiments of the present application provide a video clustering method and device to solve the problem in the prior art that the workload of labeling sample videos is large due to the need to label videos according to the content of the sample videos.

[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.

[0005] According to one aspect of an embodiment of the present application, a video clustering method is provided, including:

[0006] Acquire a video set, where the video set includes a plurality of videos to be processed;

[0007] The video frame sequence of each video in the video set is subjected to feature extraction by a feature extraction model to obtain a video semantic feature vector of each video; the feature extraction model is obtained by training an original model using image information of a plurality of sample videos and label information corresponding to the sample videos, wherein the label information includes a first label for describing the serial number of the sample video in the plurality of sample videos, and a second label for describing the cluster category to which the sample video belongs; the original model includes a first branch network and a clustering module, the first branch network is used to learn the image information of the sample video and the first label, and the clustering module is used to assist the first branch network in learning the second label; wherein the first labels of different sample videos are different;

[0008] The videos in the video set are clustered according to the video semantic feature vectors of each video, and each video in the video set is divided into at least one cluster category.

[0009] According to one aspect of an embodiment of the present application, a video clustering device is provided, the device comprising:

[0010] A video set acquisition module, used to acquire a video set, wherein the video set includes a plurality of videos to be processed;

[0011] A feature extraction module, used for extracting features from a video frame sequence of each video in the video set through a feature extraction model to obtain a video semantic feature vector of each video; the feature extraction model is obtained by training an original model using image information of a plurality of sample videos and label information corresponding to the sample videos, wherein the label information includes a first label for describing a sequence number of the sample video in a plurality of sample videos, and a second label for describing a cluster category to which the sample video belongs; the original model includes a first branch network and a clustering module, wherein the first branch network is used for learning the image information of the sample video and the first label, and the clustering module is used for assisting the first branch network in learning the second label; wherein the first labels of different sample videos are different;

[0012] The video clustering module is used to cluster the videos in the video set according to the video semantic feature vector of each video, and divide each video in the video set into at least one clustering category.

[0013] According to one aspect of an embodiment of the present application, an electronic device is provided, including: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the video clustering method described above is implemented.

[0014] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the video clustering method as described above is implemented.

[0015] In the scheme of the present application, the number corresponding to each sample video in the sample video set is used as the first label of the sample video, and the cluster category obtained by clustering the feature vector of the sample video is used as the second label of the sample video, and on this basis, the feature extraction model for extracting the video semantic features of the video is trained. Since the first labels of different sample videos are different, it is equivalent to treating each sample video as a category in the training process according to the first label of the sample video. Therefore, the first iterative training can enable the first branch network to accurately identify the features of different videos. On this basis, the first branch network is trained in combination with the second label of the sample video. Since the second label describes the cluster category to which the sample video belongs, in the process of clustering based on the semantic features of the sample video, similar video semantic feature vectors will be clustered into the same cluster category, and dissimilar video semantic feature vectors will correspond to different cluster categories. Therefore, on the basis of training based on the first label, supervision through the second label of the sample video will make the distance between the video semantic feature vectors output by the first branch network for similar videos closer and closer, while the distance between the video semantic feature vectors output for dissimilar videos will be farther and farther. After the training of the first branch network is completed, the first branch network is used as a feature extraction model to increase the distinction between video semantic feature vectors extracted for dissimilar videos, while the video semantic feature vectors extracted for similar videos are more compact. Therefore, it is convenient to cluster videos according to the obtained video semantic feature vectors, thereby ensuring the accuracy of video clustering.

[0016] Moreover, in the solution of the present application, there is no need to manually label the sample videos according to the content of the videos, but only to sequentially number the sample videos, which greatly reduces the workload of labeling the sample videos.

[0017] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0019] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied is shown.

[0020] Figure 2The flowchart of a video clustering method according to an embodiment of the present application is shown.

[0021] Figure 3 FIG. 4 is a flowchart of training a first branch network according to an embodiment.

[0022] Figure 4 is a flowchart showing training a first branch network within a training cycle according to an embodiment.

[0023] Figure 5 is a flow chart showing step 220 according to an embodiment.

[0024] Figure 6 is a schematic diagram showing training of a first branch network according to an embodiment.

[0025] Figure 7 It is a flowchart of the steps after step 230 according to an embodiment of the present application.

[0026] Figure 8 The figure is a schematic diagram of a video playback interface according to an embodiment.

[0027] Fig. 9 is a block diagram of a video clustering device according to an embodiment.

[0028] Fig.10 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown. DETAILED DESCRIPTION

[0029] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.

[0030] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0031] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0032] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0033] It should be noted that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0034] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0035] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0036] In the field of video search, in order to facilitate video search, it is generally necessary to classify the videos. In the existing video classification methods, a feature extraction model based on deep learning is generally used to extract the feature vector of the video, and then the video is classified according to the feature vector of the video. In order to ensure the accuracy of the feature vector extracted by the feature extraction model, it is necessary to train the feature extraction model with sample data. In practice, relevant personnel are required to first watch the sample video in the sample data, and then label the sample video with a label that can characterize the content of the sample video according to the content of the sample video, so as to perform supervised training on the feature extraction model based on the labeled label. In this process, since relevant personnel are required to label the sample video according to the video content, the workload of constructing the sample data is large. In order to solve this problem, a solution of the present application is proposed.

[0037] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied is shown.

[0038] like Figure 1 As shown, the system architecture may include terminal devices (such as Figure 1 The embodiment of the present invention includes one or more of the smartphone 101, tablet computer 102 and portable computer 103, which may also be a desktop computer, etc.), a network 104 and a server 105. The network 104 is used to provide a medium for a communication link between the terminal device and the server 105. The network 104 may include various connection types, such as a wired communication link, a wireless communication link, etc.

[0039] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. According to the implementation requirements, there may be any number of terminal devices, networks and servers. For example, the server 105 may be a server cluster composed of multiple servers.

[0040] The video clustering method of the present application can be executed by the server 105, wherein the user can upload the video to the server 105 through the terminal device, so that the server 105 builds a video set based on the videos uploaded by each terminal device. Then the server 105 clusters each video in the video set according to the scheme of the present application to determine the cluster category to which each video belongs.

[0041] In some embodiments of the present application, the server 105 can also push videos based on the cluster type to which the videos belong. The server 105 can obtain the cluster type of the video currently played by the terminal device, and then select the video of the cluster type from the unplayed video set as the video to be pushed to the terminal device. On this basis, it can be convenient for the user of the terminal device to browse similar videos.

[0042] The implementation details of the technical solution of the embodiment of the present application are described in detail below:

[0043] Figure 2 A flowchart of a video clustering method according to an embodiment of the present application is shown. The method can be executed by a computer device with processing capabilities, such as a laptop computer, a desktop computer, a smart phone, a server, etc., which is not specifically limited here. Figure 2 As shown, the method at least includes steps 210 to 230, which are described in detail as follows:

[0044] Step 210: Obtain a video set, where the video set includes a plurality of videos to be processed.

[0045] Step 220, extracting features from the video frame sequence of each video in the video set by a feature extraction model to obtain a video semantic feature vector of each video; the feature extraction model is obtained by training an original model using image information of multiple sample videos and label information corresponding to the sample videos, the label information including a first label for describing the serial number of the sample video in the multiple sample videos, and a second label for describing the cluster category to which the sample video belongs; the original model includes a first branch network and a clustering module, the first branch network is used to learn the image information of the sample video and the first label, and the clustering module is used to assist the first branch network in learning the second label; wherein the first labels of different sample videos are different;

[0046] The video frame sequence of the video can be obtained by frame segmentation and sampling.

[0047] In some embodiments of the present application, before step 210, the method further includes: dividing the video into frames to obtain an initial video frame sequence of the video; sampling the initial video frame sequence of the video to obtain a video frame sequence of the video. In a specific embodiment, the number of video frames in the video frame sequence can be limited by setting a sampling rate, for example, by setting a sampling rate to ensure that the video frame sequence of each video includes 8 video frames.

[0048] The semantic feature vector of a video is used to represent the semantics of the video. The semantic features of a video are the content contained and reflected in the video from a human perspective. For example, a vehicle driving on an overpass or someone laughing are semantic features abstracted from human understanding. For a video, it contains not only structural information in the spatial domain, but also contextual information in the temporal domain. Therefore, the feature extraction model needs to be able to extract features from the video in both the spatial and temporal domains.

[0049] In the scheme of the present application, the feature extraction model is obtained by training the original model. The original model includes a first branch network. After the training is completed, the trained first branch network is used as the feature extraction model. The first branch network can be constructed by a convolutional neural network. In some embodiments of the present application, the spatial domain features and temporal features of the video can be fused based on a three-dimensional (3D) convolution model to extract the characteristics of the features that characterize the content of the picture presented in the video. The feature extraction model can be a 3D convolution model. The 3D convolution model is formed by stacking multiple continuous video frames into a cube, and then applying a 3D convolution kernel in the cube. In the 3D convolution model, each feature map in the convolution layer is connected to multiple adjacent video frames in the previous layer, thereby capturing the motion information in the video.

[0050] In some embodiments of the present application, the first branch network can also be a time-shift (Temporal Shift Module, TSM) neural network model. The TSM neural network model is a neural network model that maintains the complexity of the 2D convolution model but can achieve the effect of the 3D convolution model. The TSM neural network model adds a TSM module to the 2D convolution model. The TSM module performs effective time modeling through feature maps along the time dimension. It has no redundant calculations on the basis of 2D convolution, but realizes powerful time modeling capabilities. The TSM neural network model decomposes the 2D convolution operation into two processes: displacement and weight superposition. It fuses the spatiotemporal context information without introducing a large amount of calculations. In other words, the TSM neural network model has the same spatiotemporal modeling capabilities as the 3D convolution model, while enjoying the same calculations and parameters as the 2D convolution model.

[0051] In a specific embodiment, the first branch network may be a ResNet neural network, that is, a 50-layer residual network.

[0052] In order to ensure the accuracy of the video semantic feature vector output by the feature extraction model for the video, before step 210, the original model needs to be trained with sample videos.

[0053] In an embodiment of the present application, the feature extraction model is trained by the sample videos in the training sample set. Specifically, in the scheme of the present application, the feature extraction model is alternately trained based on the image information of each sample video in the training sample set and the label information (first label and second label) of the sample video. Specifically, the original model can be trained based on the image information of the sample video and the first label of the sample video, and then the original model can be trained based on the image information of the sample video and the second label of the sample video.

[0054] The image information of the sample video is reflected by the video frame sequence of the sample video. Therefore, before training, the video frame sequence of the sample video is obtained by dividing the sample video into frames and sampling the sample video.

[0055] In the solution of the present application, the first tag is used to describe the number of the sample video in the multiple sample videos. Therefore, before the original model is trained, the sample videos in the training sample set are numbered, and different sample videos have different corresponding numbers. In some embodiments of the present application, the sample videos in the training sample set can be numbered sequentially, and one number uniquely corresponds to one sample video.

[0056] In the scheme of the present application, the second label is used to describe the cluster category to which the sample video belongs, wherein the cluster category to which the sample video belongs is determined by clustering based on the video semantic feature vector of the sample video. In other words, in order to determine the second label of the sample video, the video frame sequence of the sample video is firstly extracted through the first branch network in the original model to obtain the video semantic feature vector of the sample video; on this basis, the clustering module clusters the sample videos in the training sample set based on the video semantic feature vector of each sample video, and correspondingly determines the cluster category to which the sample video belongs.

[0057] In some embodiments of the present application, each cluster category may be numbered, and then the number corresponding to the cluster category to which the sample video belongs is used as the second label of the sample video. For distinction, the number corresponding to the cluster category may be called the second number.

[0058] The clustering module can cluster the sample videos in the training sample set according to the clustering algorithm based on the video semantic feature vectors of the sample videos. The clustering algorithm can be a K-means clustering algorithm, a mean shift clustering algorithm, a clustering algorithm using a Gaussian mixture model for maximum expectation estimation, an agglomerative hierarchical clustering algorithm, etc., which are not specifically limited here.

[0059] During the training process, the first branch network is used to extract features from the video frame sequence of the sample video to obtain the video semantic feature vector of the sample video, and then the loss function value of the target loss function of the first branch network is calculated based on the video semantic feature vector of the sample video and the first label of each sample video, and then the parameters of the first branch network are reversely adjusted based on the obtained loss function value.

[0060] After training for a period of time according to the first label of the sample video, the first branch network is trained using the second label of the sample video. Specifically, the first branch network, after parameter adjustment, extracts features from the video frame sequence of the sample video again to obtain a video semantic feature vector of the sample video, and then clusters the sample videos in the training sample set according to the video semantic feature vector of the sample video obtained again, and divides the sample videos in the training sample set into at least one clustering category, thereby determining the clustering category to which each sample video belongs, and correspondingly obtaining the second label of the sample video. On this basis, the loss function value of the target loss function is calculated based on the video semantic feature vector of the sample video obtained again and the second label of the sample video, and the parameters of the first branch network are adjusted.

[0061] For ease of description, the iterative training of the first branch network based on the first label of the sample video is referred to as the first iterative training, and the iterative training of the first branch network based on the second label of the sample video is referred to as the second iterative training.

[0062] Figure 3 is a flowchart showing training of the first branch network according to an embodiment, such as Figure 3 As shown, in the first iterative training process, after the first branch network obtains the video semantic feature vector of the sample video according to the video frame sequence of the sample video, the function value of the target loss function is calculated according to the first label of the sample video and the video semantic feature vector of the sample video, and then the parameters of the first branch network are reversely adjusted according to the loss function value of the target loss function. In the second iterative training process, after the video semantic feature vector of the sample video is obtained by feature extraction through the first branch network, the sample videos in the training sample set are clustered according to the video semantic feature vectors of each sample video, and the cluster category to which each sample video belongs is determined, and then the second label of the sample video is determined; on this basis, the loss function value of the target loss function is calculated based on the second label of the sample video and the video semantic feature vector of the sample video, and then the parameters of the first branch network are reversely adjusted according to the loss function value of the target loss function.

[0063] Please continue reading Figure 2 , step 230, clustering the videos in the video set according to the video semantic feature vector of each video, and dividing each video in the video set into at least one clustering category. In some embodiments of the present application, the video semantic feature vector of the video can be directly used as the feature vector of the video, and then clustering is performed based on the feature vector of the video.

[0064] In some embodiments of the present application, other information of the video and the video semantic feature vector of the video may be combined to generate a feature vector of the video, and then clustering may be performed based on the feature vector of the video.

[0065] In one embodiment, step 220 further includes: obtaining an additional feature vector of the video, the additional feature vector including at least one of an audio semantic feature vector, a character semantic feature vector and a title semantic feature vector; the audio semantic feature vector is obtained by extracting semantic features of the audio in the video; the character semantic feature vector is obtained by extracting semantic features of the characters in the video frame of the video; the title semantic feature vector is obtained by extracting semantic features of the title text of the video; and the video semantic feature vector of the video is fused with the additional feature vector of the video to obtain the feature vector of the video.

[0066] For a video, it may include audio, such as background sound, character dialogue, narration, etc. in the video. By extracting semantic features from the audio in the video, an audio semantic feature vector of the video can be obtained. In one embodiment, semantic features can be extracted from the audio in the video by a speech recognition model.

[0067] The title text of the video may be a title set for the video by the user who uploaded the video. The semantic feature vector of the title of the video may be obtained by extracting semantic features from the title text of the video.

[0068] For a video, there may be characters, such as text, in the video frames of a video sequence. Therefore, optical character recognition (OCR) can be performed on the video frames with characters, and then semantic features of the recognized characters can be extracted, that is, the character semantic feature vector of the characters in the video frame can be obtained. It can be understood that in the video frame sequence corresponding to the video, there may be multiple video frames in which characters exist. Therefore, the character semantic feature vector of the video is obtained by fusing the character semantic feature vectors of each video frame. Among them, one implementation method of fusing the character semantic feature vectors of each video frame can be to splice the character semantic feature vectors of each video frame.

[0069] The video semantic feature vector and the additional feature vector of the video may be concatenated to achieve fusion of the video semantic feature vector and the additional feature vector.

[0070] In the solution of this embodiment, the feature vector of the video is obtained by fusing the video semantic feature vector of the video with the additional feature vector of the video, so that the feature vector can reflect the features of the video in multiple dimensions.

[0071] In some embodiments of the present application, the videos in the video collection can be clustered by K-means clustering algorithm, mean shift clustering algorithm, clustering algorithm using Gaussian mixture model for maximum expectation estimation, agglomerative hierarchical clustering algorithm, etc., which are not specifically limited here.

[0072] In a specific embodiment of the present application, the K-means clustering algorithm is used to cluster the sample videos in the sample video set. Specifically, the total number of categories is first set, for example, K; then the sample videos in the sample video set are divided into K groups, and in each group of sample videos, the feature vector of a sample video is randomly selected as the initial cluster center of the group of sample videos; then the distance between the feature vector of each sample video and each cluster center is calculated, and each sample video is assigned to the cluster center closest to it. The sample video corresponding to the cluster center and the sample video assigned to the cluster center represent a cluster, and a cluster corresponds to a cluster category. Among them, each time a sample video is assigned to the cluster center, the cluster center of the cluster will be recalculated according to the feature vector of the existing sample video in the cluster. Repeat the above process until the clustering end condition is met.

[0073] The clustering termination condition may be that compared with the previous clustering result, the number of sample videos assigned to different cluster centers in this clustering result does not exceed a first preset number (a second preset number, for example, 0, or an integer greater than 0); the clustering termination condition may also be that compared with the previous clustering result, the number of clusters whose cluster centers have changed does not exceed a second preset number (the second preset number, for example, 0, or an integer greater than 0).

[0074] In some embodiments of the present application, the total number of categories may be set so that clustering is performed according to the total number of categories. After clustering is completed, the videos in the video set are divided into the set total number of cluster categories.

[0075] In the scheme of the present application, the number corresponding to each sample video in the sample video set is used as the first label of the sample video, and the cluster category obtained by clustering the feature vector of the sample video is used as the second label of the sample video, and on this basis, the feature extraction model for extracting the video semantic features of the video is trained. Since the first labels of different sample videos are different, it is equivalent to treating each sample video as a category in the training process according to the first label of the sample video. Therefore, the first iterative training can enable the first branch network to accurately identify the features of different videos. On this basis, the first branch network is trained in combination with the second label of the sample video. Since the second label describes the cluster category to which the sample video belongs, in the process of clustering based on the semantic features of the sample video, similar video semantic feature vectors will be clustered into the same cluster category, and dissimilar video semantic feature vectors will correspond to different cluster categories. Therefore, on the basis of training based on the first label, supervision through the second label of the sample video will make the distance between the video semantic feature vectors output by the first branch network for similar videos closer and closer, while the distance between the video semantic feature vectors output for dissimilar videos will be farther and farther. After the training of the first branch network is completed, the first branch network is used as a feature extraction model to increase the distinction between video semantic feature vectors extracted for dissimilar videos, while the video semantic feature vectors extracted for similar videos are more compact. Therefore, it is convenient to perform video clustering based on the obtained video semantic feature vectors, thereby ensuring the accuracy of video clustering.

[0076] Moreover, in the solution of the present application, there is no need to manually label the sample videos according to the content of the videos, but only to number the sample videos, which greatly reduces the workload of labeling the sample videos.

[0077] In some embodiments of the present application, before step 220, the method further includes: using image information of multiple sample videos in a training sample set and label information corresponding to the sample videos, alternating first iterative training and second iterative training for the first branch network in the original model according to a training cycle to obtain a trained first branch network; and using the trained first branch network as the feature extraction model.

[0078] Specifically, in each training cycle, follow the steps below: Figure 4 The process shown trains the first branch network in the original model:

[0079] Step 410, performing a first iterative training on the first branch network according to the first label of the sample video and the first video semantic feature vector of the sample video; the first video semantic feature vector of the sample video is obtained by extracting features of the video frame sequence of the sample video by the first branch network after the second iterative training in the previous training cycle; in the first training cycle, the video frame sequence of the sample video is extracted by the initial first branch network to obtain the corresponding first video semantic feature vector.

[0080] If the number of iterations in the first iterative training reaches a first set number, execute step 420: perform second iterative training on the first branch network according to the second label of the sample video and the second video semantic feature vector of the sample video, until the number of iterations of the second iterative training reaches a second set number; wherein the second video semantic feature vector of the sample video is obtained by performing feature extraction on the video frame sequence of the sample video by the first branch network after the first iterative training in this training cycle is completed.

[0081] As described above, the first iterative training refers to the iterative training of the first branch network according to the first label of the sample video and the feature vector of the sample video; the second iterative training refers to the iterative training of the first branch network according to the second label of the sample video and the feature vector of the sample video.

[0082] It is understandable that, during the training process, the parameters of the first branch network need to be continuously adjusted, and after the parameters are adjusted, the video semantic feature vector of the sample video is output again through the first branch network to obtain the feature vector of the sample video. In other words, the feature vector of the sample video changes with the adjustment of the parameters of the first branch network during the training process, and after the feature vector of the sample video changes, it is necessary to cluster the sample videos in the sample video set again according to the feature vectors of each sample video, and then re-divide each sample video into at least one cluster category, and re-determine the cluster category to which the sample video belongs.

[0083] For ease of description, the video semantic feature vector extracted by the first branch network for the sample video during the first iterative training process is referred to as the first video semantic feature vector; wherein, during the first iterative training process, the initially obtained first video semantic feature vector is obtained by extracting features from the video frame sequence of the sample video by the first branch network after the second iterative training in the previous training cycle. The video semantic feature vector extracted by the first branch network for the sample video during the second iterative training process is referred to as the second video semantic feature vector; wherein, during the second iterative training process, the initially obtained second video semantic feature vector is obtained by extracting features from the video frame sequence of the sample video by the first branch network after the first iterative training in this training cycle.

[0084] In a training cycle, if the number of iterations in the first iterative training does not reach the first set number at this time, the first sub-network continues to be trained in the first iterative training; conversely, if the number of iterations in the first iterative training reaches the first set number, the first sub-network is trained in the second iterative training until the number of iterations in the second iterative training reaches the second set number. The first set number and the second set number can be set according to actual needs, and of course, the first set number and the second set number can be equal or unequal.

[0085] After the first sub-network is trained for several training cycles, the first sub-network can be tested. Specifically, the test video in the test video set is input into the first sub-network, and the first sub-network outputs the corresponding video semantic feature vector for the test video, and then the test video set in the test video set is clustered according to the video semantic feature vector of the test video, and the clustering result is used as the test result. If the obtained clustering result meets the set requirements, the training of the first sub-network is terminated. Conversely, if the obtained clustering result does not meet the set requirements, the training of the first sub-network continues.

[0086] In some embodiments of the present application, Figure 4 As shown, step 410 includes:

[0087] Step 411 , extracting features from the video frame sequence of each sample video through the first branch network after the second iteration training in the previous training cycle, to obtain a first video semantic feature vector of each sample video.

[0088] Step 412: Calculate a first loss function value of a target loss function according to a first video semantic feature vector of each sample video and a first label of each sample video.

[0089] Step 413: Adjust the parameters of the first branch network based on the first loss function value.

[0090] During the first iterative training or the second iterative training of the first branch network, the parameters of the first branch network are adjusted. Before and after the parameters of the first branch network are adjusted, the video semantic feature vectors extracted by the first branch network for the same video may be different.

[0091] In some embodiments of the present application, the first video semantic feature vector of the sample video can be used as the first feature vector of the sample video, and then the loss function value of the target loss function is calculated based on the first feature vector of the sample video and the first label of the sample video.

[0092] In some other embodiments of the present application, the first feature vector of the sample video can be generated in combination with other information of the sample video. In this embodiment, step 412 further includes: obtaining an additional feature vector of the sample video, the additional feature vector including at least one of an audio semantic feature vector, a character semantic feature vector, and a title semantic feature vector; the audio semantic feature vector is obtained by performing semantic feature extraction on the audio in the sample video; the character semantic feature vector is obtained by performing semantic feature extraction on the characters in the video frame of the sample video; the title semantic feature vector is obtained by performing semantic feature extraction on the title text of the sample video; the first video semantic feature vector of the sample video is merged with the additional feature vector of the sample video to obtain the first feature vector of the sample video.

[0093] In some embodiments of the present application, the target loss function may be an Arcface loss function or a Triplet loss function. Compared with other loss functions, the Arcface loss function can make the same class more "compact", compressing the same cluster category into a tighter space, making it denser, so that the features learned by the network have more obvious angular distribution characteristics.

[0094] The function expression of Arcface loss function is:

[0095]

[0096] Wherein, N is the number of sample videos in the sample video set, m is the space margin between different first labels; s is the radius of the space; θ∈(0,π-m); x i Represents the feature vector of the i-th sample video; y i represents the first label of the i-th sample video; W yi It is the weight parameter of the feature extraction model to be adjusted.

[0097] In the scheme of this embodiment, a separate number is given to each sample video in the sample video set as the first label of the sample video, that is, each sample video is regarded as a category. In the process of training the first branch network based on the Arcface loss function, the distance between similar feature vectors will become closer and closer, while the distance between dissimilar feature vectors will become farther and farther. Therefore, after training the first branch network based on the Arcface loss function, the distinction between the video semantic vectors extracted by the trained first branch network for dissimilar videos can be increased, while the video semantic vectors extracted for similar videos are more compact, thereby facilitating the clustering of videos in the video set based on the video semantic vectors of the videos.

[0098] The expression of the Triplet loss function is:

[0099]

[0100] in, represents the feature vector of the reference sample when the i-th sample video is used as the reference sample; Represents the feature vector of the heterogeneous sample corresponding to the reference sample; α1 is the inter-class interval parameter; α2 is the intra-class interval parameter; function

[0101] In this embodiment, it is necessary to first determine the same type and different type samples of each sample video according to the video semantic feature vectors of each sample video in the sample video set. The same type samples refer to sample videos of the same type as the sample video used as the reference sample, and different type samples refer to sample videos of different types from the sample video used as the reference sample. On this basis, each sample video is used as a reference sample to determine the triple element of each reference sample, wherein the triple element includes the feature vectors corresponding to the reference sample, the same type sample of the reference sample, and the different type sample of the reference sample.

[0102] In one embodiment, the distance between any two sample videos can be calculated based on the video semantic feature vectors of each sample video. Then, based on the obtained distance between the two sample videos, the triplet element corresponding to each sample video as a reference sample is determined; and then according to the above Triplet loss function, the second loss value can be calculated.

[0103] For example, the sample videos in the sample video set can be clustered according to the feature vectors of the sample videos or the video semantic feature vectors to determine the cluster category of each sample video. On this basis, the sample video that belongs to the same cluster category as the sample video used as the reference sample and is closest to the sample video can be selected as the same class sample of the sample video; the sample video that does not belong to the same cluster category as the sample video used as the reference sample and is farthest from the sample video can be selected as the different class sample of the sample video.

[0104] Please continue reading Figure 4 In one embodiment, step 420 includes:

[0105] Step 421 , extracting features from the video frame sequence of each sample video through the first branch network after the first iteration training in this training cycle, to obtain a second video semantic feature vector of each sample video.

[0106] Step 422, clustering the sample videos in the training sample set according to the second video semantic feature vector of each sample video through the clustering module, dividing the sample videos in the training sample set into at least one cluster category, and using the second number corresponding to the cluster category to which the sample video belongs as the second label of the sample video.

[0107] As above, the second video semantic feature vector of the sample video can be used as the second feature vector of the sample video. Alternatively, the second video semantic feature vector of the sample video can be fused with the additional feature vector of the sample video, and the fused vector can be used as the second feature vector of the sample video. Then, the sample videos in the training sample set are clustered based on the second feature vector of the sample video.

[0108] Step 423: Calculate a second loss function value of the target loss function according to the second video semantic feature vector of the sample video and the second label of the sample video.

[0109] As above, the second video semantic feature vector of the sample video can be used as the second feature vector of the sample video. Alternatively, the second video semantic feature vector of the sample video can be fused with the additional feature vector of the sample video, and the fused vector can be used as the second feature vector of the sample video. Then, the function value of the target loss function is calculated based on the second feature vector of the sample video and the second label of the sample video.

[0110] It is worth mentioning that during the first iteration training and the second iteration training, the composition of the feature vector (first feature vector, second feature vector) of the sample video is consistent, that is, if during the first iteration training, the video semantic feature vector of the sample video is used as the feature vector of the sample video, then during the second iteration training, the video semantic feature vector of the sample video is also used as the feature vector of the sample video; if during the first iteration training, the fusion result of the video semantic feature vector of the sample video and the additional feature vector is used as the feature vector of the sample video, then during the second iteration training, the fusion result of the video semantic feature vector of the sample video and the additional feature vector of the same category is also used as the feature vector of the sample video. Step 424, based on the second loss function value, adjust the parameters of the first branch network.

[0111] The same as the first iteration training process, the target loss function is the Arcface loss function or the Triplet loss function. The calculation of the specific loss function value is described above and will not be repeated here. It is worth mentioning that the target loss function in the first iteration training process and the second iteration training process is the same.

[0112] The training of one training cycle is completed through the process of steps 411-424, and then the training process of the next training cycle is repeated.

[0113] In some embodiments of the present application, the total number of categories may be pre-set for each training cycle, where the total number of categories indicates the total number of clustering categories for clustering sample videos in each training cycle.

[0114] In some embodiments of the present application, step 422 further includes: obtaining the total number of categories corresponding to this training cycle; based on the total number of categories, clustering the sample videos in the sample set according to the second feature vector of each sample video, and dividing the sample videos in the training sample set into at least one clustering category; wherein the total number of categories corresponding to the next training cycle is greater than the total number of categories corresponding to this training cycle.

[0115] In the present embodiment, during the second iterative training process, the total number of categories is updated in a coarse-to-fine order so that the feature extraction model can distinguish different videos at a coarse granularity. After the current training cycle ends, the total number of categories is increased in the next training cycle. Increasing the total number of categories in each training cycle is equivalent to generating a finer-grained supervisory signal (second label), and iterating in sequence until the training end condition is reached. The total number of categories can be set with an upper limit according to the number of sample videos in the sample video and actual needs. It can be understood that the upper limit of the total number of categories set is not greater than the number of sample videos in the sample video set. When the total number of categories set is the number of sample videos in the sample video set, it is equivalent to a one-to-one correspondence between the sample videos and the clustering categories.

[0116] In this embodiment, the total number of categories set for the training cycle can be flexibly adjusted according to actual needs to adjust the discrimination granularity of the feature extraction model for the video. In addition, the feature extraction model is trained hierarchically, which can improve the convergence speed of the feature extraction model and shorten the training time of the model.

[0117] In some embodiments of the present application, the feature extraction model includes a first convolutional layer, a temporal shift layer, and a second convolutional layer; Figure 5 As shown, step 220 further includes:

[0118] Step 510: For each video, perform a two-dimensional convolution operation on each video frame in a video frame sequence of the video through the first convolution layer to obtain a first feature map of each video frame in the video frame sequence.

[0119] Step 520: Perform a timing shift operation along the time dimension based on the feature map of each video frame in the video frame sequence through the timing shift layer to obtain a second feature map of each video frame.

[0120] Step 530: Perform a two-dimensional convolution operation on the second feature map of each video frame through the second convolution layer to obtain a third feature map of each video frame.

[0121] Step 540, fully connect the third feature map of each video frame in the video frame sequence to obtain the video semantic feature vector of the video. It can be understood that in the process of training the first branch network, the first branch network still extracts the video semantic feature vector of the sample video according to the process of the above steps 510-540. Figure 6 The above steps 510-540 are described in detail.

[0122] Figure 6 is a schematic diagram showing training of a first branch network according to an embodiment, such as Figure 6As shown, the sample video is first framed and sampled to obtain a video frame sequence of the sample video, and then the video frame sequence of the sample video is input into the first branch network, and the first branch network outputs the video semantic feature vector of the sample video.

[0123] In this embodiment, the feature extraction model can be a TSM neural network model, which performs convolution processing through the 2D convolution kernel in the TSM neural network model to obtain the feature map of each video frame, and then fuses the feature maps of each video frame in the time dimension through the TSM module, and then performs a two-dimensional convolution operation again based on the fused feature map, and then fully connects the feature map of each video frame to obtain the video semantic feature vector of the video frame. On this basis, the function value of the target loss function is calculated based on the feature vector of the sample video and the first label of the sample video; and the function value of the target loss function is calculated based on the feature vector of the sample video and the second label of the sample video, and the parameters of the feature extraction model are adjusted according to the calculated function value.

[0124] exist Figure 6 In the example, the first branch network includes a first convolutional layer and a second convolutional layer, and a timing offset layer ( Figure 6 (not shown), the temporal shift layer performs a temporal shift operation on the first feature map of each video frame in the video frame sequence output by the first convolutional layer. Figure 6 The timing offset operation is described in detail.

[0125] like Figure 6 As shown, a two-dimensional convolution operation is performed on each video frame in the video frame sequence through the first convolution layer to obtain a first feature map of each video frame in the video frame sequence. Figure 6 In the above equation, C represents the channel dimension and T represents the time dimension. Figure 6 Each row of the same color in the feature map A represents the first feature map of a video frame, and each small block represents a different channel. For a video, the first feature maps of each video frame in the video frame sequence of the video are spliced ​​in time order, and the result is Figure 6 Then, the temporal shift layer performs temporal shift along the time dimension (T) based on the feature map A. Figure 6 In the middle, one channel is shifted forward and backward along the time dimension to obtain feature map B. After the timing shift is performed, the first and last blank parts can be supplemented in the channel dimension by filling with zeros. Of course, in other embodiments, a cyclic shift method can also be used to supplement the redundant parts after the timing shift to the back, so that the size of the feature map remains unchanged.

[0126] Depend on Figure 6It can be seen from the feature map B in that after the timing shift, the feature map (second feature map) of each video frame incorporates the feature information of adjacent video frames, thereby realizing the fusion of the timing features in the video frame sequence. In other words, by inserting the timing shift layer into the two-dimensional convolution model to perform the timing shift operation, the image features of each video in the video frame sequence and the timing features between each video frame in the video frame sequence are fused.

[0127] In some embodiments of the present application, Figure 7 As shown, after step 230, the method further includes:

[0128] Step 710: Obtain category information of the currently played video, where the category information is used to indicate the target cluster category to which the currently played video belongs.

[0129] Step 720: Select a target video whose clustering category is the target clustering category from the unplayed video set.

[0130] In some embodiments of the present application, the videos in the unplayed video set may be short videos.

[0131] The number of selected target videos can be set according to actual needs. In a specific embodiment, one or more target videos can be selected, which can be determined by the number of video covers that can be displayed in the display interface of the visual display terminal.

[0132] Step 730: Push the target video to the user so that the user terminal displays the target video.

[0133] In this embodiment, by determining the target cluster category to which the currently played video belongs, and then selecting a target video whose cluster category is the target cluster category from a set of unplayed videos, and pushing the selected target video to the user terminal, for the user using the terminal, since the server pushes the target video of the same cluster category as the currently played video to the user, when the user watches the video of interest, he does not need to search again according to the keywords of the currently played video, and can watch videos with the same or similar highlights as the currently played video, which is convenient for users to browse a large number of similar videos, improves video browsing efficiency, and helps to extend user retention time.

[0134] Figure 8 is a schematic diagram of a video playback interface according to an embodiment. Figure 8 As shown, the screen displayed in the P1 area is the screen of the currently played video. Among them, the P2, P3, P4, P5 and P6 areas display the cover of the pushed target video, and the user can select any target video to play based on the multiple pushed target videos displayed.

[0135] The following describes an apparatus embodiment of the present application, which can be used to execute the method in the above-mentioned embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the above-mentioned method embodiment of the present application.

[0136] Fig. 9 is a block diagram of a video clustering device according to an embodiment. Fig. 9 As shown, the video clustering device includes:

[0137] The video set acquisition module 910 is used to acquire a video set, where the video set includes a plurality of videos to be processed.

[0138] The feature extraction module 920 is used to extract features from the video frame sequence of each video in the video collection through a feature extraction model to obtain a video semantic feature vector of each video; the feature extraction model is obtained by training the original model using the image information of multiple sample videos and the label information corresponding to the sample videos, the label information includes a first label for describing the serial number of the sample video in the multiple sample videos, and a second label for describing the cluster category to which the sample video belongs: the original model includes a first branch network and a clustering module, the first branch network is used to learn the image information of the sample video and the first label, and the clustering module is used to assist the first branch network in learning the second label; wherein the first labels of different sample videos are different.

[0139] The video clustering module 930 is configured to cluster the videos in the video set according to the video semantic feature vector of each video, and divide each video in the video set into at least one clustering category.

[0140] In some embodiments of the present application, the video clustering device also includes: a training module, which is used to use the image information of multiple sample videos and the label information corresponding to the sample videos to alternately perform first iterative training and second iterative training on the first branch network in the original model according to a training cycle to obtain a trained first branch network; and use the trained first branch network as the feature extraction model.

[0141] Specifically, the training module further includes a first iterative training unit and a second iterative training unit. In each training cycle, the first branch network in the original model is trained through the processes executed by the first iterative training unit and the second iterative training unit. The first iterative training unit is used to perform a first iterative training on the first branch network according to the first label of the sample video and the first video semantic feature vector of the sample video; the first video semantic feature vector of the sample video is obtained by extracting features from the video frame sequence of the sample video by the first branch network after the second iterative training in the previous training cycle is completed; in the first training cycle, the video frame sequence of the sample video is extracted by the initial first branch network to obtain the corresponding first video semantic feature vector. The second iterative training unit is used to perform a second iterative training on the first branch network according to the second label of the sample video and the second video semantic feature vector of the sample video if the number of iterations in the first iterative training reaches the first set number, until the number of iterations of the second iterative training reaches the second set number; wherein the second video semantic feature vector of the sample video is obtained by extracting features from the video frame sequence of the sample video by the first branch network after the first iterative training in this training cycle is completed.

[0142] In some embodiments of the present application, the first iterative training unit includes: a first feature extraction unit, used to extract features from a video frame sequence of each sample video through the first branch network after the second iterative training in the previous training cycle, to obtain a first video semantic feature vector of each sample video. A first loss function value calculation unit, used to calculate a first loss function value of a target loss function according to the first video semantic feature vector of each sample video and the first label of each sample video. A first adjustment unit, used to adjust the parameters of the first branch network based on the first loss function value.

[0143] In some embodiments of the present application, the second iterative training unit includes: a second feature extraction unit, used to extract features from the video frame sequence of each sample video through the first branch network after the first iterative training in this training cycle, to obtain the second video semantic feature vector of each sample video. A first clustering unit, used to cluster the sample videos in the training sample set according to the second video semantic feature vector of each sample video through the clustering module, divide the sample videos in the training sample set into at least one clustering category, and use the second number corresponding to the clustering category to which the sample video belongs as the second label of the sample video. A second loss function value calculation unit, used to calculate the second loss function value of the target loss function according to the second video semantic feature vector of the sample video and the second label of the sample video. A second adjustment unit, used to adjust the parameters of the first branch network based on the second loss function value.

[0144] In some embodiments of the present application, the first clustering unit includes: a category total acquisition unit, used to acquire the total category number corresponding to the current training cycle. A division unit, used to cluster the sample videos in the training sample set based on the total category number and the second video semantic feature vector of each sample video by the clustering module, and divide the sample videos in the training sample set into at least one cluster category; wherein the total category number corresponding to the next training cycle is greater than the total category number corresponding to the current training cycle.

[0145] In some embodiments of the present application, the target loss function is an Arcface loss function or a Triplet loss function.

[0146] In some embodiments of the present application, the feature extraction model includes a first convolution layer, a timing offset layer, and a second convolution layer; the feature extraction module 920 includes: a first convolution unit, for each video, performing a two-dimensional convolution operation on each video frame in a video frame sequence of the video through the first convolution layer to obtain a first feature map of each video frame in the video frame sequence. A timing offset unit, for performing a timing offset operation along the time dimension based on the feature map of each video frame in the video frame sequence through the timing offset layer to obtain a second feature map of each video frame; a second convolution unit, for performing a two-dimensional convolution operation on the second feature map of each video frame through the second convolution layer to obtain a third feature map of each video frame. A fully connected unit, for fully connecting the third feature map of each video frame in the video frame sequence to obtain a video semantic feature vector of the video.

[0147] In some embodiments of the present application, the video clustering module 930 includes: an additional feature vector acquisition unit, used to acquire an additional feature vector of the video, wherein the additional feature vector includes at least one of an audio semantic feature vector, a character semantic feature vector, and a title semantic feature vector. A fusion unit, used to fuse the video semantic feature vector of the video with the additional feature vector of the video to obtain a feature vector of the video. A second clustering unit, used to perform video clustering based on the feature vectors of each video in the video set, and divide the videos in the video set into at least one clustering category.

[0148] In some embodiments of the present application, the video clustering device further includes: a category information acquisition module, used to acquire category information of the currently played video, wherein the category information is used to indicate the target cluster category to which the currently played video belongs. A selection module, used to select a target video whose cluster category is the target cluster category from a set of unplayed videos. A push module, used to push the target video to a user terminal so that the user terminal displays the target video.

[0149] Fig.10 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown.

[0150] It should be noted that Fig.10 The computer system 1000 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0151] like Fig.10 As shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 to the random access memory (RAM) 1003, such as executing the method in the above embodiment. In the RAM 1003, various programs and data required for system operation are also stored. The CPU 1001, ROM 1002 and RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0152] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read therefrom is installed into the storage section 1008 as needed.

[0153] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part 1009, and / or installed from a removable medium 1011. When the computer program is executed by a central processing unit (CPU) 1001, various functions defined in the system of the present application are executed.

[0154] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0155] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0156] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.

[0157] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the method in the above embodiment is implemented.

[0158] According to one aspect of the present application, an electronic device is also provided, which includes: a processor; a memory, in which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method in the above embodiment is implemented.

[0159] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method in each of the above optional embodiments.

[0160] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.

[0161] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation methods of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the implementation methods of the present application.

[0162] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary technical means in the art that are not disclosed in the present application.

[0163] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A video clustering method, characterized in that: The method comprises: Acquire a video set, where the video set includes a plurality of videos to be processed; The video frame sequence of each video in the video set is subjected to feature extraction by a feature extraction model to obtain a video semantic feature vector of each video; the feature extraction model is obtained by training an original model using image information of a plurality of sample videos and label information corresponding to the sample videos, the label information comprising a first label for describing the serial number of the sample video in a plurality of sample videos, and a second label for describing the cluster category to which the sample video belongs, the cluster category to which the sample video belongs is determined by clustering the video semantic feature vector of the sample video; the original model comprises a first branch network and a clustering module, the first branch network is used to learn the image information of the sample video and the first label, and the clustering module is used to assist the first branch network in learning the image information of the sample video and the second label; wherein the first labels of different sample videos are different, and the feature extraction model is obtained by alternately training the original model based on the image information of the sample video, and the first label and the second label of the sample video; The videos in the video set are clustered according to the video semantic feature vectors of each video, and each video in the video set is divided into at least one cluster category.

2. The method according to claim 1, characterized in that Before extracting features from the video frame sequence of each video in the video set by using the feature extraction model to obtain the video semantic feature vector of each video, the method further includes: Using image information of multiple sample videos in a training sample set and label information corresponding to the sample videos, alternately perform first iterative training and second iterative training on the first branch network in the original model according to a training cycle to obtain a trained first branch network; and use the trained first branch network as the feature extraction model; In each training cycle, the first branch network in the original model is trained according to the following process: The first branch network is trained for the first iteration according to the first label of the sample video and the first video semantic feature vector of the sample video; the first video semantic feature vector of the sample video is obtained by extracting features from the video frame sequence of the sample video by the first branch network after the second iteration training in the previous training cycle; in the first training cycle, the video frame sequence of the sample video is extracted by the initial first branch network to obtain the corresponding first video semantic feature vector; If the number of iterations in the first iterative training reaches a first set number, the first branch network is subjected to second iterative training according to the second label of the sample video and the second video semantic feature vector of the sample video until the number of iterations of the second iterative training reaches a second set number; wherein the second video semantic feature vector of the sample video is obtained by performing feature extraction on the video frame sequence of the sample video by the first branch network after the first iterative training in this training cycle is completed.

3. The method according to claim 2, characterized in that The performing a first iterative training on the first branch network according to the first label of the sample video and the first video semantic feature vector of the sample video includes: Extracting features from a video frame sequence of each sample video using the first branch network after the second iteration training in the previous training cycle, obtaining a first video semantic feature vector of each sample video; Calculate a first loss function value of the target loss function according to the first video semantic feature vector of each sample video and the first label of each sample video; Based on the first loss function value, a parameter of the first branch network is adjusted.

4. The method according to claim 3, characterized in that The performing a second iterative training on the first branch network according to the second label of the sample video and the second video semantic feature vector of the sample video includes: Extract features from the video frame sequence of each sample video through the first branch network after the first iteration training in this training cycle to obtain a second video semantic feature vector of each sample video; Clustering the sample videos in the training sample set according to the second video semantic feature vector of each sample video by the clustering module, dividing the sample videos in the training sample set into at least one cluster category, and using the second number corresponding to the cluster category to which the sample video belongs as the second label of the sample video; Calculating a second loss function value of the target loss function according to a second video semantic feature vector of the sample video and a second label of the sample video; Based on the second loss function value, adjust the parameters of the first branch network.

5. The method according to claim 4, characterized in that The clustering of the sample videos in the training sample set according to the second video semantic feature vector of each sample video by the clustering module, and dividing the sample videos in the training sample set into at least one cluster category, includes: Get the total number of categories corresponding to this training cycle; The clustering module clusters the sample videos in the training sample set based on the total number of categories and the second video semantic feature vector of each sample video, and divides the sample videos in the training sample set into at least one clustering category; wherein the total number of categories corresponding to the next training cycle is greater than the total number of categories corresponding to the current training cycle.

6. The method according to any one of claims 3 to 5, characterized in that: The target loss function is an Arcface loss function or a Triplet loss function.

7. The method according to claim 1, characterized in that The feature extraction model includes a first convolutional layer, a temporal shift layer and a second convolutional layer; The step of extracting features from a video frame sequence of each video in the video set by using a feature extraction model to obtain a video semantic feature vector of each video includes: For each video, performing a two-dimensional convolution operation on each video frame in a video frame sequence of the video through the first convolution layer to obtain a first feature map of each video frame in the video frame sequence; Performing a timing shift operation along the time dimension based on the feature graph of each video frame in the video frame sequence through the timing shift layer to obtain a second feature graph of each video frame; Performing a two-dimensional convolution operation on the second feature map of each video frame through the second convolution layer to obtain a third feature map of each video frame; The third feature map of each video frame in the video frame sequence is fully connected to obtain a video semantic feature vector of the video.

8. The method according to claim 1, characterized in that The step of clustering the videos in the video set according to the video semantic feature vectors of each video, and dividing each video in the video set into at least one cluster category, comprises: Acquire an additional feature vector of the video, wherein the additional feature vector includes at least one of an audio semantic feature vector, a character semantic feature vector, and a title semantic feature vector; Fusion of the video semantic feature vector of the video with the additional feature vector of the video to obtain a feature vector of the video; Video clustering is performed based on the feature vector of each video in the video set, and the videos in the video set are divided into at least one cluster category.

9. The method according to claim 1, characterized in that: After clustering the videos in the video set according to the video semantic feature vectors of each video and dividing each video in the video set into at least one cluster category, the method further includes: Acquire category information of a currently played video, where the category information is used to indicate a target cluster category to which the currently played video belongs; Selecting a target video whose clustering category is the target clustering category from the unplayed video set; Push the target video to the user terminal so that the user terminal displays the target video.

10. A video clustering device, characterized in that: The device comprises: A video set acquisition module, used to acquire a video set, wherein the video set includes a plurality of videos to be processed; A feature extraction module, used for extracting features from a video frame sequence of each video in the video set through a feature extraction model to obtain a video semantic feature vector of each video; the feature extraction model is obtained by training an original model using image information of multiple sample videos and label information corresponding to the sample videos, the label information includes a first label for describing a sequence number of the sample video in multiple sample videos, and a second label for describing a cluster category to which the sample video belongs, the cluster category to which the sample video belongs is determined by clustering the video semantic feature vector of the sample video; the original model includes a first branch network and a clustering module, the first branch network is used to learn the image information of the sample video and the first label, and the clustering module is used to assist the first branch network in learning the image information and the second label of the sample video; wherein the first labels of different sample videos are different, and the feature extraction model is obtained by alternately training the original model based on the image information of the sample video, and the first label and the second label of the sample video; The video clustering module is used to cluster the videos in the video set according to the video semantic feature vector of each video, and divide each video in the video set into at least one clustering category.

11. An electronic device, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the video clustering method according to any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the video clustering method according to any one of claims 1 to 9 is implemented.

13. A computer program product, characterized in that comprising computer instructions stored in a computer-readable storage medium; The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the video clustering method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Video classification method, video classification device, electronic equipment and storage medium

    CN111612093A