A video content recognition method and related device
By performing video segmentation and style vector clustering of video content, we can determine whether the video content contains irrelevant content, which solves the problem of difficulty in automatically identifying irrelevant content in video content in the prior art, and achieves high-accuracy automated recognition.
Patent Information
- Application Number
- CN202011137819.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-10-22
AI Technical Summary
The prior art is difficult to automatically identify content that is not related to video content in video content on the network, resulting in interference from users when watching videos.
By segmenting videos with the recognized video content, obtaining the style vector of each video clip, and performing similarity clustering, the style similarity between the first style cluster and the second style clustering is determined to determine whether the video content contains irrelevant content.
It realizes automatic recognition of video content, improves recognition accuracy, and reduces the possibility that users are interfered with by unrelated content.
Smart Images

Figure CN112270238B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to a video content recognition method and related device. Background Art
[0002] As a media provider, a user can provide video content on the network for sharing. For example, various videos shared by up hosts on common video platforms currently.
[0003] When editing video content, a media provider sometimes adds other content that has no association with the video content itself, such as advertisements, promotions, etc. Thus, when a user views this video content on the network, these other contents will be seen during the viewing process, resulting in the interruption of the user's viewing train of thought or causing user disgust.
[0004] Currently, mainly through manual screening of the video content provided on the network to exclude such video content with other added contents. However, the number of video contents uploaded to the network every day is very large, and manual screening alone cannot solve the problem, and users are still often disturbed by such video contents. Summary of the Invention
[0005] To solve the above technical problems, this application provides a video content recognition method and related device, realizing the automatic recognition of video content.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] On the one hand, the embodiments of this application provide a video content recognition method, and the method includes:
[0008] Performing video segmentation on the video content to be recognized to obtain a plurality of video segments;
[0009] Obtaining the style vectors respectively corresponding to the plurality of video segments;
[0010] Performing similarity clustering on the obtained style vectors to obtain a first style cluster and a second style cluster;
[0011] Determining the style similarity between the style vector corresponding to the first style cluster and the style vector corresponding to the second style cluster;
[0012] Determining whether the video content to be recognized contains content unrelated to the video content to be recognized according to the style similarity.
[0013] On the other hand, the embodiments of this application provide a video content recognition device, and the device includes a segmentation unit, an acquisition unit, a clustering unit, and a determination unit:
[0014] The segmentation unit is used to segment the video content to be recognized into multiple video segments;
[0015] The obtaining unit is used to obtain the style vectors corresponding to the multiple video segments respectively;
[0016] The clustering unit is used to perform similarity clustering on the obtained style vectors to obtain a first style cluster and a second style cluster;
[0017] The determining unit is used to determine the style similarity between the style vector corresponding to the first style cluster and the style vector corresponding to the second style cluster;
[0018] The determining unit is further used to determine whether the video content to be recognized contains content irrelevant to the video content to be recognized according to the style similarity.
[0019] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory:
[0020] The memory is used to store program codes and transmit the program codes to the processor;
[0021] The processor is used to execute the method described in the above aspect according to the instructions in the program codes.
[0022] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which is used to store a computer program, and the computer program is used to execute the method described in the above aspect.
[0023] On the other hand, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method described in the above aspect.
[0024] As can be seen from the above technical solution, the video content to be recognized is segmented into multiple video segments, and the style vectors corresponding to the multiple video segments are obtained. Then, similarity clustering is performed on the style vectors corresponding to the obtained video segments to obtain a first style cluster and a second style cluster, and the style similarity between the style vector corresponding to the first style cluster and the style vector corresponding to the second style cluster is determined. Since video content without irrelevant content generally has a unified video style, while in video content containing irrelevant content, it is generally difficult to unify the style of the irrelevant content with the video content. Therefore, the style similarity of the above two style clusters can reflect whether the overall style of the video content to be recognized is unified, so that it can be determined whether the video content to be recognized contains content irrelevant to the video content to be recognized based on this style similarity, realizing the automatic recognition of video content. Thus, based on the characteristic that the video style of video content is different from the video style of irrelevant content, the consideration of the video content itself is increased when recognizing irrelevant content, improving the recognition accuracy of video content. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 Schematic diagram of an application scenario of a video content recognition method provided by an embodiment of the present application;
[0027] Figure 2 Schematic diagram of a flow of a video content recognition method provided by an embodiment of the present application;
[0028] Figure 3 Schematic diagram of a flow of a first model training method provided by an embodiment of the present application;
[0029] Figure 4 Schematic diagram of a flow of a second model training method provided by an embodiment of the present application;
[0030] Figure 5 Schematic diagram of a flow of a style vector acquisition method provided by an embodiment of the present application;
[0031] Figure 6 Schematic diagram of a flow of another video content recognition method provided by an embodiment of the present application;
[0032] Figure 7 Schematic diagram of a video segment and segment boundary provided by an embodiment of the present application;
[0033] Figure 8 A flowchart of a method for determining content features of an n - order segment group provided by an embodiment of the present application;
[0034] Figure 9 A flowchart of a third model training method provided by an embodiment of the present application;
[0035] Figure 10 A schematic diagram of a module for video content recognition provided by an embodiment of the present application;
[0036] Figure 11 A schematic diagram of the structure of a video content recognition device provided by an embodiment of the present application;
[0037] Figure 12 A schematic diagram of the structure of a server provided by an embodiment of the present application;
[0038] Figure 13 A schematic diagram of the structure of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0039] The embodiments of the present application will be described below with reference to the accompanying drawings.
[0040] In the related art, video content on the network can be recognized based on manual screening. However, for a large amount of video content on the network, it requires a large amount of time and cost. Or, video content recognition tools can also be used to recognize potentially widespread irrelevant content in video content. However, in the process of algorithm design and model training, these tools only consider irrelevant content and do not consider the video content itself, resulting in a low recognition accuracy for whether video content includes irrelevant content.
[0041] Therefore, the embodiments of the present application provide a video content recognition method and related devices, which realize the automatic recognition of video content and improve the recognition accuracy of video content.
[0042] The video content recognition method provided by the embodiments of the present application is implemented based on artificial intelligence. Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results in theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machine to have the functions of perception, reasoning, and decision - making.
[0043] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0044] In the embodiments of this application, the artificial intelligence software technologies mainly involved include the above-mentioned directions such as computer vision technology and machine learning / deep learning. For example, it can involve image processing, image semantic understanding in computer vision, or deep learning in machine learning, including various artificial neural networks.
[0045] The video content recognition method provided in this application can be applied to video content recognition devices with data processing capabilities, such as terminal devices and servers. Among them, the terminal device can specifically be a smartphone, a computer, a personal digital assistant (PDA), a tablet computer, etc.; the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this here.
[0046] The video content recognition device can be equipped with the ability to implement computer vision technology. Computer vision is a science that studies how to enable machines to "see". Further speaking, it refers to using cameras and computers to replace human eyes for tasks such as target recognition, traceability, and measurement in machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. It also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0047] In the embodiments of the present application, the video content recognition device can process the video content to be recognized through technologies such as video processing, video semantic understanding, and video content / behavior recognition in computer vision.
[0048] The video content recognition device can be equipped with machine learning capabilities. Machine learning is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks.
[0049] In the video content recognition method provided by the embodiments of the present application, the models adopted mainly involve the application of neural networks, and the neural networks are used to recognize the irrelevant content that the video content may include.
[0050] The embodiments of the present application will be introduced below by taking a server as the video content recognition device.
[0051] See Figure 1 , Figure 1 which is a schematic diagram of the application scenario of the video content recognition method provided by the embodiments of the present application. In Figure 1In the application scenario shown, there is a server 100 for identifying whether the video content to be recognized contains content unrelated to the video content to be recognized. Among them, the content unrelated to the video content to be recognized refers to other content with a relatively low relevance to the main meaning conveyed by the video content to be recognized. For example, the digital product advertisement contained in the game video content is the content unrelated to the game video content.
[0052] As Figure 1 shown, the server 100 segments the video content 101 to be recognized into multiple video segments 102. For example, the video content to be recognized is segmented with a video playback length of 5s to obtain multiple video segments with a playback length of 5s.
[0053] Then, feature extraction is performed on the above multiple video segments 102 to obtain style vectors 103 corresponding to these multiple video segments 102 respectively. The style vector 103 can be understood as the style features obtained by extracting features from the video segments 102 from the video style dimension. Among them, the video style refers to the video quality presented during video playback, including but not limited to: painting style, picture quality, color tone, brightness, etc.
[0054] Furthermore, similarity clustering is performed on the style vectors 103 corresponding to the multiple video segments 102 respectively to obtain a first style cluster 104 and a second style cluster 105. Among them, the video segments corresponding to the style vectors included in the first style cluster 104 have similar video styles, and the video segments corresponding to the style vectors included in the second style cluster 105 have similar video styles.
[0055] Subsequently, the style similarity 106 between the style vector corresponding to the first style cluster 104 and the style vector corresponding to the second style cluster 105 is calculated. The style similarity 106 indicates the possibility that the video style corresponding to the first style cluster 104 is similar to the video style corresponding to the second style cluster 105.
[0056] Since the video content without unrelated content generally has a unified video style, and in the video content containing unrelated content, it is generally difficult to unify the style of the unrelated content with the video content. Therefore, the above style similarity 106 can reflect whether the overall style of the video content to be recognized is unified, and thus it can be determined whether the video content to be recognized contains content unrelated to the video content to be recognized according to the style similarity 106.
[0057] Based on the characteristic that the video style of the video content is different from the video style of the unrelated content, the automatic recognition of the video content is realized, and the consideration of the video content itself is increased when recognizing the unrelated content, thereby improving the recognition accuracy of the video content.
[0058] The following is a step-by-step introduction to the video content recognition method provided by the embodiments of the present application in conjunction with Figure 2 , and see Figure 2 . Figure 2 is a schematic flowchart of a video content recognition method provided by the embodiments of the present application. As Figure 2 shown, the video content recognition method includes the following steps:
[0059] S201: Segment the video content to be recognized to obtain multiple video segments.
[0060] In the embodiments of the present application, it is necessary to recognize whether the video to be recognized contains content irrelevant to the video content to be recognized, and the video to be recognized can be recognized as a normal video or an embedded video. Among them, a normal video refers to a video that does not contain content irrelevant to the video content to be recognized, and an embedded video refers to a video that contains content irrelevant to the video content to be recognized. And the content irrelevant to the video content to be recognized refers to other content with a relatively low relevance to the main meaning conveyed by the video content to be recognized.
[0061] For example, in a teaching video containing an advertisement for a teaching application, since the main meaning conveyed by the teaching video is knowledge, and the meaning conveyed by the advertisement for the teaching application is to promote and popularize the teaching application, which has a relatively low relevance to the knowledge conveyed by the teaching video. Therefore, the advertisement for the teaching application contained in the teaching video is content irrelevant to the teaching video.
[0062] Since the irrelevant content contained in the embedded video will affect the user's viewing experience of the video, therefore, in order to improve the user's viewing experience, content recognition can be performed on the embedded videos on the media platform.
[0063] During the recognition process, the video content to be recognized can be segmented first to obtain multiple video segments, so as to recognize the video content to be recognized based on the video segments. Among them, video segmentation can be performed on the video to be recognized according to the video playing duration. For example, the video to be recognized is segmented according to a video playing duration of 5s to obtain multiple video segments with a video playing duration of 5s. Video segmentation can also be performed on the video to be recognized according to the number of video frames. For example, the video to be recognized is segmented according to 240 video images to obtain multiple video segments containing 240 video images. In practical applications, the video segmentation method and the video segmentation granularity (the playing duration of the video segment or the number of video images contained in the video segment) can be set according to the actual application scenario, and no specific limitation is made here.
[0064] The above-mentioned video segmentation of the video content to be recognized is equivalent to making a refined division of the video content to be recognized. The content granularity for recognition is smaller, providing a data basis for subsequent recognition of the video content to be recognized based on video segments.
[0065] S202: Obtain the style vectors respectively corresponding to the multiple video segments.
[0066] It can be understood that for a normal video, the content it includes generally has a unified style. For example, it has a relatively unified painting style, similar picture quality, color tone, etc. However, for the video content and irrelevant content included in the embedded video, due to reasons such as content irrelevance, different providers, and different video recording methods, it is generally difficult to have a unified style.
[0067] Based on the characteristics that the above-mentioned normal video has a unified style while the embedded video is difficult to have a unified style, feature extraction can be performed on the multiple video segments obtained by the above-mentioned segmentation respectively to obtain the style vectors respectively corresponding to these multiple video segments. This style vector can be understood as the style feature obtained by extracting features from the video segment 102 from the video style dimension. Among them, the video style refers to the video quality presented when the video is played, including but not limited to: painting style, picture quality, color tone, brightness, etc.
[0068] In practical applications, a neural network model based on deep learning can be used to extract features from video segments to obtain the style vectors corresponding to the video segments.
[0069] The above-mentioned method of obtaining the style vectors respectively corresponding to multiple video segments is used to analyze whether the video content to be recognized has a unified video style based on the style vectors, so that it can be determined whether the video content to be recognized contains content irrelevant to the video content to be recognized according to the video style analysis result.
[0070] S203: Perform similarity clustering on the obtained style vectors to obtain a first style cluster and a second style cluster.
[0071] It can be understood that if the video content to be recognized does not contain content irrelevant to the video content to be recognized, it can be considered that the video content to be recognized has a unified video style. If the video content to be recognized contains content irrelevant to the content to be recognized, it can be considered that the video content to be recognized does not have a unified video style. During the recognition process, the video content to be recognized containing irrelevant content can be set to have two video styles, one video style corresponding to the substantial content included in the video content to be recognized, and one video style corresponding to the irrelevant content included in the video content to be recognized.
[0072] Therefore, after obtaining the style vectors corresponding to multiple video segments based on the above S202, these style vectors can be subjected to similarity clustering to obtain a first style cluster and a second style cluster. Among them, the video segments corresponding to the style vectors included in the first style cluster have similar video styles, and the video segments corresponding to the style vectors included in the second style cluster have similar video styles.
[0073] In practical applications, an unsupervised clustering algorithm can be used to cluster the style vectors to obtain a first style cluster and a second style cluster. In the embodiments of the present application, the method for clustering the style vectors is not limited in any way.
[0074] The above-mentioned similarity clustering of the style vectors realizes the analysis of the video style corresponding to the video content to be recognized, so as to determine whether the video to be recognized contains content irrelevant to the video content to be recognized according to the first style cluster and the second style cluster obtained by the similarity clustering.
[0075] S204: Determine the style similarity between the style vectors corresponding to the first style cluster and the style vectors corresponding to the second style cluster.
[0076] During the recognition process, it can be determined by determining the style similarity between the style vectors corresponding to the first style cluster and the style vectors corresponding to the second style cluster whether the video content to be recognized has a unified video style. Among them, the style similarity indicates the degree of similarity between the video styles of the video segments corresponding to the style vectors included in the first style cluster and the video styles of the video segments corresponding to the style vectors included in the second style cluster.
[0077] Specifically, the greater the style similarity, the greater the degree of similarity between the video styles of the video segments corresponding to the style vectors included in the first style cluster and the video styles of the video segments corresponding to the style vectors included in the second style cluster, indicating that the video content to be recognized is more likely to have a unified style. The smaller the style similarity, the smaller the degree of similarity between the video styles of the video segments corresponding to the style vectors included in the first style cluster and the video styles of the video segments corresponding to the style vectors included in the second style cluster, indicating that the video content to be recognized is less likely to have a unified style.
[0078] Based on the above, the video content corresponding to the style vectors included in the first style cluster has similar video styles, and the video content corresponding to the style vectors included in the second style cluster has similar video styles. In practical applications, the class center of the first style cluster can be used to represent the style vectors corresponding to the first style cluster, and the class center of the second style cluster can be used to represent the style vectors corresponding to the second style cluster. In addition, the style vectors included in the first style cluster can be averaged, and the style vectors included in the second style cluster can be averaged. The two style vectors obtained by averaging are used to represent the style vectors corresponding to the first style cluster and the second style cluster respectively. In practical applications, any of the above methods can be used to determine the style vectors corresponding to the first style cluster and the second style cluster, and no specific limitation is made here.
[0079] In practical applications, the style similarity between the style vectors corresponding to the first style cluster and the style vectors corresponding to the second style cluster can be determined based on a neural network model of deep learning. In the embodiments of this application, no specific limitation is made on the method for determining the style similarity.
[0080] Since video content that does not contain irrelevant content generally has a unified video style, and in video content that contains irrelevant content, it is generally difficult to unify the style of the irrelevant content and the video content. The above-mentioned style similarity indicates the degree of similarity between the video styles of the video segments corresponding to the style vectors included in the first style cluster and the video segments corresponding to the style vectors included in the second style cluster. Therefore, based on the characteristic of whether the video content to be recognized has a unified style, the style similarity corresponding to the video content to be recognized can be used to determine whether the video content to be recognized contains content irrelevant to the video content to be recognized.
[0081] It can be understood that in related technologies, methods for using video content recognition tools to recognize potentially widespread irrelevant content in video content only consider the irrelevant content and do not consider the video content itself in the process of algorithm design and model training, which may result in a low recognition accuracy for whether the video content includes irrelevant content.
[0082] The above-mentioned style similarity is determined based on the overall information included in the video content to be recognized. Subsequently, whether the video content to be recognized contains irrelevant content is recognized according to the style similarity. Compared with only considering the irrelevant content, this increases the consideration of the video content to be recognized itself and improves the recognition accuracy of the video content to be recognized.
[0083] S205: Determine whether the video content to be recognized contains content irrelevant to the video content to be recognized according to the style similarity.
[0084] In practical applications, the style similarity determined in S204 above can be compared with a set similarity threshold to determine whether the video content to be recognized contains content unrelated to the video content to be recognized. Specifically, the video to be recognized with a style similarity greater than the similarity threshold is determined as a normal video without unrelated content, and the video to be recognized with a style similarity not greater than the similarity threshold is determined as an embedded video containing unrelated content. For example, the similarity threshold is set to 0.5. In practical applications, the similarity threshold can be set according to specific application scenarios and is not limited here.
[0085] The video content recognition method provided in the above embodiments segments the video content to be recognized to obtain multiple video segments, and obtains the style vectors corresponding to the multiple video segments respectively. Then, similarity clustering is performed on the style vectors corresponding to the obtained video segments to obtain a first style cluster and a second style cluster, and the style similarity between the style vector corresponding to the first style cluster and the style vector corresponding to the second style cluster is determined. Since the video content without unrelated content generally has a unified video style, and in the video content containing unrelated content, it is generally difficult to unify the style of the unrelated content with the video content, therefore, the style similarity of the above two style clusters can reflect whether the overall style of the video content to be recognized is unified, so that it can be determined whether the video content to be recognized contains content unrelated to the video content to be recognized based on this style similarity, realizing the automatic recognition of video content. Based on the characteristic that the video style of the video content is different from the video style of the unrelated content in this way, the consideration of the video content itself is increased when recognizing the unrelated content, and the recognition accuracy of the video content is improved.
[0086] For the above process of determining the style similarity, an embodiment of the present application provides a possible implementation manner, that is, a first model is used to determine the style similarity between the style vector corresponding to the first style cluster and the style vector corresponding to the second style cluster. Among them, the first model is pre-trained, and in the embodiments of the present application, the model structure of the first model is not limited in any way.
[0087] It can be understood that applying the above first model to determine the style similarity between the style vector corresponding to the first style cluster and the style vector corresponding to the second style cluster requires pre-training the first model. For this reason, an embodiment of the present application provides a first model training method.
[0088] See Figure 3 , Figure 3 which is a schematic flow chart of a first model training method provided by an embodiment of the present application. As Figure 3 shown, the first model training method includes the following steps:
[0089] S301: Determine a training sample pair including a first sample and a second sample.
[0090] In practical applications, it is necessary to obtain a training sample set including training sample pairs. Among them, a training sample pair includes a first sample and a second sample. The first sample is the first sample video content without irrelevant content, that is, a positive sample; the second sample is the second sample video content with irrelevant content, that is, a negative sample. Among them, whether the positive and negative samples contain irrelevant content can be obtained through manual annotation.
[0091] S302: According to the style vectors of the video segments included in the first sample, determine the positive sample style similarity between the style vectors of the two style clusters of the first sample through the first initial model.
[0092] In practical applications, the same operations as S201 - S203 above can be performed, that is, segment the first sample video content to obtain multiple first sample video segments. Then, obtain the style vectors corresponding to these multiple first sample video segments respectively. Subsequently, perform similarity clustering on the style vectors corresponding to the first sample video segments respectively to obtain two style clusters of the first sample, so that the positive sample style similarity between the style vectors of these two style clusters can be determined by using the first initial model. This positive sample style similarity indicates the similarity degree between the video styles corresponding to the two style clusters of the first sample.
[0093] S303: According to the style vectors of the video segments included in the second sample, determine the negative sample style similarity between the style vectors of the two style clusters of the second sample through the first initial model.
[0094] In practical applications, the same operations as S201 - S203 above can be performed, segment the second sample video content to obtain multiple second sample video segments. Then, obtain the style vectors corresponding to these multiple second sample video segments respectively. Subsequently, perform similarity clustering on the style vectors corresponding to the second sample video segments respectively to obtain two style clusters of the second sample, so that the negative sample style similarity between the style vectors of these two style clusters can be determined by using the first initial model. This negative sample style similarity indicates the similarity degree between the video styles corresponding to the two style clusters of the second sample.
[0095] S304: Based on increasing the difference between the positive sample style similarity and the negative sample style similarity, train the first initial model to obtain the first model.
[0096] In practical applications, the above positive sample style similarity and negative sample style similarity can be used to train the first initial model. Among them, the first initial model can be a pre-constructed neural network model, and the model structure of the first initial model is not limited in any way in the embodiments of the present application.
[0097] During the training process of the first initial model, a loss function Loss can be designed to adjust the similarity threshold. In practical applications, the loss function can be the style similarity including irrelevant content or the style similarity without irrelevant content. However, the loss functions designed based on these two methods only consider the overall style characteristics of the video content from a single perspective, and the determined similarity threshold is not appropriate enough, thus affecting the recognition accuracy of whether the video to be recognized contains irrelevant content.
[0098] Therefore, the present application provides a possible implementation method, that is, based on increasing the difference between the positive sample style similarity and the negative sample style similarity, the first initial model is trained. The specific design of the corresponding loss function is:
[0099] Loss = style similarity including irrelevant content - style similarity without irrelevant content
[0100] During the actual training process, by minimizing the above Loss, the corresponding similarity threshold is determined, so as to determine whether the video content to be recognized contains irrelevant content according to the similarity threshold.
[0101] The above loss function determined based on the style similarity including irrelevant content and the style similarity without irrelevant content considers the overall style characteristics of the video content to be recognized from two perspectives, dynamically adjusts the similarity threshold according to the difference between the two style similarities, and the determined similarity threshold has high accuracy, thereby improving the recognition accuracy of the video content.
[0102] In view of the fact that the similarity threshold in S204 for determining whether the video content to be recognized contains irrelevant content is set artificially and has strong subjectivity, resulting in low recognition accuracy of the video content. In order to further improve the recognition accuracy of the video content, it can be determined whether the video content to be recognized contains irrelevant content by using the similarity threshold determined by training the above first model.
[0103] Specifically, if the style similarity corresponding to the video content to be recognized meets the similarity threshold determined by training the first model, it is determined that the video content to be recognized does not contain content irrelevant to the video content to be recognized; if the style similarity does not meet the similarity threshold, it is determined that the video content to be recognized contains content irrelevant to the video content to be recognized.
[0104] The similarity threshold determined by the above positive and negative samples is more objective than the similarity threshold set artificially, thus improving the recognition accuracy of video content.
[0105] For the style vectors in S202 above, an embodiment of the present application provides a possible implementation manner, that is, a second model is used to determine the style vectors corresponding to the multiple video segments obtained in S201 respectively. Wherein, the second model is pre-trained.
[0106] As Figure 4 shown, taking the video segment 401 as the input of the second model 402, the second model 402 is used to extract features from the video segment 401 to obtain the style vector 403 corresponding to the video segment. In practical applications, the style vector can be used as the output of the second model in the form of a vector. In the embodiment of the present application, the model structure of the second model may include a three-dimensional convolutional neural network (3D Convolutional Neural Networks, 3D-CNN), and the number of layers of the 3D-CNN and the fully connected layer can be set according to the actual scenario, which is not limited here. Among them, the 3D-CNN may include a convolutional layer, a pooling layer, and a fully connected layer, etc. In practical applications, the fully connected layer outputs the style vector and serves as the output of the first model.
[0107] It can be understood that applying the above second model to extract features of the style of the video segment and obtaining the style vector corresponding to the video segment requires pre-training the second model. For this reason, an embodiment of the present application provides a second model training method.
[0108] See Figure 5 , Figure 5 is a schematic flow chart of a second model training method provided by an embodiment of the present application. As Figure 5 shown, the second model training method includes the following steps:
[0109] S501: Obtain a video classification sample set including video classification samples.
[0110] In the embodiment of the present application, the model is trained in a supervised manner. Since the supervised training method requires samples with annotation labels, the annotation labels of the samples need to be manually annotated.
[0111] In order to avoid the cost required for manual annotation, in the embodiment of the present application, a general video classification data set can be used as the video classification sample set, and the video classification samples in the video classification sample set are used to train the model. Wherein, the video classification sample includes a sample video and a classification label corresponding to the sample video, and the classification label identifies the classification result of the sample video.
[0112] S502: Extract the style vectors of video segments in the video classification sample through the second initial model according to the video classification sample.
[0113] Before training the model, construct the second initial model. Among them, the model structure of the second initial model includes 3D-CNN and a fully connected layer. During the training process, the sample videos in the video classification sample can be segmented first to obtain multiple sample video segments. Then, using the sample video segments as the input of the second initial model, extract the style vectors of these multiple sample video segments through the second initial model.
[0114] S503: Determine the video classification result corresponding to the style vector extracted by the second initial model through the classification model.
[0115] Since the classification label corresponding to the sample video identifies the classification result of the sample video, therefore, the classification model can be used to determine the video classification result corresponding to the style vector extracted by the second initial model. Among them, use the style vector of the video segment as the input of the classification model, and use the video classification result corresponding to the video segment as the output of the classification model. In the embodiments of the present application, no limitation is imposed on the model structure of the classification model.
[0116] S504: Train the second initial model according to the classification label of the video classification sample and the video classification result to obtain the second model.
[0117] Based on the video classification result corresponding to the video classification sample determined in the above S503, adjust the model parameters of the second initial model with this video classification result and the classification label corresponding to the video classification sample, and use the converged second initial model as the above second model for obtaining the style vector corresponding to the video segment.
[0118] The above uses the general video classification sample set to train the second initial model to obtain the second model for extracting the style vector, avoiding the time and cost required for manually obtaining training samples and improving the efficiency of model training.
[0119] It can be understood that the video content recognition method provided by the above S201-S205 can identify whether the content to be recognized in the video to be recognized contains content that is not relevant to the content to be recognized in the video to be recognized. In order to further implement the localization of the content that is not relevant to the content to be recognized in the video to be recognized, an embodiment of the present application provides a possible implementation method, that is, for the embedded video identified by the above video content recognition method as containing content that is not relevant to the content to be recognized in the video to be recognized, see Figure 6 , the following steps can be executed:
[0120] S601: Determine a first video segment and a second video segment adjacent based on the segment boundary according to the segment boundaries between the multiple video segments and the playing order of the multiple video segments.
[0121] After splitting the video to be recognized into multiple video segments by performing the above S601, there are segment boundaries and a playing order between the multiple video segments. Therefore, the first video segment and the second video segment adjacent based on the segment boundary can be determined according to the segment boundaries and the playing order between the video segments.
[0122] For the multiple video segments corresponding to the video to be recognized, Figure 7 are represented by rectangles, where there is a segment boundary between every two video segments, such as Figure 7 The segment boundary is characterized by a dashed line. Taking Figure 7 the video boundary 701 in as an example, the video segments adjacent to this video boundary include a first video segment 702 and a second video segment 703, and the playing order of the first video segment 702 is prior to that of the second video segment 703.
[0123] For the first video segment and the second video segment adjacent based on the segment boundary, there are the following situations:
[0124] (1) Neither the first video segment nor the second video segment contains irrelevant content;
[0125] (2) Both the first video segment and the second video segment contain irrelevant content;
[0126] (3) The first video segment does not contain irrelevant content, and the second video segment contains irrelevant content;
[0127] (4) The first video segment contains irrelevant content, and the second video segment does not contain irrelevant content.
[0128] Regarding the positioning problem of the irrelevant content in the video to be recognized, it is to identify the boundary between the content of the video to be recognized and the irrelevant content. The segment boundaries corresponding to the above situations (3) and (4) are the boundaries between the content of the video to be recognized and the irrelevant content. Therefore, this positioning problem can be converted into an identification problem of the segment boundaries corresponding to the above situations (3) and (4).
[0129] S602: Obtain a first content feature of the first video segment and a second content feature of the second video segment.
[0130] Since the content irrelevant to the content of the video to be recognized has a relatively low degree of association with the content of the video to be recognized, the segment boundaries corresponding to the above situations (3) and (4) can be recognized based on this feature.
[0131] Based on the above S601, feature extraction is performed on the above first video clip and second video clip from the dimension of video content to obtain the first content feature of the first video clip and the second content feature of the second video clip.
[0132] In practical applications, the content feature extraction of the first video clip and the second video clip can be implemented based on a neural network model of deep learning. Among them, the content feature is used to identify the content included in the video, including but not limited to: images, audio, text, etc. The content feature can be represented in the form of a vector.
[0133] S603: Determine whether the segment boundary corresponds to the boundary of the irrelevant content according to the content similarity between the first content feature and the second content feature.
[0134] In practical applications, the content similarity between the first content feature and the second content feature can be determined. This content similarity identifies the similarity degree between the content included in the first video clip and the content included in the second video. The greater the content similarity, the greater the similarity degree between the content included in the first video clip and the content included in the second video clip, indicating that the segment boundary between the first video clip and the second video clip is more likely to be the boundary between the video content to be recognized and the irrelevant content; the smaller the content similarity, the smaller the similarity degree between the content included in the second video clip and the content included in the second video clip, indicating that the segment boundary between the first video clip and the second video clip is less likely to be the boundary between the video content to be recognized and the irrelevant content.
[0135] When determining whether the segment boundary is the boundary between the video content to be recognized and the irrelevant content, that is, determining whether the segment corresponds to the boundary of the irrelevant content, the above content similarity can be compared with a preset content threshold. If the content similarity meets the content threshold, it can be determined that the segment boundary corresponds to the boundary of the irrelevant content, that is, the segment boundary is determined to be the boundary between the video content to be recognized and the irrelevant content; if the content similarity does not meet the content threshold, it can be determined that the segment boundary does not correspond to the boundary of the irrelevant content, that is, the segment boundary is determined not to be the boundary between the video content to be recognized and the irrelevant content.
[0136] S604: If it corresponds, determine the video interval where the irrelevant content is located in the video content to be recognized according to the segment boundary.
[0137] According to the above S603, the segment boundary corresponding to the boundary of the irrelevant content can be determined. Based on this segment boundary, the video interval where the irrelevant content is located in the video content to be recognized can be determined, that is, the content irrelevant to the video content to be recognized in the video to be recognized is located.
[0138] It can be understood that in the above related technologies, to locate the irrelevant content in the video to be recognized, it is necessary to use strongly labeled data to train the model adopted by the video recognition tool. Among them, strongly labeled data means that it is necessary to not only label whether the video to be recognized contains content irrelevant to the content of the video to be recognized, but also label the video interval where the irrelevant content is located in the content of the video to be recognized. The labeling process is complex and requires high time and cost.
[0139] However, to locate the irrelevant content in the content of the video to be recognized based on the video content recognition method provided in the above embodiments, the data required is weakly labeled data, that is, it is only necessary to label whether the video to be recognized contains content irrelevant to the content of the video to be recognized, and this labeled data can be obtained by using the video content recognition method provided above, realizing the automatic recognition and location of the irrelevant content included in the video content, without manual labeling operations, reducing the input of time and cost, and improving the recognition efficiency of the video content.
[0140] It can be understood that the above determination of the video interval where the irrelevant content is located in the content of the video to be recognized is realized based on the first video segment and the second video segment. And the first video segment and the second video segment only include a small amount of information of the video to be recognized.
[0141] To further improve the accuracy of locating the irrelevant content in the content of the video to be recognized, the embodiments of the present application provide a possible implementation manner, which is specifically as follows:
[0142] If the content similarity between the first content feature and the second content feature determined in S603 above is defined as the first-order similarity, in Figure 7 where 708 is used to identify the first-order similarity between the first content feature and the second content feature, before the video content recognition method executes S603, the following steps are further included:
[0143] S605: Determine the first n-order segment group corresponding to the first video segment and the second n-order segment group corresponding to the second video segment.
[0144] The first n-order segment group includes the first video segment and n - 1 video segments adjacent to the first video segment, and the second video segment is not included in the first n-order segment group. The second n-order segment group includes the second video segment and n - 1 video segments adjacent to the second video segment, and the first video segment is not included in the second n-order segment group. Wherein, n is an integer greater than or equal to 2.
[0145] Taking Figure 7Taking the shown video clip and clip boundary as an example, for the first video clip 702 and the second video clip 703 adjacent to the clip boundary 701. If n is taken as 3, the first 3-order clip group corresponding to the first video clip 702 includes 3 video clips, namely the first video clip 702, the video clip 704, and the video clip 706, and the second 3-order clip group corresponding to the second video clip 703 includes 3 video clips, namely the second video clip 703, the video clip 705, and the video clip 707.
[0146] Similar to n = 3, for any integer n greater than or equal to 2, any first n-order clip group of the first video clip and any second n-order clip group of the second video clip can be determined. In the embodiments of the present application, n can take 2 and 3, that is, determine the first 2-order clip group, the first 3-order clip group corresponding to the first video clip, and the second 2-order clip group, the second 3-order clip group corresponding to the second video clip. In the actual application process, the value of n can be taken according to the specific application scenario, and no limitation is made here.
[0147] S606: Determine the n-order similarity between the content features of the first n-order clip group and the content features of the second n-order clip group.
[0148] In actual application, feature extraction can be performed on the n video clips in the first n-order clip group and the n video clips in the second n-order clip group respectively to obtain the content features corresponding to the n video clips in the first n-order clip group and the content features corresponding to the n video clips in the second n-order clip group. Then, the content features of the first n-order clip group can be determined according to the content features corresponding to the n video clips in the first n-order clip group, and the content features of the second n-order clip group can be determined according to the content features corresponding to the n video clips in the second n-order clip group. Perform the same operation as S608 above to determine the n-order similarity between the content features of the first n-order clip group and the content features of the second n-order clip group.
[0149] For determining the content features of the first n-order clip group and the content features of the second n-order clip group, in a possible implementation manner, the content features corresponding to the n video clips in the first n-order clip group can be summed to obtain the content features of the first n-order clip group; the content features corresponding to the n video clips in the second n-order clip group can be summed to obtain the content features of the second n-order clip group.
[0150] In actual application, a multi-layer perceptron can be used to implement the summation operation of the content features corresponding to the n video clips in the n-order clip group. As Figure 8As shown, taking n = 3 as an example, the content features 801 corresponding to the three video segments in the third-order segment group are used as the input of the multi-layer perceptron 802. The multi-layer perceptron 802 averages the three content features 801 and outputs the content feature 803 of the third-order segment group.
[0151] Based on the above process, the content similarity between the first n-order segment group and the second n-order segment group can be determined according to the obtained content features of the first n-order segment group and the second n-order segment group. The process of determining the content similarity is as described in S608 and will not be elaborated here.
[0152] In the embodiment of the present application, if n takes values of 2 and 3, the second-order similarity between the content features of the first second-order segment group and the second second-order segment group obtained according to the above process is Figure 7 In, the second-order similarity is identified by 709, and according to the obtained third-order similarity between the content features of the first third-order segment group and the second third-order segment group, in Figure 7 In, the third-order similarity is identified by 710.
[0153] Therefore, in the above S603, determining whether the segment boundary corresponds to the boundary of the irrelevant content can be achieved according to the first-order similarity and the n-order similarity between the first content feature and the second content feature. In practical applications, the first-order similarity and the n-order similarity can be averaged, and it is determined whether the obtained average similarity meets the content threshold, so as to determine whether the segment boundary corresponds to the boundary of the irrelevant content.
[0154] Determining whether the segment boundary corresponds to the boundary of the irrelevant content based on the first-order similarity and the n-order similarity increases the video content on which the correspondence between the segment boundary and the boundary of the irrelevant content is determined compared to only using the first-order similarity between two video segments, thereby improving the positioning accuracy of the irrelevant content in the video content to be recognized.
[0155] For the content features of video segments, the embodiment of the present application provides a possible implementation manner, that is, the content features of video segments are determined through a third model. The third model is pre-trained.
[0156] In the embodiment of the present application, the foregoing acquisition of the style vector and content features of video segments is a process of feature extraction from video content. The difference between the style vector extraction process and the content feature extraction process lies only in the difference in feature extraction dimensions. In practical applications, the model structure of the third model can adopt the same model structure as the second model, including 3D-CNN and fully connected layers, and different model parameters are set for the third model to achieve the extraction of the content features of video segments.
[0157] It can be understood that in order to extract the content features of a video clip by applying the above-mentioned third model and obtain the content features corresponding to the video clip, the third model needs to be trained in advance. For this purpose, an embodiment of the present application provides a method for training the third model.
[0158] See Figure 9 , Figure 9 which is a schematic flowchart of a method for training the third model provided by an embodiment of the present application. As Figure 9 shown, the method for training the third model includes the following steps:
[0159] S901: Obtain a video classification sample set including video classification samples.
[0160] S902: According to the video classification samples, extract the content features of the video clips in the video classification samples through the second initial model.
[0161] During the training of the second initial model, the sample videos in the video classification samples can be segmented first to obtain multiple sample video clips. Then, using the sample video clips as the input of the second initial model, the content features of these multiple sample video clips are extracted through the second initial model. Here, the model structure of the second initial model is the same as that of the second initial model in S502 described above, but the model parameters are different.
[0162] S903: Determine the video classification results corresponding to the content features extracted by the second initial model through a classification model.
[0163] In an embodiment of the present application, the content features of the video clip are used as the input of the classification model, and the video classification results corresponding to the video clip are used as the output of the classification model. Among them, the classification model is the same as the classification model in S503 described above.
[0164] S904: Train the second initial model according to the classification labels of the video classification samples and the video classification results to obtain the third model.
[0165] Based on the video classification results corresponding to the video classification samples determined in S903 above, adjust the model parameters of the second initial model with the video classification results and the classification labels corresponding to the video classification samples, and use the converged second initial model as the third model for obtaining the content features corresponding to the video clip.
[0166] The above-mentioned training of the second initial model using a general video classification sample set to obtain a third model for extracting content features avoids the time and cost required for manually obtaining training samples and improves the efficiency of model training. Moreover, the second model and the third model adopt the same model structure, avoiding the process of repeated modeling and further improving the efficiency of model training.
[0167] It can be understood that for an embedded video containing content unrelated to the video content, the user experience is poor. Therefore, for the videos to be recognized published on various video platforms, the video intervals where the unrelated content is located in the videos to be recognized can be recognized and located based on the video content recognition method provided in the above embodiments, so as to remove the unrelated content in the content of the videos to be recognized, thereby improving the user experience and enabling media publishers to only publish unrelated content through the video platform, thus increasing the revenue of the video platform.
[0168] For the video content recognition method provided in the above embodiments, in practical applications, each execution step can be integrated into different modules, and the recognition of video content can be achieved through the modules.
[0169] For better understanding, the following combines Figure 10 , and an exemplary description of the application process of video content recognition provided in the embodiments of the present application is given.
[0170] See Figure 10 , Figure 10 is a schematic diagram of a module for video content recognition provided in the embodiments of the present application. As Figure 10 shown, it includes 5 modules: Module 1001, Module 1002, Module 1003, Module 1004, and Module 1005.
[0171] Among them, the first model described above is deployed in Module 1003, and the second initial model and the classification model described above are deployed in Module 1004. Before calling Module 1001, by calling Module 1004, the second initial model is trained with two different model parameters using the classification model to obtain the second model and the third model respectively, and they are deployed in Module 1001. For the sake of distinction, Module 1001 using the second model is denoted as Module 1001-A, and Module 1001 using the third model is denoted as Module 100-B. Figure 10 Module 1004 is not shown in
[0172] In practical applications, by invoking Module 1 (1001 - A), feature extraction is performed on multiple video segments obtained after video segmentation of the video content to be recognized using the second model, and the style vectors of the video segments are obtained. Then, by invoking Module 3 (1003), based on the style vectors obtained by Module 1 (1001 - A), the style similarity of the video to be recognized is determined. Subsequently, by invoking Module 5 (1005), based on the style similarity determined by Module 3 (1003), it is determined whether the video to be recognized contains content unrelated to the video content to be recognized.
[0173] Furthermore, for the embedded video determined to contain unrelated content as described above, Module 1 (1001 - B) is invoked again to obtain the content features of the above - mentioned video segments using the third model. Then, by invoking Module 2 (1002), based on the content features obtained by Module 1 (1001 - B), the content similarity of the video to be recognized is determined. Subsequently, by invoking Module 5 (1005), based on the content similarity determined by Module 2 (1002), the video interval where the unrelated content is located in the embedded video is determined.
[0174] The above - mentioned implementation through multiple associated modules realizes the recognition and positioning of unrelated content in video content, realizes the automatic recognition of video content, and improves the recognition efficiency and accuracy of video content.
[0175] For the video content recognition method provided in the above - mentioned embodiment, the embodiment of the present application also provides a video content recognition device.
[0176] See Figure 11 , Figure 11 which is a schematic structural diagram of a video content recognition device provided by an embodiment of the present application. As Figure 11 shown, the video content recognition device 1100 includes a segmentation unit 1101, an acquisition unit 1102, a clustering unit 1103, and a determination unit 1104:
[0177] The segmentation unit 1101 is configured to perform video segmentation on the video content to be recognized, obtaining multiple video segments;
[0178] The acquisition unit 1102 is configured to obtain the style vectors respectively corresponding to the multiple video segments;
[0179] The clustering unit 1103 is configured to perform similarity clustering on the obtained style vectors, obtaining a first style cluster and a second style cluster;
[0180] The determination unit 1104 is configured to determine the style similarity between the style vector corresponding to the first style cluster and the style vector corresponding to the second style cluster;
[0181] The determining unit 1104 is further configured to determine whether the video content to be recognized contains content unrelated to the video content to be recognized according to the style similarity.
[0182] In a possible implementation manner, the determining unit 1104 is configured to determine the style similarity between the style vectors corresponding to the first style clustering and the style vectors corresponding to the second style clustering through a first model;
[0183] The determining unit 1104 is further configured to:
[0184] Determine a training sample pair including a first sample and a second sample, where the first sample is the first sample video content that does not contain unrelated content, and the second sample is the second sample video content that contains unrelated content;
[0185] According to the style vectors of the video segments included in the first sample, determine the positive sample style similarity between the style vectors of the two style clusterings of the first sample through a first initial model;
[0186] According to the style vectors of the video segments included in the second sample, determine the negative sample style similarity between the style vectors of the two style clusterings of the second sample through the first initial model;
[0187] The apparatus further includes a training unit:
[0188] The training unit is configured to train the first initial model based on increasing the difference between the positive sample style similarity and the negative sample style similarity to obtain the first model.
[0189] In a possible implementation manner, the determining unit 1104 is configured to:
[0190] If the style similarity meets the similarity threshold, determine that the video content to be recognized does not contain content unrelated to the video content to be recognized;
[0191] If the style similarity does not meet the similarity threshold, determine that the video content to be recognized contains content unrelated to the video content to be recognized;
[0192] Wherein the similarity threshold is determined by training the first model.
[0193] In a possible implementation manner, if the determining unit 1104 determines that the video content to be recognized contains content unrelated to the video content to be recognized according to the style similarity:
[0194] The determining unit 1104 is further configured to determine a first video segment and a second video segment adjacent based on the segment boundary according to the segment boundary between the multiple video segments and the playing order of the multiple video segments;
[0195] The obtaining unit 1102 is further configured to obtain a first content feature of the first video segment and a second content feature of the second video segment;
[0196] The determining unit 1104 is further configured to:
[0197] Determine whether the segment boundary corresponds to the boundary of the irrelevant content according to the content similarity between the first content feature and the second content feature;
[0198] If corresponding, determine the video interval where the irrelevant content is located in the video content to be recognized according to the segment boundary.
[0199] In a possible implementation manner, the content similarity between the first content feature and the second content feature is a first-order similarity, and the determining unit 1104 is further configured to:
[0200] Determine a first n-order segment group corresponding to the first video segment and a second n-order segment group corresponding to the second video segment; where n is an integer not less than 2;
[0201] The first n-order segment group includes the first video segment and n - 1 video segments adjacent to the first video segment, and the second video segment is not included in the first n-order segment group;
[0202] The second n-order segment group includes the second video segment and n - 1 video segments adjacent to the second video segment, and the first video segment is not included in the second n-order segment group;
[0203] Determine the n-order similarity between the content feature of the first n-order segment group and the content feature of the second n-order segment group;
[0204] The determining unit 1104 is configured to determine whether the segment boundary corresponds to the boundary of the irrelevant content according to the first-order similarity and the n-order similarity between the first content feature and the second content feature.
[0205] In a possible implementation manner, the style vectors of the multiple video segments are determined according to a second model, the content features of the multiple video segments are determined according to a third model, and the second model and the third model are trained according to the same second initial model;
[0206] The obtaining unit 1102 is further configured to obtain a video classification sample set including video classification samples;
[0207] The apparatus includes a style vector extraction unit, a content feature extraction unit, and a training unit:
[0208] The style vector extraction unit is configured to extract a style vector of a video segment in the video classification sample according to the video classification sample through the second initial model;
[0209] The determining unit 1104 is further configured to determine a video classification result corresponding to the style vector extracted by the second initial model through a classification model;
[0210] The training unit is configured to train the second initial model according to the classification label of the video classification sample and the video classification result to obtain the second model;
[0211] The content feature extraction unit is configured to extract a content feature of a video segment in the video classification sample according to the video classification sample through the second initial model;
[0212] The determining unit 1104 is further configured to determine a video classification result corresponding to the content feature extracted by the second initial model through a classification model;
[0213] The training unit is further configured to train the second initial model according to the classification label of the video classification sample and the video classification result to obtain the third model.
[0214] The video recognition content device provided in the above embodiment segments the video content to be recognized to obtain a plurality of video segments, and obtains style vectors corresponding to the plurality of video segments respectively. Then, similarity clustering is performed on the obtained style vectors corresponding to the video segments to obtain a first style cluster and a second style cluster, and the style similarity between the style vectors corresponding to the first style cluster and the style vectors corresponding to the second style cluster is determined. Since video content that does not contain irrelevant content generally has a unified video style, and in video content that contains irrelevant content, it is generally difficult for the irrelevant content to be unified with the video style, therefore, the style similarity of the foregoing two style clusters can reflect whether the overall style of the video content to be recognized is unified, so that it can be determined whether the video content to be recognized contains content irrelevant to the video content to be recognized according to the style similarity, realizing the automatic recognition of video content. Thus, based on the characteristic that the video style of video content is different from the video style of irrelevant content, the consideration of the video content itself is increased when recognizing irrelevant content, and the recognition accuracy of video content is improved.
[0215] The embodiments of the present application also provide a computer device. The computer device for video content recognition provided by the embodiments of the present application will be introduced from the perspective of hardware implementation below.
[0216] Refer to Figure 12 , Figure 12 , which is a schematic structural diagram of a server provided by the embodiments of the present application. The server 1400 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 1422 (for example, one or more processors) and a memory 1432, and one or more storage media 1430 (for example, one or more mass storage devices) for storing application programs 1442 or data 1444. Among them, the memory 1432 and the storage media 1430 may be transient storage or persistent storage. The program stored in the storage media 1430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1422 may be configured to communicate with the storage media 1430 and execute a series of instruction operations in the storage media 1430 on the server 1400.
[0217] The server 1400 may further include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input / output interfaces 1458, and / or one or more operating systems 1441, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0218] The steps executed by the server in the above embodiments may be based on the Figure 12 server structure shown.
[0219] Among them, the CPU 1422 is used to execute the following steps:
[0220] Perform video segmentation on the video content to be recognized to obtain multiple video segments;
[0221] Obtain the style vectors corresponding to the multiple video segments respectively;
[0222] Perform similarity clustering on the obtained style vectors to obtain a first style cluster and a second style cluster;
[0223] Determine the style similarity between the style vector corresponding to the first style cluster and the style vector corresponding to the second style cluster;
[0224] Determine whether the video content to be recognized contains content unrelated to the video content to be recognized according to the style similarity.
[0225] Optionally, the CPU 1422 can also execute the method steps of any specific implementation manner of the video content recognition method in the embodiments of the present application.
[0226] For the video content recognition method described above, the embodiments of the present application further provide a terminal device for video content recognition to implement and apply the above video content recognition method in practice.
[0227] See Figure 13 , Figure 13 FIG. is a schematic structural diagram of a terminal device provided by an embodiment of the present application. For ease of explanation, only parts related to the embodiments of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal device can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (Personal Digital Assistant, abbreviated as PDA), etc. Taking the terminal device as a mobile phone as an example:
[0228] Figure 13 FIG. shows a block diagram of a part of the structure of a mobile phone related to the terminal device provided by the embodiments of the present application. Refer to Figure 13 , the mobile phone includes: a radio frequency (RF) circuit 1510, a memory 1520, an input unit 1530, a display unit 1540, a sensor 1550, an audio circuit 1560, a wireless fidelity (WiFi) module 1570, a processor 1580, and a power supply 1590, etc. Those skilled in the art can understand that Figure 13 the structure of the mobile phone shown in FIG. does not constitute a limitation on the mobile phone, and may include more or fewer components than shown, or combine some components, or have different component arrangements.
[0229] Next, in conjunction with Figure 13 each component of the mobile phone will be specifically introduced:
[0230] The RF circuit 1510 can be used for receiving and transmitting information or signals during communication. Specifically, after receiving the downlink information from the base station, it is sent to the processor 1580 for processing. Additionally, the uplink data designed is sent to the base station. Generally, the RF circuit 1510 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 1510 can also communicate with the network and other devices via wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0231] The memory 1520 can be used to store software programs and modules. The processor 1580 realizes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1520. The memory 1520 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.), etc.; the data storage area can store the data created according to the use of the mobile phone (such as audio data, phone book, etc.), etc. In addition, the memory 1520 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0232] The input unit 1530 can be used to receive input digital or character information and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 1530 can include a touch panel 1531 and other input devices 1532. The touch panel 1531, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch panel 1531), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 1531 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1580, and can receive and execute the commands sent by the processor 1580. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 1531. In addition to the touch panel 1531, the input unit 1530 can also include other input devices 1532. Specifically, the other input devices 1532 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0233] The display unit 1540 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 1540 can include a display panel 1541. Optionally, the display panel 1541 can be configured in forms such as a liquid crystal display (LCD) and an organic light-emitting diode (OLED). Further, the touch panel 1531 can cover the display panel 1541. When the touch panel 1531 detects a touch operation thereon or nearby, it transmits it to the processor 1580 to determine the type of touch event. Subsequently, the processor 1580 provides corresponding visual output on the display panel 1541 according to the type of touch event. Although in Figure 13 the touch panel 1531 and the display panel 1541 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 1531 and the display panel 1541 can be integrated to realize the input and output functions of the mobile phone.
[0234] The mobile phone may further include at least one sensor 1550, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1541 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1541 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the mobile phone can also be configured with, they will not be elaborated here.
[0235] The audio circuit 1560, the speaker 1561, and the microphone 1562 can provide an audio interface between the user and the mobile phone. The audio circuit 1560 can transmit the electrical signal converted from the received audio data to the speaker 1561, and the speaker 1561 converts it into a sound signal for output; on the other hand, the microphone 1562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1560 and converted into audio data, and then the audio data is output to the processor 1580 for processing, and then sent to another mobile phone through the RF circuit 1510, for example, or the audio data is output to the memory 1520 for further processing.
[0236] WiFi belongs to short-distance wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 1570, which provides users with wireless broadband Internet access. Although Figure 13 the WiFi module 1570 is shown, it can be understood that it does not belong to the essential composition of the mobile phone and can be omitted completely within the scope of not changing the essence of the invention according to needs.
[0237] The processor 1580 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 1520, and calling the data stored in the memory 1520, it executes various functions of the mobile phone and processes data. Optionally, the processor 1580 may include one or more processing units; preferably, the processor 1580 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1580 either.
[0238] The mobile phone further includes a power source 1590 (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the processor 1580 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system.
[0239] Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated herein.
[0240] In the embodiment of the present application, the memory 1520 included in the mobile phone can store program codes and transmit the program codes to the processor.
[0241] The processor 1580 included in the mobile phone can execute the video content recognition method provided in the above embodiment according to the instructions in the program codes.
[0242] The embodiment of the present application further provides a computer-readable storage medium for storing a computer program, and the computer program is used to execute the video content recognition method provided in the above embodiment.
[0243] The embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the video content recognition method provided in various optional implementation manners of the above aspects.
[0244] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiment can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps included in the above method embodiment; and the foregoing storage medium can be at least one of the following media: read-only memory (English: read-only memory, abbreviation: ROM), RAM, magnetic disk, or optical disc, etc., which can store program codes.
[0245] It should be noted that the various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the corresponding parts in the method embodiments. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0246] As described above, it is only a specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A video content recognition method, characterized in that The method includes: Performing video segmentation on the video content to be recognized to obtain multiple video segments; Obtaining style vectors respectively corresponding to the multiple video segments; the style vector is a style feature obtained by extracting features from the video segments in terms of video style dimension; Performing similarity clustering on the obtained style vectors to obtain a first style cluster and a second style cluster; the video segments corresponding to the style vectors included in the first style cluster have similar video styles, and the video segments corresponding to the style vectors included in the second style cluster have similar video styles; Determining the style similarity between the style vector corresponding to the first style cluster and the style vector corresponding to the second style cluster; the style similarity is used to characterize whether the overall style of the video content to be recognized is unified; Determining whether the video content to be recognized contains content irrelevant to the video content to be recognized according to the style similarity, wherein if the style similarity characterizes that the overall style of the video content to be recognized is unified, it indicates that the video content to be recognized does not contain content irrelevant to the video content to be recognized, otherwise, it indicates that the video content to be recognized contains content irrelevant to the video content to be recognized.
2. The method according to claim 1, characterized in that The determining the style similarity between the style vector corresponding to the first style cluster and the style vector corresponding to the second style cluster includes: Determining the style similarity between the style vector corresponding to the first style cluster and the style vector corresponding to the second style cluster through a first model; The first model is trained in the following manner: Determining a training sample pair including a first sample and a second sample, the first sample being a first sample video content without irrelevant content, and the second sample being a second sample video content with irrelevant content; According to the style vectors of the video segments included in the first sample, determining the positive sample style similarity between the style vectors of the two style clusters of the first sample through a first initial model; According to the style vectors of the video segments included in the second sample, determining the negative sample style similarity between the style vectors of the two style clusters of the second sample through the first initial model; Training the first initial model based on increasing the difference between the positive sample style similarity and the negative sample style similarity to obtain the first model.
3. The method according to claim 2, characterized in that, The determining whether the video content to be recognized contains content irrelevant to the video content to be recognized according to the style similarity includes: If the style similarity meets the similarity threshold, determining that the video content to be recognized does not contain content irrelevant to the video content to be recognized; If the style similarity does not meet the similarity threshold, determining that the video content to be recognized contains content irrelevant to the video content to be recognized; Wherein the similarity threshold is determined by training the first model.
4. The method according to any one of claims 1 to 3, characterized in that, If it is determined according to the style similarity that the video content to be recognized contains content irrelevant to the video content to be recognized, the method further includes: Determine a first video segment and a second video segment adjacent based on the segment boundary among the multiple video segments and the playing order of the multiple video segments; Obtain a first content feature of the first video segment and a second content feature of the second video segment; Determine whether the segment boundary corresponds to the boundary of the irrelevant content according to the content similarity between the first content feature and the second content feature; If corresponding, determine the video interval where the irrelevant content is located in the video content to be recognized according to the segment boundary.
5. The method according to claim 4, characterized in that, The content similarity between the first content feature and the second content feature is a first-order similarity, and the method further includes: Determine a first n-order segment group corresponding to the first video segment and a second n-order segment group corresponding to the second video segment; where n is an integer not less than 2; Wherein, the first n-order segment group includes the first video segment and n - 1 video segments adjacent to the first video segment, and the second video segment is not included in the first n-order segment group; The second n-order segment group includes the second video segment and n - 1 video segments adjacent to the second video segment, and the first video segment is not included in the second n-order segment group; Determine the n-order similarity between the content feature of the first n-order segment group and the content feature of the second n-order segment group; The determining whether the segment boundary corresponds to the boundary of the irrelevant content according to the content similarity between the first content feature and the second content feature includes: Determine whether the segment boundary corresponds to the boundary of the irrelevant content according to the first-order similarity and the n-order similarity between the first content feature and the second content feature.
6. The method according to claim 4, wherein The style vectors of the multiple video segments are determined according to a second model, and the content features of the multiple video segments are determined according to a third model. The second model and the third model are trained from the same second initial model; The training method of the second model is as follows: Obtain a video classification sample set including video classification samples; Extract the style vectors of the video segments in the video classification samples through the second initial model according to the video classification samples; Determine the video classification result corresponding to the style vector extracted by the second initial model through a classification model; Train the second initial model according to the classification label of the video classification sample and the video classification result to obtain the second model; The training method of the third model is as follows: Obtain a video classification sample set including video classification samples; Extract the content features of the video segments in the video classification samples through the second initial model according to the video classification samples; Determine the video classification result corresponding to the content feature extracted by the second initial model through a classification model; Train the second initial model according to the classification label of the video classification sample and the video classification result to obtain the third model.
7. The method according to claim 4, wherein The method further includes: Remove the irrelevant content from the video content to be recognized according to the video interval.
8. A video content recognition device, characterized in that, The device includes a segmentation unit, an acquisition unit, a clustering unit, and a determination unit: The segmentation unit is configured to segment the video content to be recognized into multiple video segments; The acquisition unit is configured to acquire style vectors corresponding to the multiple video segments respectively; the style vector is a style feature obtained by extracting features from the video segments in terms of video style dimension; The clustering unit is configured to perform similarity clustering on the acquired style vectors to obtain a first style cluster and a second style cluster; the video segments corresponding to the style vectors included in the first style cluster have similar video styles, and the video segments corresponding to the style vectors included in the second style cluster have similar video styles; The determination unit is configured to determine the style similarity between the style vector corresponding to the first style cluster and the style vector corresponding to the second style cluster; the style similarity is used to characterize whether the overall style of the video content to be recognized is unified; The determination unit is further configured to determine whether the video content to be recognized contains content irrelevant to the video content to be recognized according to the style similarity, wherein if the style similarity characterizes that the overall style of the video content to be recognized is unified, it indicates that the video content to be recognized does not contain content irrelevant to the video content to be recognized, otherwise, it indicates that the video content to be recognized contains content irrelevant to the video content to be recognized.
9. The device according to claim 8, characterized in that, The determination unit is configured to determine the style similarity between the style vector corresponding to the first style cluster and the style vector corresponding to the second style cluster through a first model; The determination unit is further configured to: Determine a training sample pair including a first sample and a second sample, where the first sample is a first sample video content without irrelevant content, and the second sample is a second sample video content with irrelevant content; According to the style vectors of the video segments included in the first sample, determine the positive sample style similarity between the style vectors of the two style clusters of the first sample through a first initial model; According to the style vectors of the video segments included in the second sample, determine the negative sample style similarity between the style vectors of the two style clusters of the second sample through the first initial model; The device further includes a training unit: The training unit is configured to train the first initial model based on increasing the difference between the positive sample style similarity and the negative sample style similarity to obtain the first model.
10. The device according to claim 9, characterized in that, The determination unit is configured to: If the style similarity meets the similarity threshold, determine that the video content to be recognized does not contain content irrelevant to the video content to be recognized; If the style similarity does not meet the similarity threshold, determine that the video content to be recognized contains content irrelevant to the video content to be recognized; Wherein the similarity threshold is determined by training the first model.
11. The device according to any one of claims 8-10, characterized in that, If the determination unit determines that the video content to be recognized contains content irrelevant to the video content to be recognized according to the style similarity: The determining unit is further configured to determine a first video segment and a second video segment adjacent based on the segment boundary according to the segment boundary between the multiple video segments and the playing order of the multiple video segments; The obtaining unit is further configured to obtain a first content feature of the first video segment and a second content feature of the second video segment; The determining unit is further configured to: Determine whether the segment boundary corresponds to the boundary of the irrelevant content according to the content similarity between the first content feature and the second content feature; If corresponding, determine the video interval where the irrelevant content is located in the video content to be recognized according to the segment boundary.
12. The device according to claim 11, wherein, The content similarity between the first content feature and the second content feature is the first-order similarity, and the determining unit is further configured to: Determine a first n-order segment group corresponding to the first video segment and a second n-order segment group corresponding to the second video segment; where n is an integer not less than 2; Wherein, the first n-order segment group includes the first video segment and n-1 video segments adjacent to the first video segment, and the second video segment is not included in the first n-order segment group; The second n-order segment group includes the second video segment and n-1 video segments adjacent to the second video segment, and the first video segment is not included in the second n-order segment group; Determine the n-order similarity between the content feature of the first n-order segment group and the content feature of the second n-order segment group; The determining unit is configured to determine whether the segment boundary corresponds to the boundary of the irrelevant content according to the first-order similarity and the n-order similarity between the first content feature and the second content feature.
13. The device according to claim 11, characterized in that, The style vectors of the multiple video segments are determined according to a second model, the content features of the multiple video segments are determined according to a third model, and the second model and the third model are trained from the same second initial model; The obtaining unit is further configured to obtain a video classification sample set including video classification samples; The apparatus includes a style vector extraction unit, a content feature extraction unit, and a training unit: The style vector extraction unit is configured to extract the style vector of the video segment in the video classification sample through the second initial model according to the video classification sample; The determining unit is further configured to determine the video classification result corresponding to the style vector extracted by the second initial model through a classification model; The training unit is configured to train the second initial model according to the classification label of the video classification sample and the video classification result to obtain the second model; The content feature extraction unit is configured to extract the content feature of the video segment in the video classification sample through the second initial model according to the video classification sample; The determining unit is further configured to determine the video classification result corresponding to the content feature extracted by the second initial model through a classification model; The training unit is further configured to train the second initial model according to the classification label of the video classification sample and the video classification result to obtain the third model.
14. A computer device, characterized in that, The computer device includes a processor and a memory: The memory is used to store program codes and transmit the program codes to the processor; The processor is used to execute the method according to any one of claims 1-7 based on the instructions in the program codes.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1-7.
16. A computer program product, characterized in that, The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method according to any one of claims 1-7.
Citation Information
Patent Citations
Transformer substation personnel behavior recognition method based on monitoring video time sequence action positioning and anomaly detection
CN111291699A