Live content understanding method, device, storage medium and electronic device

By performing equal-time slices and feature extraction of live broadcast content on the live broadcast content of online videos, combined with classification and clustering strategies, the problems of difficult and inaccurate recommendations in the existing technology are solved, and higher content understanding accuracy and user needs satisfaction are achieved.

CN115665441BActive Publication Date: 2025-05-23BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211282884.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-19
Publication Date
2025-05-23
Estimated Expiration
2042-10-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively understand and distribute online video live content that meets user needs, especially in different implementation scenarios, resulting in insufficient accuracy and applicability of content recommendations.

Method used

By performing isochronous slice processing on the live video stream to be understood, the feature vectors of the video clips are extracted, and the content understanding results are determined using preset classification strategies or clustering strategies. Classification strategies use category tags for supervised learning, while clustering strategies use clustering algorithms for unsupervised processing.

Benefits of technology

It improves the accuracy of live broadcast content understanding, can better meet the needs of different users, and is suitable for different implementation scenarios, reduces user differences, and improves the accuracy of content recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115665441B_ABST
    Figure CN115665441B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, device, storage medium and electronic device for understanding the content of a live broadcast, the method comprising: obtaining a video stream of a live broadcast to be understood; slicing the video stream with equal length to obtain multiple video segments; extracting a feature vector of at least one video segment among the multiple video segments; and determining a content understanding result of the live broadcast to be understood by using a preset strategy based on the feature vector of at least one video segment; wherein the preset strategy comprises a classification strategy or a clustering strategy, the classification strategy is used to determine the content understanding result of the live broadcast to be understood by at least using a category label, the category label is obtained by predicting the feature vector of the video segment using a label prediction model obtained by training with manually annotated training samples, the clustering strategy is used to determine the content understanding result of the live broadcast to be understood by using a clustering algorithm to perform clustering processing based on the feature vectors of all video segments, the accuracy of content understanding can meet the needs of different landing scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of electronic information technology, and in particular, to a live content understanding method, device, storage medium and electronic device. Background Art

[0002] Live streaming means that users can watch live audio and video events happening remotely through the Internet. It has highly interactive product attributes, which makes live streaming have strong social functions and high product stickiness. The content of live streaming is rich and varied, and different scenarios are implemented, and different users have different needs, which makes the distribution and recommendation of live streaming very important.

[0003] However, the related art makes it difficult to understand the content of live online videos, which results in an inability to distribute and recommend live online videos that meet the needs of users, and is also unable to be applied to different landing scenarios. Summary of the invention

[0004] This summary is provided to introduce concepts in a brief form that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] In a first aspect, the present disclosure provides a method for understanding live content, comprising:

[0006] Get the video stream to be understood.

[0007] Slicing the video stream with equal length to obtain multiple video segments;

[0008] extracting a feature vector of at least one video segment among the plurality of video segments;

[0009] Determining a content understanding result of the live broadcast to be understood by adopting a preset strategy according to the feature vector of the at least one video segment;

[0010] Among them, the preset strategy includes a classification strategy or a clustering strategy. The classification strategy is used to determine the content understanding result of the live broadcast to be understood by at least using category labels. The category labels are obtained by predicting the feature vectors of the video clips using a label prediction model trained with manually annotated training samples. The clustering strategy is used to determine the content understanding result of the live broadcast to be understood by clustering the feature vectors of all the video clips using a clustering algorithm.

[0011] In a second aspect, the present disclosure provides a live content understanding device, comprising:

[0012] An acquisition module is used to acquire the video stream to be broadcasted;

[0013] A slicing module, used for slicing the video stream with equal length to obtain multiple video segments;

[0014] An extraction module, configured to extract a feature vector of at least one video segment among the plurality of video segments;

[0015] A determination module, configured to determine a content understanding result of the live broadcast to be understood by adopting a preset strategy according to a feature vector of the at least one video segment;

[0016] Among them, the preset strategy includes a classification strategy or a clustering strategy. The classification strategy is used to determine the content understanding result of the live broadcast to be understood by at least using category labels. The category labels are obtained by predicting the feature vectors of the video clips using a label prediction model trained with manually annotated training samples. The clustering strategy is used to determine the content understanding result of the live broadcast to be understood by clustering the feature vectors of all the video clips using a clustering algorithm.

[0017] In a third aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.

[0018] In a fourth aspect, the present disclosure provides an electronic device, including:

[0019] a storage device having a computer program stored thereon;

[0020] A processing device is used to execute the computer program in the storage device to implement the steps of the method described in the first aspect.

[0021] Through the above technical solution, the video stream of the live broadcast to be understood is divided into multiple video segments of equal length, and the content understanding result of the live broadcast to be understood is determined based on the feature vector of at least one video segment. Since the video segments are short and the amount of content information is small, it is not easy to cause interference, so it is not easy to cause disagreements among different users, thereby improving the accuracy of content understanding; in addition, the classification strategy is a supervised method, and the clustering strategy is an unsupervised method. The two strategies have different advantages, so they can meet the needs of different landing scenarios.

[0022] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale. In the drawings:

[0024] Figure 1 The figure is a flowchart of a method for understanding live content according to an exemplary embodiment of the present disclosure.

[0025] Figure 2 The flowchart of determining a feature vector of a video clip according to an exemplary embodiment of the present disclosure is shown.

[0026] Figure 3 The figure is a block diagram of a device for understanding live content according to an exemplary embodiment of the present disclosure.

[0027] Figure 4 It is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0029] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0030] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0031] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0032] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0033] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0034] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0035] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0036] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0037] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0038] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0039] As mentioned in the background technology, the duration of online video live broadcast is long, and the content is very difficult to understand. For example, for a 1-hour video, 10 minutes of live broadcast content is the host dancing, 30 minutes of live broadcast content is the host singing, and the rest of the live broadcast content is the host interacting with the users watching the live broadcast. For different types of users who operate or pay attention to dancing, singing, and language interaction, the standards for understanding the live content are different. Therefore, different users locate the category of the live video differently, so the prediction effect of the classification model obtained by training the classification model using manually labeled video samples will be affected, especially for the landing scenarios of refined operations and ecological monitoring types, and the prediction effect will directly affect the distribution and recommendation of the live broadcast room.

[0040] In addition, for online video live streaming, the current method of understanding content cannot be applied to different landing scenarios.

[0041] In view of this, the embodiments of the present disclosure disclose a live content understanding method, device, storage medium and electronic device to improve the accuracy of content understanding and meet the needs of different landing scenarios.

[0042] The present disclosure is further explained below in conjunction with the accompanying drawings.

[0043] Figure 1 is a flowchart of a method for understanding live content according to an exemplary embodiment of the present disclosure. The method for understanding live content can be applied to an electronic device, and the electronic device can be, for example, a mobile phone, a tablet, etc. Figure 1 , the live content understanding method may include the following steps.

[0044] Step S101, obtaining a video stream to be broadcast live.

[0045] It is worth noting that after the content understanding is performed on the video stream of the live broadcast to be understood and a content understanding result is obtained, the live broadcast to be understood can be distributed and recommended to users who are interested in the live broadcast to be understood based on the content understanding result.

[0046] In some embodiments, the video stream to be broadcast live can be pulled from a live broadcast source station. The live broadcast source station here can be understood as the address where the anchor publishes the live broadcast video.

[0047] Step S102: Slice the video stream into segments of equal length to obtain multiple video segments.

[0048] For example, the duration of the video clip may be 20 seconds or 25 seconds.

[0049] In some embodiments, the video content of two adjacent video segments may be continuous. For example, the video segments of the live broadcast to be understood include video segment 1 and video segment 2. Video segment 1 may be a video segment consisting of the 1st to 20th seconds of the live broadcast to be understood, and video segment 2 may be a video segment consisting of the 21st to 40th seconds of the live broadcast to be understood. The duration of video segment 1 and video segment 2 are both 20 seconds.

[0050] In some embodiments, the video content of two adjacent video clips may be non-continuous, and the two adjacent video clips with non-continuous video content may have partially overlapping video content. For example, the video clips of the live broadcast to be understood include video clip 1 and video clip 2, video clip 1 may be a video clip consisting of the 1st to 20th seconds of the live broadcast to be understood, and video clip 2 may also be a video clip consisting of the 16th to 35th seconds, the duration of video clip 1 and video clip 2 are both 20 seconds, and the overlapping video content in video clip 1 and video clip 2 is the video content from the 16th to 20th seconds of the live broadcast to be understood.

[0051] Step S103: extracting a feature vector of at least one video segment from among the multiple video segments.

[0052] In some embodiments, the feature vector of the video clip may be a unimodal feature vector, or a multimodal feature vector obtained by fusing multiple unimodal feature vectors. In the case where the feature vector of the video clip is a multimodal feature vector, Figure 1 The step S103 shown can be implemented in the following manner: extracting a multimodal feature vector of each video clip, wherein the multimodal features include an audio feature vector, a host voice feature vector, a subtitle feature vector, and a picture feature vector; fusing the multimodal feature vector of each video clip to obtain a feature vector of each video clip.

[0053] The audio feature vector may be an audio feature vector corresponding to the background audio of the video clip, where the background audio may be audio other than the host's voice. The audio feature vector may be detected using AED (Acoustic Events Detection), the main purpose of which is to detect whether a target sound event occurs in a continuous audio stream, such as the sound of an abnormal device failure or the sound of wild animals.

[0054] The host's voice feature vector can be extracted using ASR (Automatic Speech Recognition), which is a technology that converts human speech into text, so that the host's voice feature vector can be obtained based on the text.

[0055] Among them, the subtitle feature vector can use OCR (optical character recognition) technology to extract the subtitles in the live broadcast to be understood, so that the subtitle feature vector can be obtained based on the recognized text.

[0056] Among them, the picture feature vector can also be called an image feature vector, and the picture feature vector can be a HOG (Histogram of Oriented Gradient) feature, an LBP (Local Binary Pattern) feature, a Haar (Haar-like) feature, and the like.

[0057] Reference Figure 2 ,exist Figure 2 In , different line types with arrows are used to characterize the data flow direction of different video clips. Among them, video clip 1, video clip 2, video clip 2, ..., video clip N are different video clips obtained by dividing the video stream to be understood as live broadcast into equal lengths in chronological order, and the time length of each video clip is 20 seconds. For each video clip, the ASR feature vector, AED feature vector, subtitle feature vector and screen feature vector of each video clip are extracted respectively, and then the four feature vectors extracted from the video clip are fused to obtain the feature vector corresponding to the video clip, which is a multimodal feature vector obtained by fusion processing based on multiple single-mode feature vectors. Taking video clip 1 as an example, the ASR feature vector of video clip 1, the AED feature vector of video clip 1, the subtitle feature vector of video clip 1 and the screen feature vector of video clip 1 are extracted, and then the ASR feature vector of video clip 1, the AED feature vector of video clip 1, the subtitle feature vector of video clip 1 and the screen feature vector of video clip 1 are fused to obtain the feature vector of video clip 1.

[0058] By fusing the feature vectors of multiple single-modal feature vectors of the video clip, a feature representation of the video clip is obtained from multiple dimensions, thereby increasing the amount of information used to characterize the features of the video clip, and further achieving a comprehensive understanding of the video content.

[0059] Step S104: according to the feature vector of at least one video clip, a preset strategy is adopted to determine the content understanding result of the live broadcast to be understood.

[0060] Among them, the preset strategies include classification strategies or clustering strategies. The classification strategy is used to determine the content understanding result of the live broadcast to be understood by at least using category labels. The category labels are obtained by predicting the feature vectors of the video clips through a label prediction model obtained by training with manually annotated training samples. The content understanding result here is a specific category, such as singing category, dancing category, interaction category, etc.; the clustering strategy is used to determine the content understanding result of the live broadcast to be understood by clustering the feature vectors of all video clips using a clustering algorithm. The content understanding result here is a clustering identifier, which cannot be used to explain a specific category, but other live broadcasts similar to the live broadcast to be understood can be determined based on the clustering identifier.

[0061] It is worth noting that classification is to define categories in advance, and the number of categories is fixed, that is, the granularity of classification is fixed. The classification model needs to be trained by manually annotated classification training corpus, which can intuitively explain the categories. It is a supervised method and can also be called a white-box method. Clustering does not have pre-determined categories, and the number of categories is uncertain. Although the categories cannot be explained intuitively, clustering can effectively control the granularity, and clustering does not require manual annotation and pre-trained classifiers. It is an unsupervised method and can also be called a black-box method. Since the advantages and disadvantages of classification and clustering are different, in actual use, classification is more suitable for landing scenarios such as categories and classification systems that have been determined, refined operations, and ecological monitoring; while clustering is more suitable for landing scenarios where there is no classification system and the number of categories is uncertain, for example, similarity recommendations.

[0062] Through the above method, the video stream of the live broadcast to be understood is divided into multiple video segments of equal length, and the content understanding result of the live broadcast to be understood is determined based on the feature vector of each video segment. Since the video segments are short and the amount of content information is small, it is not easy to cause interference, so it is not easy to cause disagreements among different users, thereby improving the accuracy of content understanding; in addition, different strategies are adopted to obtain the content understanding result of the live broadcast to be understood. The classification strategy belongs to the supervised method, and the clustering strategy belongs to the unsupervised method. The two strategies have different advantages and can be applied to video content understanding in different landing scenarios.

[0063] The following are based on classification strategy and clustering strategy. Figure 1 Step S104 is shown for explanation.

[0064] When the preset strategy includes a classification strategy, Figure 1The shown step S104 can be implemented in the following manner: input the feature vector of each video clip into the label prediction model for processing to obtain the category label corresponding to the feature vector of each video clip; obtain the total duration of video clips with the same category label in the live broadcast to be understood, and obtain the total duration of all video clips with the same category label in other live broadcasts that are within the same live broadcast duration as the live broadcast to be understood; sort the obtained total duration of the live broadcast to be understood and other live broadcasts with the same category label in order from small to large to obtain a first sorting result; determine a target first sorting result in which the total duration corresponding to the live broadcast to be understood among all first sorting results exceeds a first preset percentile, and determine the category label corresponding to the total duration corresponding to the live broadcast to be understood in the target first sorting result as the content understanding result of the live broadcast to be understood.

[0065] The label prediction model can be pre-trained using training samples, where the training samples can be obtained through manual annotation, where the manual annotation includes the annotation of positive training samples and the annotation of negative training samples. The label prediction model can take the feature vector of the video feature as input, output the predicted probability that the video segment belongs to each category label, and take the category label with the highest probability as the category label corresponding to the feature vector of the video segment, and the category label is used to represent the specific category.

[0066] Among them, determining the content understanding result of the live broadcast to be understood by percentile is a statistical method. Percentile is used to compare the relative status of individuals in a group, and the first preset percentile can be set according to actual conditions.

[0067] Among them, when the total duration of the same category label of multiple live broadcasts is arranged, for the target live broadcast, the later the target live broadcast is arranged, the category label representing the target live broadcast can be positioned as the category label for this arrangement.

[0068] For example, when the category labels of the video clips include singing label, dancing label and interactive label, it is necessary to count the total duration of the video clips with singing label, the total duration of the video clips with dancing label and the total duration of the video clips with interactive label; then, count the total duration of all video clips with singing label, dancing label and interactive label in other live broadcasts that are in the same live broadcast duration as the live broadcast to be understood. If there are 50 other live broadcasts, the total duration of the singing label in the live broadcast to be understood and other live broadcasts is sorted in order from small to large to obtain the first sorting result. The first sorting result is the total duration of all video clips with singing label in 51 live broadcasts within the same live broadcast duration. If in the first sorting result, the total duration corresponding to the live broadcast to be understood exceeds the first preset percentile, the category label of the live broadcast to be understood can be determined as the singing label. At the same time, the first sorting result can be called the target first sorting result.

[0069] Among them, the type of other live broadcasts may be real-time live broadcasts or historical live broadcasts. It is worth noting that, since the total duration of video clips belonging to different tag categories in the historical live broadcasts can be counted and cached offline, the first ranking result can be determined more quickly based on the historical live broadcasts.

[0070] In the above method, since the category label belongs to the supervised mode, manual category labeling is required, and its classification is accurate, which can be applied to the landing scene of refined operation, for example; in addition, in live broadcast, the proportion of anchor language interaction is very high, and the number of category labels based on the segmented video clips in the live broadcast to be understood will lead to a large number of language interaction categories, thereby reducing the recall rate of scarce long-tail vertical categories. Therefore, the recall rate of scarce long-tail vertical categories can be guaranteed through the above percentile-based method.

[0071] When the preset strategy includes a classification strategy, Figure 1 The step S104 shown can be implemented in the following manner: input the feature vector of each video clip into the label prediction model for processing to obtain the category label corresponding to the feature vector of each video clip; at least input all category labels into the trained category prediction model to obtain the content understanding result of the live broadcast to be understood.

[0072] Among them, the label prediction model here can refer to the above embodiment, and this embodiment will not be described in detail here.

[0073] Among them, the category prediction model can be implemented by the XGB (extreme gradient boosting) algorithm.

[0074] Here, in addition to inputting the category label into the trained category prediction model, at least one of the feature vector of the video clip of the live broadcast to be understood, the feature vector of the title of the live broadcast to be understood, and the user feature vector of the host of the live broadcast to be understood can also be input into the trained category prediction model together with the category label to increase the amount of information of the category prediction model, thereby improving the accuracy of the category prediction model.

[0075] Through the above method, since the category label belongs to the supervised mode and needs to be manually labeled, the classification is accurate and can be applied to the implementation scenarios such as refined operations; and compared with the above percentile-based method, the model-based method is more efficient.

[0076] When the preset strategy includes a clustering strategy, Figure 1 The shown step S104 can be implemented in the following manner: clustering the feature vectors of all video clips using a clustering algorithm to obtain a clustering result, wherein the clustering result includes at least one cluster identifier representing a cluster; obtaining the total duration of video clips in the live broadcast to be understood that belong to the same cluster identifier, and obtaining the total duration of all video clips in other live broadcasts that belong to the same cluster identifier within the same live broadcast duration as the live broadcast to be understood; sorting the obtained total durations of the live broadcast to be understood and the other live broadcasts that belong to the same cluster identifier in order from small to large to obtain a second sorting result; determining a target second sorting result in which the total duration corresponding to the live broadcast to be understood among all second sorting results exceeds a second preset percentile, and determining the cluster identifier corresponding to the total duration corresponding to the live broadcast to be understood in the target second sorting result as the content understanding result of the live broadcast to be understood.

[0077] Among them, other types of live broadcasts can refer to the above embodiments, and this embodiment is not limited here.

[0078] Among them, the clustering algorithm can be Kmeans algorithm, density-based spatial clustering algorithm, etc.

[0079] Similar to the classification label, when the total duration of the same cluster identifier of multiple live broadcasts is arranged, for the target live broadcast, the later the target live broadcast is arranged, the category label representing the target live broadcast can be positioned as the cluster identifier for this arrangement.

[0080] For example, when the cluster identifiers of the video clips include cluster identifier 1, cluster identifier 2, and cluster identifier 3, it is necessary to count the total duration of the video clips belonging to cluster identifier 1, the total duration of the video clips belonging to cluster identifier 2, and the total duration of the video clips belonging to cluster identifier 3; then, count the total duration of all video clips belonging to cluster identifier 1, cluster identifier 2, and cluster identifier 3 in other live broadcasts that are within the same live broadcast duration as the live broadcast to be understood. If there are 20 other live broadcasts, the total duration of the live broadcast to be understood and other live broadcasts that belong to cluster identifier 3 are sorted in order from small to large to obtain a second sorting result. The second sorting result is the total duration of all video clips belonging to cluster identifier 3 in the 21 live broadcasts within the same live broadcast duration. If in the second sorting result, the total duration corresponding to the live broadcast to be understood exceeds the second preset percentile, the content understanding result of the live broadcast to be understood can be determined as cluster identifier 3. At the same time, the second sorting result can be called the target second sorting result.

[0081] Through the above method, since clustering belongs to the unsupervised mode, the granularity of clustering can be flexibly controlled, and there is no need for manual category labeling, which reduces labor costs. In addition, the method of determining the above content understanding results can also ensure the recall rate of scarce long-tail vertical categories.

[0082] When the preset strategy includes a clustering strategy, Figure 1 The shown step S104 can be implemented in the following manner: perform weighted averaging on the feature vectors of all video clips to obtain a mean vector used to describe the cluster to which the live broadcast to be understood belongs; perform distance calculation between the mean vector and the center vector of a preset cluster, and determine the cluster identifier of the preset cluster corresponding to the center vector with the shortest distance to the mean vector as the content understanding result of the live broadcast to be understood.

[0083] It is worth noting that the preset cluster is obtained in an offline manner, and the center vector of the preset cluster is used to describe the cluster, so that the cluster identifier of the preset cluster corresponding to the center vector with the shortest distance to the mean vector is determined as the content understanding result of the live broadcast to be understood.

[0084] For example, the distance calculation may be a Euclidean distance calculation, a Manhattan distance calculation, a Chebyshev distance calculation, a Mahalanobis distance calculation, and the like.

[0085] Through the above method, since clustering belongs to an unsupervised mode, the granularity of clustering can be flexibly controlled, and there is no need for manual category labeling, which reduces labor costs and can be applied to scenarios such as recommendation based on similarity. In addition, compared with the above percentile-based method, the content understanding result of the live broadcast to be understood is determined by the mean vector, which can reduce the amount of calculation and thus improve efficiency.

[0086] Figure 3is a block diagram of a live content understanding device according to an exemplary embodiment of the present disclosure, referring to Figure 3 ,include:

[0087] An acquisition module 301 is used to acquire a video stream to be broadcasted;

[0088] A slicing module 302, configured to slice the video stream into equal lengths to obtain a plurality of video segments;

[0089] An extraction module 303, configured to extract a feature vector of at least one video segment from the plurality of video segments;

[0090] A determination module 304 is configured to determine a content understanding result of the live broadcast to be understood by adopting a preset strategy according to the feature vector of the at least one video segment;

[0091] Among them, the preset strategy includes a classification strategy or a clustering strategy. The classification strategy is used to determine the content understanding result of the live broadcast to be understood by at least using category labels. The category labels are obtained by predicting the feature vectors of the video clips using a label prediction model trained with manually annotated training samples. The clustering strategy is used to determine the content understanding result of the live broadcast to be understood by clustering the feature vectors of all the video clips using a clustering algorithm.

[0092] In some embodiments, the preset strategy includes the classification strategy, and the determination module 304 includes:

[0093] A first processing submodule, configured to input the feature vector of each video clip into the label prediction model for processing to obtain a category label corresponding to the feature vector of each video clip;

[0094] The first acquisition submodule is used to acquire the total duration of video segments with the same category label in the live broadcast to be understood, and to acquire the total duration of all video segments with the same category label in other live broadcasts within the same live broadcast duration as the live broadcast to be understood;

[0095] A first sorting submodule is used to sort the total duration of the live broadcast to be understood and the other live broadcasts belonging to the same category label in ascending order to obtain a first sorting result;

[0096] The first determination submodule is used to determine a target first sorting result in which the total duration corresponding to the live broadcast to be understood is at a first preset percentile among all first sorting results, and determine the category label corresponding to the total duration corresponding to the live broadcast to be understood in the target first sorting result as the content understanding result of the live broadcast to be understood.

[0097] In some embodiments, the preset strategy includes the classification strategy, and the determination module 304 includes:

[0098] A second processing submodule is used to input the feature vector of each video clip into the label prediction model for processing to obtain a category label corresponding to the feature vector of each video clip;

[0099] The second determination submodule is used to input at least all of the category labels into a trained category prediction model to obtain a content understanding result of the live broadcast to be understood.

[0100] In some embodiments, the preset strategy includes the clustering strategy, and the determination module 304 includes:

[0101] A first clustering submodule, configured to cluster the feature vectors of all the video clips using the clustering algorithm to obtain a clustering result, wherein the clustering result includes at least one cluster identifier representing a cluster;

[0102] The second acquisition submodule is used to obtain the total duration of the video segments belonging to the same cluster identifier in the live broadcast to be understood, and to obtain the total duration of all video segments belonging to the same cluster identifier in other live broadcasts within the same live broadcast duration as the live broadcast to be understood;

[0103] A second sorting submodule is used to sort the total duration of the live broadcast to be understood and the other live broadcasts belonging to the same cluster identifier in ascending order to obtain a second sorting result;

[0104] The third determination submodule is used to determine a target second sorting result in which the total duration corresponding to the live broadcast to be understood exceeds a second preset percentile in all second sorting results, and determine the cluster identifier corresponding to the total duration corresponding to the live broadcast to be understood in the target second sorting result as the content understanding result of the live broadcast to be understood.

[0105] In some embodiments, the preset strategy includes the clustering strategy, and the determination module 304 includes:

[0106] A third processing submodule is used to perform weighted average processing on the feature vectors of all the video clips to obtain a mean vector for describing the cluster to which the live broadcast to be understood belongs;

[0107] The fourth determination submodule is used to calculate the distance between the mean vector and the center vector of the preset cluster, and determine the cluster identifier of the preset cluster corresponding to the center vector with the shortest distance to the mean vector as the content understanding result of the live broadcast to be understood.

[0108] In some embodiments, the extraction module 303 includes:

[0109] An extraction submodule, configured to extract a multimodal feature vector of each of the video clips, wherein the multimodal features include an audio feature vector, a host voice feature vector, a subtitle feature vector, and a picture feature vector;

[0110] The fusion submodule is used to perform fusion processing on the multimodal feature vectors of each of the video clips to obtain the feature vector of each of the video clips.

[0111] The implementation of each module of the device can refer to the relevant method embodiments, and this embodiment will not be described in detail here.

[0112] Reference below Figure 4 , which shows a schematic diagram of the structure of an electronic device 400 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0113] like Figure 4 As shown, the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0114] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 The electronic device 400 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0115] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable storage medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0116] It should be noted that the computer-readable storage medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable storage medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0117] In some embodiments, the electronic device may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0118] The computer-readable storage medium may be included in the electronic device, or may exist independently without being installed in the electronic device.

[0119] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device is enabled to: obtain a video stream of a live broadcast to be understood; slice the video stream with equal lengths to obtain multiple video segments; extract a feature vector of at least one video segment among the multiple video segments; and determine a content understanding result of the live broadcast to be understood by using a preset strategy based on the feature vector of the at least one video segment; wherein the preset strategy includes a classification strategy or a clustering strategy, and the classification strategy is used to determine the content understanding result of the live broadcast to be understood by at least using a category label, and the category label is obtained by predicting the feature vector of the video segment using a label prediction model obtained by training with manually annotated training samples, and the clustering strategy is used to determine the content understanding result of the live broadcast to be understood by using a clustering algorithm to perform clustering processing based on the feature vectors of all the video segments.

[0120] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0121] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0122] The modules involved in the embodiments described in the present disclosure may be implemented by software or hardware. The name of the module does not limit the module itself in some cases. For example, the acquisition module may also be described as a "module for acquiring a video stream to be understood live".

[0123] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0124] In the context of the present disclosure, a machine-readable storage medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable storage medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0125] According to one or more embodiments of the present disclosure, Example 1 provides a method for understanding live content, including:

[0126] Get the video stream to be understood.

[0127] Slicing the video stream with equal length to obtain multiple video segments;

[0128] extracting a feature vector of at least one video segment among the plurality of video segments;

[0129] Determining a content understanding result of the live broadcast to be understood by adopting a preset strategy according to the feature vector of the at least one video segment;

[0130] Among them, the preset strategy includes a classification strategy or a clustering strategy. The classification strategy is used to determine the content understanding result of the live broadcast to be understood by at least using category labels. The category labels are obtained by predicting the feature vectors of the video clips using a label prediction model trained with manually annotated training samples. The clustering strategy is used to determine the content understanding result of the live broadcast to be understood by clustering the feature vectors of all the video clips using a clustering algorithm.

[0131] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein the preset strategy includes the classification strategy, and the method of determining the content understanding result of the live broadcast to be understood by using the preset strategy according to the feature vector of the at least one video segment includes:

[0132] Inputting the feature vector of each video clip into the label prediction model for processing to obtain a category label corresponding to the feature vector of each video clip;

[0133] Obtain the total duration of video segments with the same category label in the live broadcast to be understood, and obtain the total duration of all video segments with the same category label in other live broadcasts within the same live broadcast duration as the live broadcast to be understood;

[0134] Sorting the total duration of the live broadcast to be understood and the other live broadcasts that belong to the same category label in ascending order to obtain a first sorting result;

[0135] Determine a target first sorting result in which the total duration corresponding to the live broadcast to be understood exceeds a first preset percentile among all first sorting results, and determine the category label corresponding to the total duration corresponding to the live broadcast to be understood in the target first sorting result as the content understanding result of the live broadcast to be understood.

[0136] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 1, wherein the preset strategy includes the classification strategy, and the determining, based on the feature vector of the at least one video segment, of the content understanding result of the live broadcast to be understood by using the preset strategy includes:

[0137] Inputting the feature vector of each video clip into the label prediction model for processing to obtain a category label corresponding to the feature vector of each video clip;

[0138] At least all the category labels are input into the trained category prediction model to obtain the content understanding result of the live broadcast to be understood.

[0139] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 1, wherein the preset strategy includes the clustering strategy, and the determining, based on the feature vector of the at least one video segment, of the content understanding result of the live broadcast to be understood by using the preset strategy includes:

[0140] Using the clustering algorithm to cluster the feature vectors of all the video clips to obtain a clustering result, wherein the clustering result includes at least one cluster identifier representing a cluster;

[0141] Obtain the total duration of video segments in the live broadcast to be understood that belong to the same cluster identifier, and obtain the total duration of all video segments in other live broadcasts that are in the same live broadcast duration as the live broadcast to be understood and belong to the same cluster identifier;

[0142] Sorting the total duration of the live broadcast to be understood and the other live broadcasts that belong to the same cluster identifier in ascending order to obtain a second sorting result;

[0143] Determine a target second sorting result in which the total duration corresponding to the live broadcast to be understood exceeds a second preset percentile among all second sorting results, and determine the cluster identifier corresponding to the total duration corresponding to the live broadcast to be understood in the target second sorting result as the content understanding result of the live broadcast to be understood.

[0144] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 1, wherein the preset strategy includes the clustering strategy, and the method of determining the content understanding result of the live broadcast to be understood by using the preset strategy according to the feature vector of the at least one video segment includes:

[0145] Performing weighted averaging processing on the feature vectors of all the video clips to obtain a mean vector for describing the cluster to which the live broadcast to be understood belongs;

[0146] A distance calculation is performed between the mean vector and a center vector of a preset cluster, and a cluster identifier of the preset cluster corresponding to the center vector having the shortest distance to the mean vector is determined as a content understanding result of the live broadcast to be understood.

[0147] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 1, wherein extracting a feature vector of at least one video segment from the multiple video segments includes:

[0148] Extracting a multimodal feature vector of each of the video clips, wherein the multimodal features include an audio feature vector, a host voice feature vector, a subtitle feature vector, and a picture feature vector;

[0149] The multimodal feature vectors of each of the video clips are fused to obtain a feature vector of each of the video clips.

[0150] According to one or more embodiments of the present disclosure, Example 7 provides a live content understanding device, including:

[0151] include:

[0152] An acquisition module is used to acquire the video stream to be broadcasted;

[0153] A slicing module, used for slicing the video stream with equal length to obtain multiple video segments;

[0154] An extraction module, configured to extract a feature vector of at least one video segment among the plurality of video segments;

[0155] A determination module, configured to determine a content understanding result of the live broadcast to be understood by adopting a preset strategy according to a feature vector of the at least one video segment;

[0156] Among them, the preset strategy includes a classification strategy or a clustering strategy. The classification strategy is used to determine the content understanding result of the live broadcast to be understood by at least using category labels. The category labels are obtained by predicting the feature vectors of the video clips using a label prediction model trained with manually annotated training samples. The clustering strategy is used to determine the content understanding result of the live broadcast to be understood by clustering the feature vectors of all the video clips using a clustering algorithm.

[0157] According to one or more embodiments of the present disclosure, Example 8 provides the method of Example 7, including:

[0158] The preset strategy includes the classification strategy, and the determination module includes:

[0159] A first processing submodule, configured to input the feature vector of each video clip into the label prediction model for processing to obtain a category label corresponding to the feature vector of each video clip;

[0160] The first acquisition submodule is used to acquire the total duration of video segments with the same category label in the live broadcast to be understood, and to acquire the total duration of all video segments with the same category label in other live broadcasts within the same live broadcast duration as the live broadcast to be understood;

[0161] A first sorting submodule is used to sort the total duration of the live broadcast to be understood and the other live broadcasts belonging to the same category label in ascending order to obtain a first sorting result;

[0162] The first determination submodule is used to determine a target first sorting result in which the total duration corresponding to the live broadcast to be understood is at a first preset percentile among all first sorting results, and determine the category label corresponding to the total duration corresponding to the live broadcast to be understood in the target first sorting result as the content understanding result of the live broadcast to be understood.

[0163] According to one or more embodiments of the present disclosure, Example 9 provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any of the methods described in Examples 1-6 when executed by a processing device.

[0164] According to one or more embodiments of the present disclosure, Example 10 provides an electronic device, including:

[0165] a storage device having a computer program stored thereon;

[0166] A processing device is used to execute the computer program in the storage device to implement the steps of any one of the methods described in Examples 1-6.

[0167] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.

[0168] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0169] Although the subject matter has been described in language specific to structural features and / or method logic actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims. Regarding the device in the above embodiment, the specific manner in which each module performs the operation has been described in detail in the embodiment related to the method, and will not be elaborated here.

Claims

1. A method for understanding live content, It is characterized in that include: Get the video stream to be understood. Slicing the video stream with equal length to obtain multiple video segments; extracting a feature vector of at least one video segment among the plurality of video segments; Determining a content understanding result of the live broadcast to be understood by adopting a preset strategy according to the feature vector of the at least one video segment; The preset strategy includes a classification strategy or a clustering strategy. The classification strategy is used to determine the content understanding result of the live broadcast to be understood by at least using a category label. The category label is obtained by predicting the feature vector of the video clip using a label prediction model trained by manually annotated training samples. The clustering strategy is used to determine the content understanding result of the live broadcast to be understood by clustering the feature vectors of all the video clips using a clustering algorithm. The preset strategy includes the classification strategy, and the preset strategy is used to determine the content understanding result of the live broadcast to be understood based on the feature vector of the at least one video clip, including: inputting the feature vector of each video clip into the label prediction model for processing to obtain the category label corresponding to the feature vector of each video clip; obtaining the total duration of video clips with the same category label in the live broadcast to be understood, and obtaining the total duration of all video clips with the same category label in other live broadcasts that are within the same live broadcast duration as the live broadcast to be understood; sorting the obtained total durations of the live broadcast to be understood and the other live broadcasts with the same category label in order from small to large to obtain a first sorting result; determining a target first sorting result in which the total duration corresponding to the live broadcast to be understood exceeds a first preset percentile among all first sorting results, and determining the category label corresponding to the total duration corresponding to the live broadcast to be understood in the target first sorting result as the content understanding result of the live broadcast to be understood.

2. The method according to claim 1, It is characterized in that The extracting a feature vector of at least one video segment from the plurality of video segments comprises: Extracting a multimodal feature vector of each of the video clips, wherein the multimodal features include an audio feature vector, a host voice feature vector, a subtitle feature vector, and a picture feature vector; The multimodal feature vectors of each of the video clips are fused to obtain a feature vector of each of the video clips.

3. A device for understanding live content, It is characterized in that include: An acquisition module is used to acquire the video stream to be broadcasted; A slicing module, used for slicing the video stream with equal length to obtain multiple video segments; An extraction module, configured to extract a feature vector of at least one video segment among the plurality of video segments; A determination module, configured to determine a content understanding result of the live broadcast to be understood by adopting a preset strategy according to a feature vector of the at least one video segment; The preset strategy includes a classification strategy or a clustering strategy. The classification strategy is used to determine the content understanding result of the live broadcast to be understood by at least using a category label. The category label is obtained by predicting the feature vector of the video clip using a label prediction model trained by manually annotated training samples. The clustering strategy is used to determine the content understanding result of the live broadcast to be understood by clustering the feature vectors of all the video clips using a clustering algorithm. The preset strategy includes the classification strategy, and the determination module includes: A first processing submodule, configured to input the feature vector of each video clip into the label prediction model for processing to obtain a category label corresponding to the feature vector of each video clip; The first acquisition submodule is used to acquire the total duration of video segments with the same category label in the live broadcast to be understood, and to acquire the total duration of all video segments with the same category label in other live broadcasts within the same live broadcast duration as the live broadcast to be understood; A first sorting submodule is used to sort the total duration of the live broadcast to be understood and the other live broadcasts belonging to the same category label in ascending order to obtain a first sorting result; The first determination submodule is used to determine a target first sorting result in which the total duration corresponding to the live broadcast to be understood is at a first preset percentile among all first sorting results, and determine the category label corresponding to the total duration corresponding to the live broadcast to be understood in the target first sorting result as the content understanding result of the live broadcast to be understood.

4. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processing device, the steps of the method according to claim 1 or 2 are implemented.

5. An electronic device, It is characterized in that include: a storage device having a computer program stored thereon; A processing device, used to execute the computer program in the storage device to implement the steps of the method according to claim 1 or 2.

Citation Information

Patent Citations

  • Live video classification and preview selection

    WO2017176808A1