Label-based Video Analysis Method, Device, Computer Equipment and Storage Medium

A video analysis method using a predefined tag system to label and analyze video data multidimensionally addresses the lack of accuracy in traditional metrics, providing actionable insights for video content optimization.

CN114491154BActive Publication Date: 2025-07-15特赞(上海)信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111041840.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-06
Publication Date
2025-07-15
Estimated Expiration
2041-09-06

AI Technical Summary

Technical Problem

Traditional video analysis methods cannot accurately analyze the video delivery effect and cannot find out the insight of the video, resulting in inaccurate analysis.

Method used

By obtaining the video data to be analyzed, the video data is annotated according to the pre-constructed video tag system, multi-dimensional label analysis is performed, and the target analysis results are generated, including effect data analysis and label data analysis, and the importance of video content labels is determined in combination with the label analysis model to generate video analysis results.

Benefits of technology

It improves the accuracy of video analysis, helps enterprises optimize their video delivery strategies, and improves the return on video delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114491154B_ABST
    Figure CN114491154B_ABST
Patent Text Reader

Abstract

The present application discloses a tag-based video analysis method, apparatus, computer device, and storage medium. The method includes: obtaining video data to be analyzed; annotating the video data to be analyzed according to a pre-constructed video tag system to obtain tagged video data; performing multi-dimensional tag analysis on the tagged video data to obtain a target analysis result; and generating a video analysis result according to the target analysis result. The present application can improve the accuracy of video analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more particularly, to a tag-based video analysis method, apparatus, computer device, and storage medium. Background Art

[0002] With the explosive growth of marketing content and the enrichment of online channels, brands have an increasing demand for video content. The volume of video production is increasing day by day, and the requirements are also getting higher and higher. Enterprises have found in video information flow delivery that the content production methods of different agencies and different KOLs (Key Opinion Leaders) are different, and the effects vary greatly. In the traditional method, the video delivery effect is analyzed by recording the number of users entering the video projection area and the number of video viewers, and the video delivery is feedback-optimized based on this. However, the traditional method cannot find the insights of the video, and the video analysis is not accurate enough. Summary of the Invention

[0003] The main purpose of the present application is to provide a tag-based video analysis method, apparatus, computer device, and storage medium that can improve the accuracy of video analysis.

[0004] To achieve the above object, according to one aspect of the present application, a tag-based video analysis method is provided.

[0005] The tag-based video analysis method according to the present application includes:

[0006] Obtain video data to be analyzed;

[0007] Annotate the video data to be analyzed according to a pre-constructed video tag system to obtain tagged video data;

[0008] Perform multi-dimensional tag analysis on the tagged video data to obtain a target analysis result;

[0009] Generate a video analysis result according to the target analysis result.

[0010] Further, the performing multi-dimensional tag analysis on the tagged video data to obtain a target analysis result includes:

[0011] Perform effect data analysis on the tagged video data to obtain an effect data analysis result;

[0012] Perform tag data analysis on the video content tags corresponding to the tagged video data to obtain a tag data analysis result;

[0013] Generate a target analysis result according to the effect data analysis result and the tag data analysis result.

[0014] Further, analyzing the effect data of the marked video data to obtain an effect data analysis result, including:

[0015] Calculating the click-through rate, consumption, and conversion number of the marked video data;

[0016] Generating an effect data analysis result based on the click-through rate, placement consumption, and conversion number.

[0017] Further, analyzing the label data of the marked video data to obtain a label data analysis result, including:

[0018] Flattening the video content labels corresponding to the marked video data;

[0019] Counting the basic label information corresponding to the flattened video content labels, and obtaining a label data analysis result based on the basic label information.

[0020] Further, generating a target analysis result according to the effect data analysis result and the label data analysis result, including:

[0021] Classifying the video content labels corresponding to the marked video data based on the effect data analysis result and the label data analysis result to obtain multiple label categories;

[0022] Performing a time series analysis on the video content labels corresponding to each label category to obtain the time periods corresponding to each label category;

[0023] Based on the effect data analysis result, analyzing the importance degree of the video content labels corresponding to the marked video data, and generating a target analysis result according to the importance degree of the video content labels, the time periods corresponding to each label category, and the effect data analysis result.

[0024] Further, analyzing the importance degree of the video content labels corresponding to the marked video data based on the effect data analysis result, including:

[0025] Aligning the effect data analysis result with the video content labels corresponding to the marked video data to obtain combined data;

[0026] Invoking a pre-constructed label analysis model, inputting the combined data and the video content labels into the label analysis model respectively, and determining the importance degree of the video content labels.

[0027] Further, generating a video analysis result according to the target analysis result, including:

[0028] Determine the key information corresponding to the video data to be analyzed according to the target analysis result;

[0029] Generate a video analysis result according to the key information.

[0030] To achieve the above object, according to another aspect of the present application, there is provided a tag-based video analysis device.

[0031] The tag-based video analysis device according to the present application includes:

[0032] A communication module for acquiring video data to be analyzed;

[0033] A labeling module for labeling the video data to be analyzed according to a pre-constructed video tag system to obtain labeled video data;

[0034] A tag analysis module for performing multi-dimensional tag analysis on the labeled video data to obtain a target analysis result;

[0035] A result generation module for generating a video analysis result according to the target analysis result.

[0036] A computer device includes a memory and a processor, the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the steps in the above-mentioned method embodiments are implemented.

[0037] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0038] The above-mentioned tag-based video analysis method, device, computer device and storage medium obtain video data to be analyzed, label the video data to be analyzed according to a pre-constructed video tag system to obtain labeled video data, and realize the tagging of video content. Thus, multi-dimensional tag analysis is performed on the labeled video data to obtain a target analysis result, and the relationship between video content and delivery effect is analyzed based on video tags. Furthermore, a video analysis result is generated according to the target analysis result, effectively improving the accuracy of video analysis and being beneficial to improving the return rate of video delivery. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The drawings constituting a part of this application are used to provide a further understanding of this application, making other features, objects and advantages of this application more obvious. The schematic embodiments and descriptions of the drawings of this application are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0040] Figure 1It is an application environment diagram of a label - based video analysis method in an embodiment;

[0041] Figure 2 It is a schematic flowchart of a label - based video analysis method in an embodiment;

[0042] Figure 3 It is a schematic diagram of the label structure of a video label system in an embodiment;

[0043] Figure 4 It is a schematic diagram of video data with a single label in an embodiment;

[0044] Figure 5 It is a schematic diagram of a single label data in an embodiment;

[0045] Figure 6 It is a schematic diagram of the result of effect data analysis in an embodiment;

[0046] Figure 7 It is a scatter plot of the position distribution corresponding to the first - level label in an embodiment;

[0047] Figure 8 It is a schematic flowchart of the step of generating a target analysis result according to the result of effect data analysis and the result of label data analysis in an embodiment;

[0048] Figure 9 It is a distribution diagram of label categories in an embodiment;

[0049] Figure 10 It is a schematic diagram of the time period corresponding to each label category when the content is stratified into oral broadcast and subtitles in an embodiment;

[0050] Figure 11 It is a schematic diagram of combined data in an embodiment;

[0051] Figure 12 It is a structural block diagram of a label - based video analysis device in an embodiment;

[0052] Figure 13 It is an internal structure diagram of a computer device in an embodiment. Specific embodiments

[0053] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.

[0054] It should be noted that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0055] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will detail this application with reference to the drawings and in conjunction with the embodiments.

[0056] The tag-based video analysis method provided in this application can be applied to, for example, Figure 1 the application environment shown. Among them, the terminal 102 communicates with the server 104 through a network. The server 104 obtains the video analysis request sent by the terminal 102, parses the video analysis request to obtain the video data to be analyzed, annotates the video data to be analyzed according to the pre-constructed video tag system to obtain the tagged video data, and then performs multi-dimensional tag analysis on the tagged video data to obtain the target analysis result, and further generates the video analysis result according to the target analysis result. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0057] In one embodiment, as Figure 2 shown, a tag-based video analysis method is provided. Taking the method applied to the Figure 1 server as an example, it includes the following steps 202 to 208:

[0058] Step 202, obtain the video data to be analyzed.

[0059] The video data to be analyzed refers to the data that needs to be analyzed for video content. The video data to be analyzed may include multiple video data, and each video data corresponds to a video ID (Identity document).

[0060] Specifically, obtaining the video data to be analyzed includes: obtaining the original video data; performing data cleaning on the original video data to obtain the video data to be analyzed. The original video data refers to the unprocessed video data. It is necessary to roughly review the original video data and then perform data cleaning on the data, and select the data that can be analyzed as the video data to be analyzed. For example, by roughly reviewing the original data, it is found that there are 145 pieces of data in the account, including videos and effect data. Perform data cleaning, select the data of the cherry blossom SKU, exclude the data with no consumption or extremely low consumption, and duplicate removal processing can also be performed on the videos. Finally, 69 pieces of data of the cherry blossom SKU are obtained, which is the video data to be analyzed.

[0061] Step 204: Label the video data to be analyzed according to the pre-constructed video tag system to obtain the labeled video data.

[0062] The server stores a pre-constructed video tag system. The tag structure of the video tag system can be a tag tree. The tag tree can be as Figure 3 shown.

[0063] The server can perform tagging processing on the video content of the video data to be analyzed according to the pre-constructed video tag system, and label the video data to be analyzed according to the video content tags. Specifically, extract the video features of the video data to be analyzed, match the video features with the preset tags in the pre-constructed video tag system, determine the video content tags corresponding to the video data to be analyzed, and label the video data to be analyzed according to the video content tags. A single piece of labeled video data can be as Figure 4 shown. The video content tags of a single piece of labeled video data can include content stratification, first-level tags, second-level tags, third-level tags, etc. It is also possible to separately count the video content tags corresponding to the labeled video data to obtain each piece of tag data, including the appearance layer, as well as tags, second-level tags, and the original text. A single piece of tag data can be as Figure 5 shown.

[0064] Step 206: Perform multi-dimensional tag analysis on the labeled video data to obtain the target analysis result.

[0065] The multi-dimensional tag analysis includes effect dimension analysis and tag dimension analysis, and analyzes the relationship between the effect dimension analysis result and the tag dimension analysis result, so as to obtain the target analysis result.

[0066] In one embodiment, multi-dimensional label analysis is performed on the labeled video data to obtain a target analysis result, including: performing effect data analysis on the labeled video data to obtain an effect data analysis result; performing label data analysis on the video content labels corresponding to the labeled video data to obtain a label data analysis result; and generating a target analysis result based on the effect data analysis result and the label data analysis result.

[0067] Effect dimension analysis refers to performing effect data analysis on the labeled video data. Effect data analysis refers to analyzing the video delivery effect of the labeled video data, and the video delivery effect can be determined by calculating the click-through rate, conversion number, and consumption of the video.

[0068] Label dimension analysis refers to performing label data analysis on the labeled video data. Label data analysis can include analysis of video duration distribution, analysis of the number of times each video content label is mentioned, analysis of the position distribution of each video content label in the video, etc.

[0069] After obtaining the effect data analysis result and the label data analysis result, the relationship between the video content label and the delivery effect is determined based on the above analysis results. Specifically, the target analysis result can be generated by analyzing the quality of the label, the time when the good / bad label appears, and the importance of the label. The target analysis result can represent the relationship between the video content label and the delivery effect.

[0070] Step 208, generating a video analysis result based on the target analysis result.

[0071] The target analysis result represents the relationship between the video content label and the delivery effect. Thus, the key information of the video data to be analyzed can be extracted based on the target analysis result to obtain the insight of the video data to be analyzed, and then the video analysis result can be obtained. Enterprises can perform targeted optimization processing on video delivery according to the video analysis result, which can effectively improve the video delivery reporting rate.

[0072] In this embodiment, by obtaining the video data to be analyzed and labeling the video data to be analyzed according to the pre-constructed video label system, the labeled video data is obtained, and the video content is labeled. Then, multi-dimensional label analysis is performed on the labeled video data to obtain a target analysis result, and the relationship between the video content and the delivery effect is analyzed based on the video label, and then the video analysis result is generated based on the target analysis result, which effectively improves the accuracy of video analysis and is beneficial to improving the return rate of video delivery.

[0073] In one embodiment, effect data analysis is performed on the labeled video data to obtain an effect data analysis result, including: calculating the click-through rate, consumption, and conversion number of the labeled video data; and generating an effect data analysis result based on the click-through rate, delivery consumption, and conversion number.

[0074] Effect data analysis can be called effect data EDA (Exploratory Data Analysis). The click-through rate, conversion number, and consumption of the video can be calculated, and graphs can be plotted based on the click-through rate, conversion number, and consumption to obtain the effect data analysis result. In this embodiment, a two-dimensional coordinate system graph can be constructed with the ID of the labeled video data as the x-axis and the consumption cost or duration as the y-axis. The sorting basis for each labeled video data on the x-axis can be any one of the click-through rate, conversion number, or consumption. As Figure 6 shown, the click-through rate can be selected as the sorting basis to determine the order of each labeled video data on the x-axis, and the consumption cost can be used as the y-axis to construct a coordinate system graph, so that the relationship between the click-through rate and the consumption cost of the video can be obtained.

[0075] In one embodiment, label data analysis is performed on the labeled video data to obtain a label data analysis result, including: flattening the video content labels corresponding to the labeled video data; counting the basic label information corresponding to the flattened video content labels, and obtaining the label data analysis result based on the basic label information.

[0076] Label data analysis is called label data EDA (Exploratory Data Analysis). Label data analysis can include analyzing the basic label information corresponding to the video content labels, including the analysis of the video duration distribution, the analysis of the number of times each video content label is mentioned, the analysis of the position distribution of each video content label in the video, etc.

[0077] Specifically, first flatten the video content labels corresponding to the labeled video data. The video content labels corresponding to each video are in an Excel file, and the labels are in a two-dimensional table and include a time dimension. For convenient analysis, the data is flattened. The flattening process uses the secondary label as the smallest label unit for data analysis, and the tertiary label and subsequent labels are used as the specific values of the smallest label unit.

[0078] Exemplarily, the video content labels corresponding to the labeled video data can be as shown in the following table:

[0079]

[0080] Among them, content stratification indicates the carrier position where the label appears, such as the label appears in the voiceover, picture, etc.; the first-level label indicates the root label of the label tree, which can be further divided into second-level labels; the second-level label indicates the leaf label in the label tree; the third-level label, etc., indicates the specific value in the upper-level label. Using the leaf label as the smallest label unit for data analysis, the third-level label is used as the specific value of the smallest label unit to flatten the video content labels. The flattened data is as follows:

[0081]

[0082] After the data flattening process, each row is a video data, and the columns are content labels. The cell contains the label value, the time when the label appears, and the duration.

[0083] Analyze the video duration distribution of the labeled video data. Specifically, construct a video duration distribution histogram based on the labeled video data, calculate the average video duration according to the histogram, and determine the shortest video duration and the longest video duration.

[0084] The server can also count the number of times each video content label is mentioned, and determine the labels that are mentioned more or less according to the counted number of times the labels are mentioned. For example, by counting the number of times the labels are mentioned, in the content stratification labels, more labels appear in the voiceover and subtitle layers, which is 137 times, accounting for about 44%, followed by the picture display layer, which is 74, accounting for about 24%, and the labels mentioned in the caption and content creativity layers are the least. Among the level labels, labels of types such as conversion stimulation, product display, efficacy description, brand information, and product basic information are mentioned more, while labels of types such as user psychology, preferential selling points, and influencer call are mentioned less.

[0085] Furthermore, separately count the position distributions of various labels such as content stratification labels, first-level labels, second-level labels, and third-level labels that appear in the video. Specifically, perform a position distribution analysis on each type of label separately, obtain the time points and durations when each label in each type of label appears in the video. Represent the duration with scatter points, and construct a position scatter plot corresponding to each type of label according to the time points and durations when each label in each type of label appears in the video, so as to determine the distribution characteristics of each type of label according to the position scatter plot. The distribution characteristics can include at the beginning / middle / end, labels that are mentioned more and have a longer duration, labels that are mentioned less, labels that are mentioned, etc. Exemplarily, as Figure 7 shown, it is the position distribution scatter plot corresponding to the first-level label. According to this figure, it can be obtained that: at the beginning, conversion stimulation, conversion purpose, brand information, product basic information, and body text labels are mentioned more and have a longer duration; at the end, pain point description, efficacy description, influencer recommendation, etc. are all mentioned.

[0086] For another example, for the position scatter plot corresponding to the content stratification class label, the analysis shows that: at the beginning, the content labels mainly appear in the package frame brand area layer, the paragraph layer, the voiceover and subtitle layer, and the picture display layer. Among them, the content labels in the package frame brand area layer and the paragraph layer have a longer duration, while the labels in the voiceover and subtitle layer have a shorter duration. At the end, there are labels in the voiceover and subtitle layer and the paragraph layer. Around 40 - 50 seconds in the middle, the label appearance rate is relatively low, and it is considered to check whether there are statistical errors in the code.

[0087] After analyzing the basic information of the labels such as the analysis of the video duration distribution, the number of times each video content label is mentioned, and the position distribution of each video content label in the video, the analysis results are used as the label data analysis results.

[0088] In this embodiment, by flattening the video content labels corresponding to the labeled video data, it is beneficial to subsequent analysis of the basic information of the labels. Statistically analyze the basic information of the labels corresponding to the flattened video content labels, and obtain the label data analysis results based on the basic information of the labels. Through the analysis of the video duration distribution, the number of times each video content label is mentioned, and the position distribution of each video content label in the video, it is possible to comprehensively and accurately analyze the appearance time and other situations of the video content labels in the video.

[0089] In one embodiment, as Figure 8 shown, the steps of generating the target analysis result according to the effect data analysis result and the label data analysis result include:

[0090] Step 802, based on the effect data analysis result and the label data analysis result, classify the video content labels corresponding to the labeled video data to obtain multiple label categories.

[0091] Step 804, perform a time series analysis on the video content labels corresponding to each label category to obtain the time periods corresponding to each label category.

[0092] Step 806, based on the effect data analysis result, analyze the importance of the video content labels corresponding to the labeled video data, and generate the target analysis result according to the importance of the video content labels, the time periods corresponding to each label category, and the effect data analysis result.

[0093] Based on the effect data analysis result and the label data analysis result, analyze the quality of the labels, the time when good / bad labels appear, and the importance of the labels, and generate the target analysis result according to the analysis data. The target analysis result can represent the relationship between the video content labels and the placement effect.

[0094] The quality of analysis tags can be based on the results of effect data analysis and tag data analysis, calculating the hit count of tags, the rising / falling hit count, and the tag ctr (Click-Through-Rate). Among them, the hit count represents the number of times the tag appears in the video, the rising / falling hit count represents the number of times the tag appears during the period when the video ctr rises / falls, and the tag ctr represents the weighted average ctr of all videos containing the tag. Thus, the video content tags corresponding to the tagged video data can be classified according to the hit count of the tag, the rising / falling hit count, and the tag ctr, obtaining multiple tag categories. Further, those with a tag ctr higher than the median ctr of all tags and a rising hit count greater than the falling hit count are classified as category A. Those with a tag ctr lower than the median ctr of all tags and a rising hit count greater than the falling hit count are classified as category B. Those with a tag ctr higher than the median ctr of all tags and a rising hit count less than the falling hit count are classified as category C. Those with a tag ctr lower than the median ctr of all tags and a rising hit count less than the falling hit count are classified as category D, obtaining 4 categories of tags, with category A tags being the best and category D tags being the worst. The distribution diagram of tag categories can be as Figure 9 shown. Category A tags may include voiceover and subtitles - fragrance - pleasant smell (no aspect) - pleasant smell, voiceover and subtitles - conversion purpose - guide to purchase - buy now immediately. Category B tags may include voiceover and subtitles - efficacy description - fluffy - naturally fluffy. Category C tags may include voiceover and subtitles - efficacy description - fluffy - twice as fluffy. Category D tags may include voiceover and subtitles - efficacy description - fluffy - fluffy.

[0095] The server then performs a time series analysis on the video content tags corresponding to each tag category. Specifically, it views the time series data of the tags appearing in the four regions A, B, C, and D corresponding to each content layer, and obtains the positions of the better tags (AB) for each content layer. For example, when the content layer is voiceover and subtitles, the schematic diagram of the time periods corresponding to each tag category can be as Figure 10 shown. By performing a time series analysis on the video content tags corresponding to each tag category, the time when the better tags should appear can be obtained.

[0096] In one embodiment, based on the results of effect data analysis, the importance of the video content tags corresponding to the tagged video data is analyzed, including: aligning the results of effect data analysis with the video content tags corresponding to the tagged video data to obtain combined data; calling a pre-constructed tag analysis model, and inputting the combined data and the video content tags into the tag analysis model respectively to determine the importance of the video content tags.

[0097] The alignment method can be to splice multiple video content tags to obtain a tag column. The column value is the time when the tag appears and the duration, separated by the "&" symbol. Multiple time points in a video are separated by the "、" symbol. For example, Figure 11 As shown, the column is the tag name, which is spliced by the appearance layer, first-level tag, second-level tag, and third-level tag. The column values of 291 tags are the time when the tag appears and the duration, separated by the "&" symbol. After aligning with the effect data analysis results, there are a total of 19 pieces of data. By combining the data, the relationship between the number of tags and the change in video CTR can be analyzed.

[0098] A tag analysis model is pre-constructed in the server. The tag analysis model can be composed of a click-through rate prediction model and an attention model. The combined data and the video content tags are respectively input into the tag analysis model to output the importance of the video content tags. By aligning the effect data analysis results with the video content tags corresponding to the tagged video data, it is beneficial to analyze the importance of the tags in the follow-up. And by analyzing the importance of the tags through the pre-constructed tag analysis model, the importance of the tags can be accurately predicted.

[0099] Furthermore, the importance of the video content tags, the time periods corresponding to each tag category, and the effect data analysis results are used as the target analysis results.

[0100] In this embodiment, based on the effect data analysis results and the tag data analysis results, the video content tags corresponding to the tagged video data are classified to obtain multiple tag categories, and good and bad tags can be determined. Performing a time series analysis on the video content tags corresponding to each tag category to obtain the time periods corresponding to each tag category can determine the time when good and bad tags should appear. Based on the effect data analysis results, analyzing the importance of the video content tags corresponding to the tagged video data, and generating the target analysis results according to the importance of the video content tags, the time periods corresponding to each tag category, and the effect data analysis results can obtain the influence degree of the tags on the delivery effect.

[0101] In one embodiment, the tag analysis model includes a click-through rate prediction model and an attention model. Invoking the pre-constructed tag analysis model and inputting the combined data and the video content tags into the tag analysis model respectively to determine the importance of the video content tags includes: inputting the combined data into the click-through rate prediction model to determine the influence degree of each tag in the combined data on the click-through rate; inputting the video content tags into the attention model to output the attention distribution corresponding to the click-through rate of the tagged video data; and determining the importance of the video content tags according to the attention distribution and the influence degree of each tag in the combined data on the click-through rate.

[0102] The server can input the combined data into the click-through rate prediction model to determine the influence degree of each tag pair in the combined data on the click-through rate. The click-through rate prediction model can be trained using multiple regression models. For example, the regression models can include KNeighborsUnif, KNeighborsDist, LightGBMXT, LightGBM, RandomForestMSE, CatBoost, ExtraTreesMSE, NeuralNetFastAI, XGBoost, NeuralNetMXNet, LightGBMLarge, WeightedEnsemble_L2. Specifically, the data features considered during the training process are: the number of times a tag appears in a video. The combined data can be sorted according to this data feature, and the click-through rate of the sorted data can be predicted through the regression model. By comparing the predicted click-through rate with the actual click-through rate, the final model is determined as the click-through rate prediction model. Through the click-through rate prediction model, budget processing is performed on the combined data, and the influence degree of each tag on the click-through rate is output, including the importance score (importance) and significance score (p_valune) of each tag.

[0103] In video content tags, the tags are composed of time series. The features that affect the video click-through rate (CTR) do not exist independently. It is very likely that the influence is caused by the combined pattern of several tag orders. Based on the above problems, the server can convert the time series tags into natural language tasks and use the attention model to output the attention distribution. The attention model in this embodiment is an attention model based on the tag system, and the attention model can perform video CTR attribution processing. Specifically, the video content tags, the start time of each tag, and the tag duration are input into the attention model, and the input data is processed into token data, including: tag token, tag start time token, and tag duration ratio token. Feature extraction is performed on each token data to obtain the tag embedding, tag start time embedding, and tag duration ratio embedding corresponding to each tag. The tag embedding, tag start time embedding, and tag duration ratio embedding corresponding to each tag are fused to obtain the tag fusion feature corresponding to each tag, and a fusion sequence is obtained. The context feature of the fusion sequence is extracted through the GRU network model, and the attention distribution attention weight of the video CTR score is obtained using Attention Pooling. Furthermore, the importance of all video content tags is determined according to the attention distribution and the influence degree of each tag in the combined data on the click-through rate. Important tags can be selected from all video content tags for video analysis. For example, 20 can be selected as important tags, as shown in the following table:

[0104]

[0105] Furthermore, a feature can be added to each tag in the attention model: a weight that decreases with time, simulating the real attention tendency of people watching videos, which is beneficial to improving the effectiveness of tag importance analysis.

[0106] In this embodiment, the click-through rate prediction model can accurately determine the influence degree of each tag in the combined data on the click-through rate, while the attention model can output the attention distribution corresponding to the click-through rate of the tagged video data, and the influence degree of the combination of tag orders on the click-through rate can be obtained. Thus, the importance of video content tags is determined according to the attention distribution and the influence degree of each tag in the combined data on the click-through rate. Therefore, the influence of tags on the video delivery effect can be accurately predicted.

[0107] In one of the embodiments, a video analysis result is generated according to the target analysis result, including: determining the key information corresponding to the video data to be analyzed according to the target analysis result; generating a video analysis result according to the key information.

[0108] The target analysis results include the importance of video content tags, the time periods corresponding to each tag category, and the effect data analysis results. The server extracts the key information corresponding to the video data to be analyzed based on the target analysis results to obtain the insight of each video data. For example, the insight of each video data may include: Insight 1: The user loss is serious after 3 seconds, and the first 3 seconds and the first 10 seconds are the golden periods for video display content. Specifically, the difference between high click-through rate and low click-through rate lies in whether key data is intensively displayed within 10 seconds, and the product display brand information should be emphasized in the first 3 seconds. Insight 2: All video effects will peak within 3 seconds. Optimizing the first 3 seconds has an overall improvement effect on the content, and optimizing the user loss from 3 to 30 seconds is also an effective direction. Insight 3: There is a tendency for the conversion rate to decrease when the video duration is too long, and 20 - 30 seconds is a more reasonable video duration. Specifically, using more mixed-cut videos of 20 - 30 seconds will have a more stable effect than longer voice-over videos.

[0109] Furthermore, it is also possible to count the ranking information of preset video content types, including: average click-through rate ranking, average conversion rate ranking, comprehensive index ranking, existing video material quantity ranking, click-through rate variance ranking, conversion rate variance ranking. The comprehensive index calculation formula is: 60% * click-through rate + 30% * conversion rate - 10% * average click cost. The preset video content types may include drama, influencer voice-over, mixed cut, and single-person voice-over - star, etc. Determine the video placement strategy based on the statistical ranking information of the preset video content types and the insight of each video data. For example, if the insight is that there is a tendency for the conversion rate to decrease when the video duration is too long, and 20 - 30 seconds is a more reasonable video duration, then more mixed-cut videos of 20 - 30 seconds can be used, which will have a more stable effect than longer voice-over videos.

[0110] Summarize the insights of multiple video data to obtain the video analysis results. For example, the summary table may be as follows:

[0111]

[0112] In this embodiment, since the target analysis results include the importance of video content tags, the time periods corresponding to each tag category, and the effect data analysis results, the key information corresponding to the video data to be analyzed is determined based on the target analysis results, and then the video analysis results are generated. It can improve the accuracy of video analysis, quickly determine the content tags with better video placement effects, and the production methods of video content.

[0113] Note that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0114] In one embodiment, as Figure 12 shown, a tag-based video analysis device is provided, including: a communication module 1202, a labeling module 1204, a tag analysis module 1206, and a result generation module 1208, where:

[0115] The communication module 1202 is configured to obtain video data to be analyzed.

[0116] The labeling module 1204 is configured to label the video data to be analyzed according to a pre-constructed video tag system to obtain labeled video data.

[0117] The tag analysis module 1206 is configured to perform multi-dimensional tag analysis on the labeled video data to obtain a target analysis result.

[0118] The result generation module 1208 is configured to generate a video analysis result according to the target analysis result.

[0119] In one embodiment, the tag analysis module 1206 is further configured to perform effect data analysis on the labeled video data to obtain an effect data analysis result; perform tag data analysis on the video content tags corresponding to the labeled video data to obtain a tag data analysis result; and generate a target analysis result according to the effect data analysis result and the tag data analysis result.

[0120] In one embodiment, the tag analysis module 1206 is further configured to calculate the click-through rate, consumption, and conversion number of the labeled video data; and generate an effect data analysis result according to the click-through rate, placement consumption, and conversion number.

[0121] In one embodiment, the tag analysis module 1206 is further configured to flatten the video content tags corresponding to the labeled video data; count the basic tag information corresponding to the flattened video content tags, and obtain a tag data analysis result according to the basic tag information.

[0122] In one embodiment, the label analysis module 1206 is further configured to classify the video content labels corresponding to the labeled video data based on the effect data analysis result and the label data analysis result, to obtain multiple label categories; perform a timing analysis on the video content labels corresponding to each label category, to obtain the time periods corresponding to each label category; analyze the importance of the video content labels corresponding to the labeled video data based on the effect data analysis result, and generate a target analysis result according to the importance of the video content labels, the time periods corresponding to each label category, and the effect data analysis result.

[0123] In one embodiment, the label analysis module 1206 is further configured to align the effect data analysis result with the video content labels corresponding to the labeled video data, to obtain combined data; call a pre-constructed label analysis model, and input the combined data and the video content labels into the label analysis model respectively, to determine the importance of the video content labels.

[0124] In one embodiment, the result generation module 1208 is further configured to determine the key information corresponding to the video data to be analyzed according to the target analysis result; generate a video analysis result according to the key information.

[0125] For the specific limitations of the label-based video analysis device, reference may be made to the limitations of the label-based video analysis method in the foregoing text, which will not be elaborated herein. Each module in the foregoing label-based video analysis device may be implemented in whole or in part by software, hardware, and their combination. The foregoing modules may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the foregoing modules.

[0126] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 13 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Wherein, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data of a label-based video analysis method. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a label-based video analysis method.

[0127] Those skilled in the art can understandFigure 13 The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0128] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the above various embodiments are implemented.

[0129] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above various embodiments are implemented.

[0130] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application may include non-volatile and / or volatile memories. Non-volatile memories may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0131] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0132] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A tag-based video analysis method, characterized in that Including: Obtain the video data to be analyzed; Annotate the video data to be analyzed according to the pre-constructed video tag system to obtain the tagged video data; Perform multi-dimensional tag analysis on the tagged video data to obtain the target analysis result; The multi-dimensional tag analysis includes effect dimension analysis and tag dimension analysis, and analyze the relationship between the effect dimension analysis result and the tag dimension analysis result to obtain the target analysis result; Dimension tag analysis includes: Perform effect data analysis on the tagged video data; Calculate the click-through rate, consumption, and conversion number of the tagged video data; Generate an effect data analysis result according to the click-through rate, placement consumption, and conversion number; Perform tag data analysis on the video content tags corresponding to the tagged video data to obtain a tag data analysis result; Tag dimension analysis includes: Flatten the video content tags corresponding to the tagged video data; Statistically analyze the basic tag information corresponding to the flattened video content tags, and obtain a tag data analysis result according to the basic tag information; Classify the video content tags corresponding to the tagged video data based on the effect data analysis result and the tag data analysis result to obtain multiple tag categories; Perform time series analysis on the video content tags corresponding to each tag category to obtain the time period corresponding to each tag category; Analyze the relationship between the effect dimension analysis result and the tag dimension analysis result to obtain the target analysis result, including: Based on the effect data analysis result, analyze the importance of the video content tags corresponding to the tagged video data, and generate a target analysis result according to the importance of the video content tags, the time period corresponding to each tag category, and the effect data analysis result; Generate a video analysis result according to the target analysis result.

2. The method according to claim 1, wherein The analysis of the importance of the video content tags corresponding to the tagged video data based on the effect data analysis result includes: Align the effect data analysis result with the video content tags corresponding to the tagged video data to obtain combined data; Call the pre-constructed tag analysis model, input the combined data and the video content tags into the tag analysis model respectively, and determine the importance of the video content tags.

3. The method according to claim 1, wherein The generation of the video analysis result according to the target analysis result includes: Determine the key information corresponding to the video data to be analyzed according to the target analysis result; Generate a video analysis result according to the key information.

4. A tag-based video analysis device, characterized in that, The device includes: A communication module for obtaining the video data to be analyzed; An annotation module for annotating the video data to be analyzed according to the pre-constructed video tag system to obtain the tagged video data; A tag analysis module for performing multi-dimensional tag analysis on the tagged video data to obtain the target analysis result; The tag analysis module is further used for: The multi-dimensional tag analysis includes effect dimension analysis and tag dimension analysis, and analyze the relationship between the effect dimension analysis result and the tag dimension analysis result to obtain the target analysis result; Dimension tag analysis includes: Perform effect data analysis on the marked video data; Calculate the click-through rate, consumption, and conversion count of the marked video data; Generate an effect data analysis result based on the click-through rate, placement consumption, and conversion count; Perform label data analysis on the video content labels corresponding to the marked video data to obtain a label data analysis result; Label dimension analysis, including: Flatten the video content labels corresponding to the marked video data; Statistically analyze the basic label information corresponding to the flattened video content labels, and obtain a label data analysis result based on the basic label information; Classify the video content labels corresponding to the marked video data based on the effect data analysis result and the label data analysis result to obtain multiple label categories; Perform time series analysis on the video content labels corresponding to each label category to obtain the time periods corresponding to each label category; Analyze the relationship between the effect dimension analysis result and the label dimension analysis result to obtain a target analysis result, including: Based on the effect data analysis result, analyze the importance of the video content labels corresponding to the marked video data, and generate a target analysis result based on the importance of the video content labels, the time periods corresponding to each label category, and the effect data analysis result; A result generation module for generating a video analysis result based on the target analysis result.

5. A computer device, comprising a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Label generation method and device and computer readable storage medium

    CN111708913A

  • Video analysis method and device, electronic equipment and readable storage medium

    CN112948635A