Analysis Method, Device, Computer Equipment and Storage Medium of Tags

The method improves video analysis accuracy by extracting and analyzing video content tags using a tag analysis model, enabling precise identification of effective video content elements.

CN114443901BActive Publication Date: 2025-07-15特赞(上海)信息科技有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111041764.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-06
Publication Date
2025-07-15
Estimated Expiration
2041-09-06

AI Technical Summary

Technical Problem

Traditional video analysis methods cannot accurately determine the video content elements with good video delivery effect, resulting in inaccurate analysis.

Method used

By obtaining the content labels of video data, performing effect data analysis and alignment, calling pre-constructed label analysis models, including click-through rate prediction models and attention models, determining the importance of video content labels, and analyzing the importance of labels in combination with click-through rate and attention distribution.

Benefits of technology

It improves the accuracy of video analysis, can accurately predict the importance and influence of tags, and optimizes the video delivery effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114443901B_ABST
    Figure CN114443901B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, computer device, and storage medium for analyzing tags. The method includes: obtaining video content tags of tagged video data; performing effect data analysis on the tagged video data to obtain an effect data analysis result; aligning the effect data analysis result and the video content tags to obtain combined data; and calling a pre-constructed tag analysis model, and inputting the combined data and the video content tags into the tag analysis model respectively to determine the importance level of the video content tags. The present application can improve the accuracy of video analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and in particular, to a method, an apparatus, a computer device, and a storage medium for analyzing tags. Background Art

[0002] With the explosive growth of marketing content and the enrichment of online channels, brands have an increasing demand for video content. The volume of video production is increasing day by day, and the requirements are also getting higher and higher. Enterprises have found in video information flow delivery that the content production methods of different agencies and different KOLs (Key Opinion Leaders) are different, and the effects also vary greatly. In the traditional method, the delivery effect of a video is analyzed by recording the number of users entering the video projection area and the number of video viewing users, and the video delivery is feedback-optimized based on this. However, the traditional method cannot determine the video content elements with better video delivery effects, resulting in inaccurate video analysis. Summary of the Invention

[0003] The main purpose of the present application is to provide a method, an apparatus, a computer device, and a storage medium for analyzing tags, which can improve the accuracy of video analysis.

[0004] To achieve the above object, according to one aspect of the present application, a method for analyzing tags is provided.

[0005] The method for analyzing tags according to the present application includes:

[0006] Obtaining video content tags of tagged video data;

[0007] Performing effect data analysis on the tagged video data to obtain an effect data analysis result;

[0008] Aligning the effect data analysis result and the video content tags to obtain combined data;

[0009] Invoking a pre-constructed tag analysis model, and inputting the combined data and the video content tags into the tag analysis model respectively to determine the importance degree of the video content tags.

[0010] Further, the tag analysis model includes a click-through rate prediction model and an attention model; inputting the combined data into the tag analysis model to determine the importance degree of the video content tags includes:

[0011] Determining the influence degree of each tag in the combined data on the click-through rate through the click-through rate prediction model;

[0012] Inputting the video content tags into the attention model, and outputting the attention distribution corresponding to the click-through rate of the tagged video data;

[0013] Determine the importance level of the video content tags according to the attention distribution and the influence degree of each tag pair on the click-through rate in the combined data.

[0014] Further, the inputting the video content tags into the attention model and outputting the attention distribution corresponding to the click-through rate of the tagged video data includes:

[0015] Input the video content tags into the attention model, and process the combined data into token data through the attention model;

[0016] Perform dimensionality reduction processing on each token data;

[0017] Fuse the dimensionality-reduced token data to obtain a fused sequence;

[0018] Extract the context features of the fused sequence, and calculate the attention distribution corresponding to the click-through rate of the tagged video data according to the context features.

[0019] Further, the performing effect data analysis on the tagged video data to obtain an effect data analysis result includes:

[0020] Calculate the click-through rate, consumption, and conversion number of the tagged video data;

[0021] Generate an effect data analysis result according to the click-through rate, delivery consumption, and conversion number.

[0022] Further, the method further includes:

[0023] Flatten the video content tags;

[0024] Statistically analyze the basic tag information corresponding to the flattened video content tags, and obtain a tag data analysis result according to the basic tag information.

[0025] Further, the method further includes:

[0026] Classify the video content tags corresponding to the tagged video data based on the effect data analysis result and the tag data analysis result to obtain multiple tag categories;

[0027] Perform time series analysis on the video content tags corresponding to each tag category to obtain the time period corresponding to each tag category;

[0028] Generate a video analysis result according to the importance level of the video content tags, the time period corresponding to each tag category, and the effect data analysis result.

[0029] Further, the video content tags for obtaining the labeled video data include:

[0030] Obtain the video data to be extracted, and extract the video features of the video data to be extracted;

[0031] Obtain the pre-constructed video tag system;

[0032] Perform multi-dimensional processing on the video features to obtain target features;

[0033] Match the target features with the preset tags in the video tag system to determine the video content tags corresponding to the video data to be extracted.

[0034] To achieve the above object, according to another aspect of the present application, there is provided an analysis device for tags.

[0035] The analysis device for tags according to the present application includes:

[0036] A tag acquisition module, configured to acquire the video content tags of the labeled video data;

[0037] An effect analysis module, configured to perform effect data analysis on the labeled video data to obtain an effect data analysis result;

[0038] An alignment module, configured to align the effect data analysis result and the video content tags to obtain combined data;

[0039] A tag analysis module, configured to call a pre-constructed tag analysis model, and input the combined data and the video content tags into the tag analysis model respectively to determine the importance degree of the video content tags.

[0040] A computer device includes a memory and a processor, the memory stores a computer program that can run on the processor, and when the processor executes the computer program, the steps in the above-mentioned method embodiments are implemented.

[0041] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0042] The above-mentioned tag analysis method, device, computer device and storage medium, by aligning the effect data analysis result and the video content tags corresponding to the labeled video data, is beneficial to subsequent analysis of the importance degree of the tags, and by analyzing the importance degree of the tags through a pre-constructed tag analysis model, the importance degree of the tags can be accurately predicted. Description of the Drawings

[0043] The accompanying drawings, which form a part of this application, are used to provide a further understanding of this application, making other features, objectives, and advantages of this application more obvious. The schematic embodiments and descriptions of the accompanying drawings of this application are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0044] Figure 1 is an application environment diagram of the label analysis method in an embodiment;

[0045] Figure 2 is a schematic flowchart of the label analysis method in an embodiment;

[0046] Figure 3 is a schematic diagram of the label structure of the video label system in an embodiment;

[0047] Figure 4 is a schematic diagram of the video data with a single label in an embodiment;

[0048] Figure 5 is a schematic diagram of a single label data in an embodiment;

[0049] Figure 6 is a schematic diagram of the analysis result of the effect data in an embodiment;

[0050] Figure 7 is a schematic diagram of the combined data in an embodiment;

[0051] Figure 8 is a schematic flowchart of the step of determining the importance degree of the video content label by inputting the combined data into the label analysis model in an embodiment;

[0052] Figure 9 is a scatter plot of the position distribution corresponding to the first-level label in an embodiment;

[0053] Figure 10 is a distribution diagram of the label categories in an embodiment;

[0054] Figure 11 is a schematic diagram of the time period corresponding to each label category when the content is stratified into voiceover and subtitles in an embodiment;

[0055] Figure 12 is a structural block diagram of the label analysis device in an embodiment;

[0056] Figure 13 is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0057] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0058] It should be noted that the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0059] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will detail this application with reference to the accompanying drawings and in conjunction with the embodiments.

[0060] The method for analyzing tags provided by this application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through a network. The server 104 obtains the tag analysis request sent by the terminal 102, parses the tag analysis request to obtain the video content tags of the tagged video data, performs effect data analysis on the tagged video data to obtain the effect data analysis result, aligns the effect data analysis result with the video content tags to obtain combined data, and then calls a pre-constructed tag analysis model, inputs the combined data and the video content tags into the tag analysis model respectively, and determines the importance degree of the video content tags. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0061] In one embodiment, as Figure 2 shown, a method for analyzing tags is provided. Taking the case where this method is applied to the Figure 1 server as an example, it includes the following steps 202 to step 208:

[0062] Step 202, obtain the video content tags of the tagged video data.

[0063] The tagged video data refers to the video data that has been annotated.

[0064] Specifically, the video content tags for obtaining the tagged video data include: obtaining the video data to be extracted, and extracting the video features of the video data to be extracted; obtaining the pre-constructed video tag system; performing multi-dimensional processing on the video features to obtain target features; matching the target features with the preset tags in the video tag system to determine the video content tags corresponding to the video data to be extracted. The video features of the video data to be extracted can be extracted through a feature extraction network, and the feature extraction network can be a network for video feature extraction such as the Inception-ResNet-v2 convolutional neural network model, the C3D network, etc. The video features are used to represent the content information of the video image, and the video features can include features in the time dimension and features in the space dimension.

[0065] The server stores a pre-constructed video tag system. The video tag system is established by experts. The video tag system can include content stratification, definition information of first-level tags, definition information of second-level tags, and definition information of third-level tags. The tag structure of the video tag system can be a tag tree, and the tag tree can be as Figure 3 shown.

[0066] The multi-dimensional processing can include stratification processing and classification processing. Performing multi-dimensional processing on the video features includes: performing stratification processing on the video features according to the video tag system to obtain the stratified features corresponding to the video data to be extracted, and performing classification processing on the stratified features according to the video tag system to obtain the target features.

[0067] Since the video data to be extracted includes content stratification and video content at different levels, in order to determine the content stratification corresponding to the data to be extracted, the video features can be subjected to stratification processing to obtain stratified features, and the stratified tags corresponding to the video data to be extracted can be determined through the stratified features. Then, classification processing is performed on the stratified features according to the video tag system, so as to obtain the target features. By performing stratification and classification processing on the video features, more accurate and detailed video features can be obtained, so as to match the corresponding video content tags according to the target features, improving the extraction accuracy of the video tags.

[0068] Match the target features with the preset tags in the pre-constructed video tag system, and use the successfully matched tags as the video content tags corresponding to the video data to be extracted, so as to label the video data to be extracted according to the video content tags. A single tagged video data can be as Figure 4 shown. The video content tags of a single tagged video data can include content stratification, first-level tags, second-level tags, third-level tags, etc. It is also possible to separately count the video content tags corresponding to the tagged video data to obtain each tag data, including the appearance layer, as well as the tag, second-level tag, and original text. A single tag data can be as Figure 5 shown.

[0069] By extracting the video features of the video data to be extracted, performing multi-dimensional processing on the video features to obtain target features, and matching the target features with the preset tags in the video label system, the video content tags corresponding to the video data to be extracted are determined. By performing multi-dimensional processing on the video features, more accurate and detailed video features can be obtained, so that the corresponding video content tags can be matched according to the target features, improving the extraction accuracy of video tags.

[0070] Step 204, perform effect data analysis on the labeled video data to obtain an effect data analysis result.

[0071] Effect data analysis can be called effect data EDA (Exploratory Data Analysis). The effect data analysis result can be obtained by calculating the click-through rate, conversion number, and consumption of the video, and plotting graphs based on the click-through rate, conversion number, and consumption. In this embodiment, a two-dimensional coordinate system graph can be constructed with the ID of the labeled video data as the x-axis and the consumption cost or duration as the y-axis. The sorting basis for each labeled video data on the x-axis can be selected from any one of the click-through rate, conversion number, or consumption. As Figure 6 shown, the click-through rate can be selected as the sorting basis to determine the order of each labeled video data on the x-axis, and the consumption cost is used as the y-axis to construct a coordinate system graph, so that the relationship between the click-through rate and the consumption cost of the video can be obtained.

[0072] Step 206, align the effect data analysis result and the video content tags to obtain combined data.

[0073] The alignment method can be to splice multiple video content tags to obtain a tag column, where the column value is the time and duration when the tag appears, separated by the "&" symbol, and multiple time points in one video are separated by the "," symbol. As Figure 7 shown, the column is the tag name, which is spliced by the appearance layer, first-level tag, second-level tag, and third-level tag. The 291 tag column values are the time and duration when the tag appears, separated by the "&" symbol. After aligning with the effect data analysis result, there are a total of 19 data. The relationship between the number of tags and the change in video CTR can be analyzed through the combined data.

[0074] Step 208, call the pre-constructed tag analysis model, input the combined data and the video content tags into the tag analysis model respectively, and determine the importance of the video content tags.

[0075] A label analysis model is pre-built in the server. The label analysis model can be composed of a click-through rate prediction model and an attention model. The combined data and the video content labels are respectively input into the label analysis model, so as to output the importance degree of the video content labels.

[0076] In this embodiment, by aligning the effect data analysis results with the video content labels corresponding to the labeled video data, it is beneficial to analyze the importance degree of the labels in the subsequent process. And by analyzing the importance degree of the labels through the pre-built label analysis model, the importance degree of the labels can be accurately predicted.

[0077] In one embodiment, the label analysis model includes a click-through rate prediction model and an attention model. As Figure 8 shown, the steps of determining the importance degree of the video content labels by inputting the combined data into the label analysis model include:

[0078] Step 802, determine the influence degree of each label in the combined data on the click-through rate through the click-through rate prediction model.

[0079] Step 804, input the video content labels into the attention model, and output the attention distribution corresponding to the click-through rate of the labeled video data.

[0080] Step 806, determine the importance degree of the video content labels according to the attention distribution and the influence degree of each label in the combined data on the click-through rate.

[0081] The label analysis model includes a click-through rate prediction model and an attention model.

[0082] The server can input the combined data into the click-through rate prediction model to determine the influence degree of each tag pair in the combined data on the click-through rate. The click-through rate prediction model can be trained using multiple regression models. For example, the regression models can include KNeighborsUnif, KNeighborsDist, LightGBMXT, LightGBM, RandomForestMSE, CatBoost, ExtraTreesMSE, NeuralNetFastAI, XGBoost, NeuralNetMXNet, LightGBMLarge, WeightedEnsemble_L2. Specifically, the data features considered during the training process are: the number of times a tag appears in a video. The combined data can be sorted according to this data feature, and the click-through rate of the sorted data can be predicted through the regression model. By comparing the predicted click-through rate with the actual click-through rate, the final model is determined as the click-through rate prediction model. Through the click-through rate prediction model, budget processing is performed on the combined data, and the influence degree of each tag on the click-through rate is output, including the importance score (importance) and significance score (p_valune) of each tag.

[0083] In the video content tags, the tags are composed of time series, and the features affecting the video click-through rate (CTR) effect do not exist independently. It is very likely that the influence is caused by the combined pattern of several tag orders. Based on the above problems, the server can convert the time series tags into natural language tasks and use the attention model to output the attention distribution. The attention model in this embodiment is an attention model based on the tag system, and the attention model can perform video CTR attribution processing.

[0084] In one of the embodiments, the video content tags are input into the attention model, and the attention distribution corresponding to the click-through rate of the tagged video data is output, including: inputting the video content tags into the attention model, and processing the combined data into token data through the attention model; performing dimensionality reduction processing on each token data; fusing the dimensionality-reduced token data to obtain a fused sequence; extracting the context features of the fused sequence, and calculating the attention distribution corresponding to the click-through rate of the tagged video data according to the context features.

[0085] Specifically, input the video content tags, the start time of each tag, and the tag duration into the attention model, and process the input data into token data, including: tag token, tag start time token, and tag duration ratio token. Perform dimensionality reduction processing on each token data to obtain the tag embedding, tag start time embedding, and tag duration ratio embedding corresponding to each tag. Integrate the tag embedding, tag start time embedding, and tag duration ratio embedding corresponding to each tag to obtain the tag fusion feature corresponding to each tag, thereby obtaining a fusion sequence. Extract the context features of the fusion sequence through the GRU network model, and use Attention Pooling to obtain the attention distribution attention weight of the CTR score of the video. Furthermore, determine the importance of all video content tags according to the attention distribution and the influence degree of each tag in the combined data on the click-through rate. Important tags can be selected from all video content tags for video analysis. For example, 20 tags can be selected as important tags, as shown in the following table:

[0086]

[0087] Furthermore, a feature can be added to each tag in the attention model: a weight that decreases over time, simulating the real attention tendency of people watching videos, which is beneficial to improving the effectiveness of tag importance analysis.

[0088] In this embodiment, the click-through rate prediction model can accurately determine the influence degree of each tag in the combined data on the click-through rate, while the attention model can output the attention distribution corresponding to the click-through rate of the tagged video data, and the influence degree of the combination of tag orders on the click-through rate can be obtained. Thus, the importance of video content tags can be determined according to the attention distribution and the influence degree of each tag in the combined data on the click-through rate. Therefore, the influence of tags on the video delivery effect can be accurately predicted.

[0089] In one embodiment, perform effect data analysis on the tagged video data to obtain the effect data analysis result, including: calculating the click-through rate, consumption, and conversion number of the tagged video data; generating the effect data analysis result according to the click-through rate, delivery consumption, and conversion number.

[0090] Effect data analysis can be called effect data EDA (Exploratory Data Analysis). By calculating the click-through rate, conversion number, and consumption of the video, plotting graphs based on the click-through rate, conversion number, and consumption, the effect data analysis results can be obtained. In this embodiment, a two-dimensional coordinate system graph can be constructed with the ID of the labeled video data on the x-axis and the consumption cost or duration on the y-axis. The sorting basis for each labeled video data on the x-axis can be selected from any one of the click-through rate, conversion number, or consumption.

[0091] In one embodiment, the above method further includes: performing label data analysis on the labeled video content tags. Specifically, flatten the video content tags corresponding to the labeled video data; count the basic label information corresponding to the flattened video content tags, and obtain the label data analysis results based on the basic label information.

[0092] Label data analysis is called label data EDA (Exploratory Data Analysis). Label data analysis can include analyzing the basic label information corresponding to the video content tags, including analysis of video duration distribution, analysis of the number of times each video content tag is mentioned, analysis of the position distribution of each video content tag in the video, etc.

[0093] First, the video content tags corresponding to the labeled video data can be flattened. The video content tags corresponding to each video are in an Excel file, and the tags are in a two-dimensional table and include a time dimension. For convenient analysis, the data is flattened. The flattening process uses the secondary label as the smallest label unit for data analysis, and the tertiary label and subsequent labels as the specific values of the smallest label unit.

[0094] Exemplarily, the video content tags corresponding to the labeled video data can be as shown in the following table:

[0095]

[0096] Among them, content stratification indicates the carrier position where the label appears, such as the label appears in the voiceover, picture, etc.; the primary label indicates the root label of the label tree, which can be further divided into secondary labels; the secondary label indicates the leaf label in the label tree; the tertiary label, etc., indicates the specific value in the superior label. Using the leaf label as the smallest label unit for data analysis and the tertiary label as the specific value of the smallest label unit, the video content tags are flattened. The flattened data is as follows:

[0097]

[0098] After data flattening, each row represents a video data and each column represents a content label. The cell contains the label value, the time when the label appears, and the duration.

[0099] Analyze the video duration distribution of the labeled video data. Specifically, construct a video duration distribution histogram based on the labeled video data, calculate the average video duration according to the histogram, and determine the shortest and longest video durations.

[0100] The server can also count the number of times each video content label is mentioned, and determine the labels that are mentioned more or less based on the counted number of times the labels are mentioned. For example, by counting the number of times the labels are mentioned, among the content hierarchical labels, the more mentioned labels appear in the oral broadcast and subtitle layer, 137 times, accounting for about 44%, followed by the picture display layer, 74 times, accounting for about 24%, and the labels mentioned in the caption and content creativity layer are the least. Among the level labels, labels of types such as conversion stimulus, product display, efficacy description, brand information, and product basic information are mentioned more, while labels of types such as user psychology, preferential selling points, and influencer call are mentioned less.

[0101] Furthermore, separately count the position distributions of various labels such as content hierarchical labels, first-level labels, second-level labels, and third-level labels that appear in the video. Specifically, perform a position distribution analysis on each type of label separately to obtain the time points and durations when each label in each type of label appears in the video. Represent the duration with scatter points, and construct a position scatter plot corresponding to each type of label according to the time points and durations when each label in each type of label appears in the video, so as to determine the distribution characteristics of each type of label based on the position scatter plot. The distribution characteristics can include at the beginning / middle / end, labels that are mentioned more and have a longer duration, labels that are mentioned less, and labels that are mentioned. Exemplarily, as Figure 9 shown, it is the position distribution scatter plot corresponding to the first-level label. According to this figure, it can be obtained that at the beginning, conversion stimulus, conversion purpose, brand information, product basic information, and body text labels are mentioned more and have a longer duration; at the end, pain point description, efficacy description, influencer recommendation, etc. are all mentioned.

[0102] Another example is for the position scatter plot corresponding to the content hierarchical type labels. The analysis shows that at the beginning, the content labels mainly appear in the package frame brand area layer, paragraph layer, oral broadcast and subtitle layer, and picture display layer. Among them, the content labels in the package frame brand area layer and paragraph layer have a longer duration, and the labels in the oral broadcast and subtitle layer have a shorter duration. At the end, there are labels that appear in the oral broadcast and subtitle layer and paragraph layer. The label appearance rate is less around 40 - 50 seconds in the middle. Consider checking whether there are statistical errors in the code.

[0103] After analyzing the basic information of tags such as the analysis of the video duration distribution, the analysis of the number of times each video content tag is mentioned, and the analysis of the position distribution of each video content tag in the video, the analysis results are used as the tag data analysis results.

[0104] In this embodiment, by flattening the video content tags corresponding to the tagged video data, it is beneficial to perform the analysis of the basic information of tags in the subsequent process. The basic information of tags corresponding to the flattened video content tags is counted, and the tag data analysis results are obtained according to the basic information of tags. Through the analysis of the video duration distribution, the analysis of the number of times each video content tag is mentioned, and the analysis of the position distribution of each video content tag in the video, it is possible to comprehensively and accurately analyze and obtain the occurrence time and other situations of the video content tags in the video.

[0105] In one embodiment, the above method further includes: generating a video analysis result based on the effect data analysis result, the tag data analysis result, and the importance of the video content tags. Specifically, based on the effect data analysis result and the tag data analysis result, the video content tags corresponding to the tagged video data are classified to obtain multiple tag categories; the time series analysis is performed on the video content tags corresponding to each tag category to obtain the time periods corresponding to each tag category; and the video analysis result is generated according to the importance of the video content tags, the time periods corresponding to each tag category, and the effect data analysis result.

[0106] Based on the effect data analysis result and the tag data analysis result, analyze the quality of the tags, the time when good / bad tags appear, and the importance of the tags, and generate a target analysis result according to the analysis data. The target analysis result can represent the relationship between the video content tags and the placement effect.

[0107] The quality of analysis tags can be based on the results of effect data analysis and tag data analysis, calculating the hit count of tags, the rising / falling hit count, and the tag ctr (Click-Through-Rate). Among them, the hit count indicates the number of times the tag appears in the video, the rising / falling hit count indicates the number of times the tag appears during the period when the video ctr rises / falls, and the tag ctr represents the weighted average ctr of all videos containing the tag. Thus, the video content tags corresponding to the tagged video data can be classified according to the hit count of tags, the rising / falling hit count, and the tag ctr, resulting in multiple tag categories. Further, tags with a tag ctr higher than the median ctr of all tags and a rising hit count greater than the falling hit count are classified as category A. Tags with a tag ctr lower than the median ctr of all tags and a rising hit count greater than the falling hit count are classified as category B. Tags with a tag ctr higher than the median ctr of all tags and a rising hit count less than the falling hit count are classified as category C. Tags with a tag ctr lower than the median ctr of all tags and a rising hit count less than the falling hit count are classified as category D, obtaining 4 categories of tags, with category A tags being the best and category D tags being the worst. The distribution diagram of tag categories can be as Figure 10 shown. Category A tags may include voiceover and subtitles - fragrance - pleasant smell (no aspect) - pleasant smell, voiceover and subtitles - conversion purpose - guide to purchase - buy now immediately. Category B tags may include voiceover and subtitles - efficacy description - fluffy - naturally fluffy. Category C tags may include voiceover and subtitles - efficacy description - fluffy - twice as fluffy. Category D tags may include voiceover and subtitles - efficacy description - fluffy - fluffy.

[0108] The server then performs a time series analysis on the video content tags corresponding to each tag category. Specifically, it examines the time series data of the tags appearing in the four regions A, B, C, and D corresponding to each content layer, and obtains the positions of the better tags (AB) for each content layer. For example, when the content layer is voiceover and subtitles, the schematic diagram of the time periods corresponding to each tag category can be as Figure 11 shown. By performing a time series analysis on the video content tags corresponding to each tag category, the time when the better tags should appear can be obtained.

[0109] Furthermore, a video analysis result is generated based on the importance of the video content tags, the time periods corresponding to each tag category, and the results of the effect data analysis.

[0110] In this embodiment, based on the result of the effect data analysis and the result of the label data analysis, the video content labels corresponding to the labeled video data are classified to obtain multiple label categories, and good and bad labels can be determined. By performing a time series analysis on the video content labels corresponding to each label category to obtain the time periods corresponding to each label category, it is possible to determine the time when good and bad labels should appear. Based on the result of the effect data analysis, the importance of the video content labels corresponding to the labeled video data is analyzed, and a video analysis result is generated according to the importance of the video content labels, the time periods corresponding to each label category, and the result of the effect data analysis, and the influence degree of the label on the placement effect can be obtained.

[0111] In one embodiment, generating a video analysis result according to the importance of the video content labels, the time periods corresponding to each label category, and the result of the effect data analysis may include: using the importance of the video content labels, the time periods corresponding to each label category, and the result of the effect data analysis as the target analysis result, and extracting the key information corresponding to the video data to be extracted according to the target analysis result to obtain the insight of each video data. For example, the insight of each video data may include: Insight 1: The user loss is serious after 3 seconds, and the first 3 seconds and the first 10 seconds are the golden periods for the video display content. Specifically, the difference between high click-through rate and low click-through rate lies in whether key data is densely displayed within 10 seconds, and the product display brand information should be emphasized in the first 3 seconds. Insight 2: All video effects will peak within 3 seconds. Optimizing the first 3 seconds has an overall improvement effect on the content, and optimizing the user loss from 3 to 30 seconds is also an effective direction. Insight 3: The conversion rate tends to decrease when the video duration is too long, and 20 - 30 seconds is a more reasonable video duration. Specifically, using more mixed-cut videos of 20 - 30 seconds will have a more stable effect than longer voice-over videos.

[0112] Furthermore, it is also possible to count the ranking information of the preset video content types, including: average click-through rate ranking, average conversion rate ranking, comprehensive index ranking, existing video material quantity ranking, click-through rate variance ranking, conversion rate variance ranking. The comprehensive index calculation formula is: 60% * click-through rate + 30% * conversion rate - 10% * average click cost. The preset video content types may include plot, influencer voice-over, mixed cut, and single-person voice-over - star, etc. Determine the video placement strategy according to the counted ranking information of the preset video content types and the insight of each video data. For example, if the insight is that the conversion rate tends to decrease when the video duration is too long, and 20 - 30 seconds is a more reasonable video duration, then more mixed-cut videos of 20 - 30 seconds can be used, which will have a more stable effect than longer voice-over videos.

[0113] Summarize the insights of multiple video data to obtain video analysis results. For example, the summary table can be as follows:

[0114]

[0115] In this embodiment, determine the key information corresponding to the video data to be analyzed from the importance level of video content tags, the time periods corresponding to each tag category, and the effect data analysis results, and then generate video analysis results. This can improve the accuracy of video analysis, quickly determine the content tags with better video placement effects, and the production methods of video content.

[0116] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0117] In one embodiment, as Figure 12 shown, a tag analysis device is provided, including: a tag acquisition module 1202, an effect analysis module 1204, an alignment module 1206, and a tag analysis module 1208, where:

[0118] The tag acquisition module 1202 is used to acquire the video content tags of the tagged video data.

[0119] The effect analysis module 1204 is used to perform effect data analysis on the tagged video data to obtain effect data analysis results.

[0120] The alignment module 1206 is used to align the effect data analysis results and the video content tags to obtain combined data.

[0121] The tag analysis module 1208 is used to call a pre-constructed tag analysis model, input the combined data and the video content tags into the tag analysis model respectively, and determine the importance level of the video content tags.

[0122] In one embodiment, the tag analysis model includes a click-through rate prediction model and an attention model; the tag analysis module 1208 is further used to determine the influence degree of each tag in the combined data on the click-through rate through the click-through rate prediction model; input the video content tags into the attention model, and output the attention distribution corresponding to the click-through rate of the tagged video data; determine the importance level of the video content tags according to the attention distribution and the influence degree of each tag in the combined data on the click-through rate.

[0123] In one embodiment, the label analysis module 1208 is further configured to input the video content labels into an attention model, process the combined data into token data through the attention model; perform dimensionality reduction processing on each token data; fuse the token data after dimensionality reduction to obtain a fused sequence; extract the context features of the fused sequence, and calculate the attention distribution corresponding to the click-through rate of the labeled video data according to the context features.

[0124] In one embodiment, the effect analysis module 1204 is further configured to calculate the click-through rate, consumption, and conversion number of the labeled video data; generate an effect data analysis result according to the click-through rate, placement consumption, and conversion number.

[0125] In one embodiment, the above device further includes: a label analysis module, configured to flatten the video content labels; count the basic label information corresponding to the flattened video content labels, and obtain a label data analysis result according to the basic label information.

[0126] In one embodiment, the above device further includes: a video analysis module, configured to classify the video content labels corresponding to the labeled video data based on the effect data analysis result and the label data analysis result to obtain multiple label categories; perform time series analysis on the video content labels corresponding to each label category to obtain the time period corresponding to each label category; generate a video analysis result according to the importance degree of the video content labels, the time period corresponding to each label category, and the effect data analysis result.

[0127] In one embodiment, the label acquisition module 1202 is further configured to acquire video data to be extracted, extract the video features of the video data to be extracted; acquire a pre-constructed video label system; perform multi-dimensional processing on the video features to obtain target features; match the target features with the preset labels in the video label system to determine the video content labels corresponding to the video data to be extracted.

[0128] For the specific limitations of the label analysis device, reference can be made to the limitations of the label analysis method in the foregoing text, which will not be elaborated here. Each module in the above label analysis device can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0129] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 13As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data of an analysis method for a kind of label. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes an analysis method for a kind of label.

[0130] Those skilled in the art can understand that Figure 13 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0131] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it realizes the steps in the above-mentioned various embodiments.

[0132] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it realizes the steps in the above-mentioned various embodiments.

[0133] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0134] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0135] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for analyzing a label, characterized in that Including: A video content tag for obtaining tagged video data; Performing effect data analysis on the tagged video data to obtain an effect data analysis result; Aligning the effect data analysis result and the video content tag to obtain combined data; Invoking a pre-constructed tag analysis model, and inputting the combined data and the video content tag into the tag analysis model respectively to determine the importance degree of the video content tag; The tag analysis model includes a click-through rate prediction model and an attention model; Inputting the combined data into the tag analysis model to determine the importance degree of the video content tag, including: Determining the influence degree of each tag in the combined data on the click-through rate through the click-through rate prediction model; Inputting the video content tag into the attention model, and outputting an attention distribution corresponding to the click-through rate of the tagged video data; Determining the importance degree of the video content tag according to the attention distribution and the influence degree of each tag in the combined data on the click-through rate.

2. The method according to claim 1, characterized in that, The inputting the video content tag into the attention model and outputting an attention distribution corresponding to the click-through rate of the tagged video data includes: Inputting the video content tag into the attention model, and processing the combined data into token data through the attention model; Performing dimensionality reduction processing on each token data; Fusing the dimensionality-reduced token data to obtain a fusion sequence; Extracting context features of the fusion sequence, and calculating an attention distribution corresponding to the click-through rate of the tagged video data according to the context features.

3. The method according to claim 1, wherein The performing effect data analysis on the tagged video data to obtain an effect data analysis result includes: Calculating the click-through rate, consumption, and conversion number of the tagged video data; Generating an effect data analysis result according to the click-through rate, delivery consumption, and conversion number.

4. The method according to claim 1, wherein The method further includes: Performing a flattening process on the video content tag; Counting the basic tag information corresponding to the flattened video content tag, and obtaining a tag data analysis result according to the basic tag information.

5. The method according to claim 4, wherein The method further includes: Classifying the video content tags corresponding to the tagged video data based on the effect data analysis result and the tag data analysis result to obtain multiple tag categories; Performing a time series analysis on the video content tags corresponding to each tag category to obtain a time period corresponding to each tag category; Generating a video analysis result according to the importance degree of the video content tag, the time period corresponding to each tag category, and the effect data analysis result.

6. The method according to any one of claims 1 to 5, characterized in that The obtaining a video content tag of tagged video data includes: Obtaining video data to be extracted, and extracting video features of the video data to be extracted; Obtaining a pre-constructed video tag system; Performing multi-dimensional processing on the video features to obtain target features; Matching the target features with preset tags in the video tag system to determine the video content tag corresponding to the video data to be extracted.

7. An analysis device for a label, characterized in that, The device includes: A tag acquisition module for obtaining a video content tag of tagged video data; An effect analysis module for performing effect data analysis on the labeled video data to obtain an effect data analysis result; An alignment module for aligning the effect data analysis result and the video content label to obtain combined data; A label analysis module for calling a pre-constructed label analysis model, inputting the combined data and the video content label into the label analysis model respectively, and determining the importance degree of the video content label; The label analysis model includes a click-through rate prediction model and an attention model; inputting the combined data into the label analysis model to determine the importance degree of the video content label includes: Determining the influence degree of each label in the combined data on the click-through rate through the click-through rate prediction model; Inputting the video content label into the attention model and outputting the attention distribution corresponding to the click-through rate of the labeled video data; Determining the importance degree of the video content label according to the attention distribution and the influence degree of each label in the combined data on the click-through rate.

8. A computer device, comprising a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Video analysis method and device, electronic equipment and readable storage medium

    CN112948635A