Analysis method for quantitatively analyzing interaction and sales indexes of content labels

By analyzing the content database of social media platforms and e-commerce product databases and calculating the interaction and sales indicators of content labels, the problem of difficulty in evaluating the impact of marketing content on product sales in the existing technology is solved, and more accurate marketing content creation and resource conservation are achieved.

WO2025123691A1PCT designated stage expired Publication Date: 2025-06-19NINT (SHANGHAI) CO LTD

Patent Information

Application Number
PCT/CN2024/108400
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-14
Filing Date
2024-07-30
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The prior art is difficult to quantify the impact of content labels on commodity sales of marketing content, resulting in inaccurate creation of marketing content and difficulty in assessing the indirect impact of content on consumer purchasing decisions.

Method used

By obtaining the content database of social media platforms and the e-commerce product database, the content in the content database is labeled, the content interaction indicators and product sales indicators of the content tag are calculated, and visual analysis is carried out.

Benefits of technology

Quantitative evaluation of content labels in content interaction and product sales is achieved, making marketing content creation simpler, feasible and accurate, saving manpower and resource costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024108400_19062025_PF_FP_ABST
    Figure CN2024108400_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the field of content marketing, and disclosed is an analysis method for quantitatively analyzing interaction and sales indexes of content labels. In order to solve the problem that creation of existing marketing content is difficult and inaccurate, the present invention provides an analysis method for quantitatively analyzing the interaction effect of content labels on marketing content and the contribution of the content labels to product sales, comprising: obtaining a content database, labeling content in the content database to obtain content labels, and adding product SPU labels to the content database; obtaining an e-commerce product database, and adding product SPU labels to the product database; calculating content interaction indexes of all the content labels and product sales indexes of all the content labels; and performing visual analysis on content interaction indexes and product sales indexes. According to the present invention, the effects of the content labels on content interaction and product sales are quantitatively evaluated on the basis of the content interaction indexes and the product sales indexes, and the content labels are quantitatively compared, such that the creation of the marketing content is simple, feasible and accurate, thereby saving costs.
Need to check novelty before this filing date? Find Prior Art

Description

A method for quantitatively analyzing the interaction and sales indicators of content tags Technical Field

[0001] The present invention belongs to the technical field of content marketing, and more specifically, relates to an analysis method for quantitatively analyzing the interaction and sales indicators of content tags. Background Art

[0002] Product marketing content on social media drives sales on e-commerce platforms within social media platforms. Marketing content itself is highly variable. To quantitatively analyze its elements, it can be broken down into structured content tags (content elements). These tags can then be analyzed for metrics related to content engagement (including exposure, likes, reposts, reviews, and favorites). Brands and content creators alike are interested in obtaining quantitative assessments of content tags' impact on content engagement and product sales to guide their marketing content creation. However, quantifying and analyzing the impact of marketing content tags on product sales has been a challenging topic within the content marketing industry.

[0003] Marketing content has both direct and indirect impacts on product sales. Some marketing content directly links to products on the site, allowing consumers to click and purchase directly. However, some consumers may not purchase immediately and may instead make their own purchase later. Another type of marketing content is "seeding content" and does not directly provide product links. In this case, the impact of marketing content on consumer purchasing decisions is even more difficult to assess. Typically, platform backends provide sales data for products linked to products, but this data is held only by the platform itself and is not publicly available. It is difficult for organizations outside the platform to analyze industry data. Furthermore, it is often difficult for brands and content creators to quantify the indirect impact of marketing content on consumer purchases.

[0004] Corresponding improvements have been made to address the above issues. For example, Chinese patent application number CN202110787436.X, published on October 29, 2021, discloses a smart bank multi-channel collaborative marketing system and method, including: collecting data from proprietary channels and marketing channels to obtain user information; analyzing user information to obtain valid information; matching valid information with bank products; and using the bank's proprietary channels to deliver product marketing content. The patent's shortcomings are: the accuracy of targeted delivery is poor, and the matching accuracy needs to be improved.

[0005] Another example is Chinese patent application number CN202110319997.7, published on June 15, 2021. The patent discloses a self-service video marketing management system, including: a web crawler module for crawling the corresponding marketing video playback volume, likes and comments on major video playback platforms based on preset marketing video feature parameters; an audience group positioning module for analyzing the account information of the likers and commenters, thereby locating the audience group characteristics adapted to the marketing video and building an audience group feature configuration model; a comment analysis module for mining valuable comment data and processing and analyzing the comment data; a video content improvement suggestion module for generating corresponding video content improvement suggestions based on the processing and analysis results of the comment data; a marketing video targeted delivery module for achieving targeted delivery of marketing videos based on the positioning results of the audience group positioning module. The shortcoming of this patent is that although it can achieve targeted delivery of marketing videos based on the positioning results, the accuracy is poor.

[0006] Summary of the Invention

[0007] 1. Problems to be solved

[0008] To address the difficulty and inaccuracy of existing marketing content creation, this paper provides a method for quantitatively analyzing the interaction and sales metrics of content tags. This method quantitatively evaluates the impact of content tags on content interaction and product sales using content interaction and product sales metrics, and quantitatively compares content tags, making the creation of marketing content simple, feasible, and accurate, saving manpower and resource costs.

[0009] 2. Technical solution

[0010] To solve the above problems, the present invention adopts the following technical solutions.

[0011] A method for quantitatively analyzing interaction and sales indicators of content tags includes the following steps:

[0012] Obtaining a content database of a social media platform, labeling the content in the content database to obtain content tags, and adding a product SPU (Standard Product Unit) tag to the content database;

[0013] Obtain the e-commerce product database of the e-commerce platform within the social media platform and add product SPU tags to the product database;

[0014] Calculate content engagement metrics for all content tags and calculate product sales metrics for all content tags;

[0015] Conduct visual analysis of content interaction metrics and product sales metrics.

[0016] Furthermore, the calculation of the content interaction index includes the following steps:

[0017] Determine the product category in the content database and count the number of contents n of all content tags k in the product category within the set time period to obtain the content set C k ={content1, content2,…, contentn};

[0018] Calculate the interaction value E of each content in the content collection i :

[0019] E i = Number of likes for each content + Number of reposts for each content + Number of favorites for each content + Number of comments for each content; or

[0020] E i = Number of likes for each content, number of reposts for each content, number of collections for each content, or number of reviews for each content;

[0021] Calculate the content interaction index Y for content tag k k :

[0022] or

[0023] Furthermore, the calculation of the commodity sales index includes the following steps:

[0024] Determine the product category in the content database and count the number of contents n of all content tags k in the product category within the set time period to obtain the content set C k ={content1, content2,…, contentn};

[0025] Determine the m product SPUs corresponding to a single content i in the content set, and obtain the product SPU set P corresponding to a single content i i ={SPU1, SPU2, ..., SPU m};

[0026] Determine the sales volume of a single product SPU on each date within a set time period;

[0027] Determine a date d, a single content i for a single product SPU j Sales contribution value:

[0028] Among them, Sales(j,d) is the product SPU j Sales on date d; W i For a single content i on date d, the product SPU j The contribution weight of sales; u is the number of SPUs mentioned within the effective time window for date dj The total number of contents;

[0029] Determine the relationship between a single content i and the product P mentioned in the content i ={SPU1, SPU2, ..., SPU m Total sales contribution:

[0030] Where t is the number of days in the time window from the release date that the calculated content affects the product sales;

[0031] Determine the content set C corresponding to the content tag k k Cumulative value of impact on product sales:

[0032] Calculate the sales index X of content tag k k : or Where p is the content set C corresponding to the content label k k The corresponding total SPU number of duplicate products.

[0033] Furthermore, visual analysis of content engagement metrics and product sales metrics includes the following steps:

[0034] Create a two-dimensional coordinate system, with the sales indicator corresponding to the content tag as the X-axis and the interaction indicator corresponding to the content tag as the Y-axis;

[0035] Determine the mean of the sales index and interaction index of all content tags, divide them into four areas based on the mean of the sales index and interaction index, and place the sales index and interaction index into the four areas respectively.

[0036] Furthermore, the tagging of the content in the content database to obtain the content tag specifically includes the following steps:

[0037] Build a product knowledge graph;

[0038] Build a content category database; obtain a media content database, use the product knowledge graph to filter out the content data related to each product category from the media content database, and build a content category database;

[0039] Extract information from the content category database to construct a category content tag tree, and tag the data in the category database according to the category content tag tree. This step specifically includes the following steps:

[0040] First, the RaNER model is used to extract person entities, category entities, brand entities, and product attribute entities from the content category database. Then, a large language model is used in combination with information extraction prompts and thought chain summary prompts to perform semantic recognition on the category database and extract person entities, internet buzzword entities, user pain point entities, product feature entities, and applicable entities. Finally, the entity results are fused to obtain the final entity.

[0041] The final entity is converted into a word vector using a text vectorization model. A clustering algorithm is then used to obtain several categories of word vectors. The words in each category are then grouped into one or more tags using a large language model. The keyword types output by the large language model are used to construct a tree-structured content tag tree, along with keywords for each tag.

[0042] Label the content text in the category database according to the category content label tree.

[0043] Furthermore, tagging the content text includes the following steps:

[0044] When labeling the content text, determine whether the content text has been subjected to entity extraction; if entity extraction has been performed, the entity words become candidate tags; if entity extraction has not been performed, perform entity extraction and add the tags corresponding to the identified entity words to the candidate tag set;

[0045] And for the keywords or regular expressions corresponding to each tag in the tag tree, keyword matching and regular expression matching are used, and the matched tags are also added to the candidate tag set of the content text;

[0046] Use the large language model to use the discriminant prompt to judge all candidate tags that have been screened out for the content text to determine whether the candidate tags match the meaning of the corresponding content text; if they match, confirm the candidate tag; if not, make corrections.

[0047] Furthermore, building a content category database specifically includes the following steps:

[0048] Collect content information from various social media platforms to form a media content database;

[0049] Use the product knowledge graph to perform text matching on the text information in the media content database to establish a preliminary screening database for product categories;

[0050] Convert the image-type content and video-type content in the initial screening database into text content respectively;

[0051] Perform fine screening and classification on the initial screening database to determine whether the text content is relevant to the product category.

[0052] Furthermore, the media content database only stores original text description information: for graphic content, the content title, content text and picture link content are stored; for video content, the content title and video link content are stored.

[0053] 3. Beneficial effects

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] (1) The present invention obtains the content database of the social media platform and the e-commerce product database within the social media platform, labels the content in the content database, and then calculates the content interaction index and product sales index of each content tag based on the e-commerce product database, and then performs a visual analysis on them; the entire data source is only the data publicly available on the social media page, and does not rely on tracking user behavior links, which greatly broadens the scope of users; at the same time, the content interaction index and product sales index are used to quantitatively evaluate the impact of content tags on content interaction and product sales, thereby making a quantitative comparison of content tags. When brands or content creators need to create product marketing content, they can choose tags with relatively high quality in the category, making the creation of marketing content simple, feasible and accurate, saving labor costs and resource costs;

[0056] (2) The present invention constructs a commodity knowledge graph and uses the entities and relationships in the commodity knowledge graph to construct a content category database from the media content database, so that the construction efficiency of the content category database is fast and work efficiency is effectively improved; after the construction of the category database is completed, information is extracted from it to construct a content tag tree, and when performing information extraction, the RaNER model is used to identify concrete entities, and then the large language model is used in combination with the information extraction prompt and the thinking chain summary prompt to identify abstract entities. Different types of entities are identified and extracted using different models, which effectively compensates for the problems of inaccurate extraction and recognition and incomplete entity recall caused by entity extraction using a single model; finally, the content text is labeled using the content tag tree; the entire process is simple and not cumbersome, the entity extraction accuracy is high, and the overall labeling accuracy is high; at the same time, the efficiency is high, reducing labor costs and time costs;

[0057] (3) When labeling a content text, the present invention first determines whether entity extraction should be performed on the content text to improve work efficiency and save time; entity words that have been entity extracted become candidate tags, and entity words that have not been entity extracted become candidate tags after entity extraction. At the same time, the obtained candidate tag set is judged to avoid the possibility of potential errors, and the recall rate is improved as much as possible (reducing omissions). At the same time, the semantic understanding ability of the large language model is used to filter out entities that are matched by keywords but are semantically incorrect, and to filter out entity words that do not exist in the original text that may be given by the large language model as much as possible, thereby improving the precision rate;

[0058] (4) When constructing a content category database, the present invention includes first performing a preliminary screening of the media content database to obtain a preliminary screening database, then converting the pictures and video contents in the preliminary screening database into text contents respectively, and finally performing fine screening and classification to obtain a final category database; the processing flow of first performing a low-cost and fast preliminary screening, narrowing the data range and then performing fine screening reduces costs while improving efficiency while ensuring overall accuracy; and only the original text description information is stored in the media content database, which greatly reduces storage costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] FIG1 is a schematic flow diagram of the present invention;

[0060] FIG2 is a schematic diagram of the process of calculating the sales index of content tags;

[0061] FIG3 is a schematic diagram showing the calculation process of the contribution value of content to product sales. DETAILED DESCRIPTION

[0062] The present invention is further described below with reference to specific embodiments and accompanying drawings.

[0063] Please refer to Figures 1 and 2. Figure 1 is a flowchart of the present application; Figure 2 is a flowchart of the sales index calculation process of content tags.

[0064] In this embodiment, as shown in FIG1 , a method for quantitatively analyzing the interaction and sales indicators of content tags includes the following steps:

[0065] S1: Obtain the content database of the social media platform and label the content in the database to obtain content tags. The content database stores content information and the amount of interaction (including the number of likes, comments, reposts, and favorites) by date. It also stores the product SPUs identified in the content, the content tags assigned to the content according to the structured tag tree, and the relationship between the product SPUs and the content tags describing them.

[0066] Obtain an e-commerce product database of an e-commerce platform within a social media platform: The e-commerce product database is an e-commerce product database of clean brands, categories, and SPUs (Standard Product Units) of the e-commerce platform within the social media platform.

[0067] In a specific example, the content data in the content database is shown in Table 1 below; the interactive data of the content data by date is shown in Table 2; the product SPU corresponding to the content data and the content tag describing the product SPU are shown in Table 3 below:

[0068] Table 1 Content database table

[0069] Table 2 Interactive data table of content data by date

[0070] Table 3 Data table of commodity SPU corresponding to content data and content tags describing the commodity SPU

[0071] As can be seen from Tables 1 to 3, Content 1 mentions two products, namely SPU1 and SPU2. Two content labels are identified in the description of SPU1, namely Efficacy-Whitening and Efficacy-Freckle Removal; the description of SPU2 mentions the content label Efficacy-Moisturizing.

[0072] In a specific example, an e-commerce product database stores product information, including cleaned fields such as brand, category, and SPU. It also stores product sales and sales figures by date. A product ID corresponds to a product link on the e-commerce platform, and an SPU may correspond to multiple product links. The e-commerce product data in the database is shown in Table 4 below, and the sales figures for the product data by date are shown in Table 5 below.

[0073] Table 4 E-commerce product data table in the e-commerce product database

[0074] Table 5 Sales data of commodity data by date

[0075] In a specific embodiment, labeling the content in the content database according to the structured tag tree to obtain content tags specifically includes the following steps:

[0076] S11: Build product knowledge graph;

[0077] In a specific embodiment, the construction of a product knowledge graph can be achieved by collecting product information from various e-commerce platforms (product information includes platform, product ID, product name, specifications, product parameters, etc.) to form a product database; using the RaNER model to perform entity recognition on the product text information in the product database to obtain brand, category, product attribute entities and related keywords, and using the co-occurrence relationship of entities in the products to construct a product knowledge graph; the structure of the product knowledge graph is in the form of a triple of (entity, relationship, entity).

[0078] S12: Constructing a content category database; obtaining a media content database, and using the product knowledge graph to filter out content data related to each product category from the media content database to construct a content category database;

[0079] Specifically, building a content category database includes the following steps:

[0080] S121: Collect content information from various social media platforms to form a media content database. The content on social media primarily includes text, images, and videos. In this step, the media content database stores only original text description information: for text, images, and images, the title, text, and image link are stored; for video, the title and video link are stored. The media content database stores only original text description information and does not include multimedia data such as images and videos, thereby significantly reducing storage costs.

[0081] S122: Use the product knowledge graph to perform text matching on the text information in the media content database to establish a preliminary screening database for product categories. It should be noted that since the content information on social media includes topics from various fields, including those related to or unrelated to products, it is necessary to screen out content related to products in a certain category. At the same time, since the media content database contains massive amounts of data (more than 1 billion items), and the content types include not only text but also video types, which need to be converted to text for further processing. The computing cost of converting videos to text is high and time-consuming, and using AI models to classify massive amounts of data is also high. Therefore, text matching is first performed on the text information to obtain a preliminary screening database. Low-cost and rapid preliminary screening narrows the data scope, ensuring both efficiency and reducing costs.

[0082] In the initial screening stage, keywords of various entities in the product knowledge graph, including keywords of related entities such as brands, categories, and attributes, are used to perform text matching on the text information in the media content database. This step only matches the video titles of the video content, thereby quickly filtering out content data related to products in a certain category from massive content data and establishing an initial screening database for a certain category. A category initial screening database can be established for all target categories required for the business, such as a category initial screening database for beauty and personal care products.

[0083] S123: Convert the image-type content and video-type content in the preliminary screening database into text content respectively; since the preliminary screening database in step S122 contains image-type and video-type content, only the text information such as the title is determined. In order to further analyze the text information related to the products contained in the images and videos, it is necessary to convert the content in the images and videos into text;

[0084] Specifically, for image content, OCR technology is used to convert the text in the image into text content; for video content, OCR technology and ASR technology are used respectively to convert the video into text content: before performing OCR processing, the video is framed at a certain time interval, such as one frame per second, to convert the video into a group of images, and OCR technology is used to convert the text in the images into text content. The position and size information of the text is used to filter out minor text information in the background as much as possible, while retaining important content text such as subtitles; and ASR technology is used to convert the video voice into text content, thereby adding an (OCR text, ASR text) field to each video content; the two text contents are combined to obtain the final text content of the video content; since the text and voice on the video screen each contain some language information, and each also lacks some language information, and OCR and ASR technologies may have certain error rates, converting a video into text using OCR and ASR at the same time and using them together later is conducive to a more complete analysis of the language information in the video.

[0085] S124: Perform fine screening and classification on the primary screening database to determine whether the text content is relevant to the product category. After obtaining the text content of the image content and video content in the primary screening database in step S123, it is necessary to further determine whether the content is relevant to the product data of this category; because the primary screening uses keywords to quickly match to narrow the data range in massive data and recall potentially relevant content, there will be ambiguous keywords. For example, the milk brand "Guangming" may match "sunny" with keywords. Therefore, it is necessary to further fine screen and classify the category primary screening database. The traditional method is to manually label the data first, and then train a supervised text classification model to determine whether the content text is related to a certain category of goods. Due to the large number of product categories that need to be processed, the cost of manually labeling data and then training the classification model is high.

[0086] In step S124, the fine screening classification specifically includes: after matching the primary screening data in the primary screening database with brand and category keywords, if both brand and category entities are matched, it is determined that the text content is related to the product category data; if both brand and category entities are matched, there is no need to use a model to judge, thereby increasing the speed and reducing computing power consumption; for only brand or category keywords, due to the high possibility of ambiguity, a large language model is used to use prompt words such as "Does the above text describe a skin care product?" to build a text classifier, classify the content text into the corresponding category, and then build a content category database to complete.

[0087] S13: Extract information from the content category database to construct a category content label tree, and label the data in the category database according to the category content label tree. This step specifically includes the following steps:

[0088] S131: First, use the RaNER model to extract person entities, category entities, brand entities, and product attribute entities from the category database. Then, use the large language model combined with information extraction prompts and thought chain summary prompts to perform semantic recognition on the category database and extract person entities, network hot word entities, user pain point entities, and product feature entities. Finally, fuse the entity results to obtain the final entity.

[0089] In this step, RaNER extracts category information, brand information, and some product attribute-related information relatively accurately. However, its extraction performance on some semantically rich content texts is average, and its ability to understand the semantics of long content texts is insufficient. For example, the meaning of words such as "desert dryness" and "for skin" is unclear. Therefore, a large language model (LLM model) is used for entity recognition, and multiple prompts are used for extraction. The LLM model can be used for various NLP tasks, including entity recognition. The LLM model is an artificial intelligence model designed to understand and generate human language. They are trained on a large amount of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, etc. However, there are still some problems when using LLM extraction, such as incomplete entity recall. This solution uses multiple types of prompts to perform entity recognition on a piece of content text and fuses the results. The combined use of multiple types of prompts can achieve better label word extraction effects than a single prompt, and further ensure the accuracy of entity extraction, thereby ensuring the accuracy of subsequent content label tree construction; and the entity words extracted from the content text through the RaNER model and LLM model, including the correspondence between content ID and entity word ID, are stored in the word extraction database table. In the subsequent content labeling according to the structured label tree, the result can be reused.

[0090] S132: The final entities (entity words and entity word types, plus the attribute words extracted from the products, such as entity words of the function and efficacy type) are converted into word vectors through the text vectorization model; then, a clustering algorithm is used to obtain several categories of word vectors (word vectors with high similarity are clustered into one category); due to the limitations of the clustering algorithm, after simple word vector clustering, words in the same category may still correspond to different meanings. Therefore, the semantic understanding ability of the large language model is then used to summarize the words in each category into one or more tags, and the keyword types output by the large language model are used to construct a tree-structured content tag tree and keywords for each tag;

[0091] S133: Tag the content text in the category database according to the category content tag tree. Tagging the content text in step S133 includes the following steps:

[0092] S1331: When labeling the content text, determine whether entity extraction has been performed on the content text (i.e., whether extraction has been performed using the RaNER and LLM models). If entity extraction has been performed, the entity words become candidate tags. If entity extraction has not been performed, entity extraction is performed (including RaNER entity recognition and recognition using the LLM large model using the information extraction prompt and hierarchical link summary prompt). The tags corresponding to the recognized entity words are added to the candidate tag set.

[0093] S1332: Keyword matching and regular expression matching are performed on the keywords or regular expressions corresponding to each tag in the tag tree, and the matched tags are also added to the candidate tag set of the content text;

[0094] S1333: The candidate tag set obtained through the first two steps may have potential errors. For the tags identified such as keywords, there may be some semantic errors. Therefore, this step uses the large language model to use the discriminant Prompt to judge all the candidate tags that have been screened out by the content text. The prompt word is such as "Judge whether the content text mentions the following content tags..." to determine whether the candidate tag matches the meaning of the corresponding content text; if it matches, the candidate tag is confirmed, and if it does not match, it is corrected. In this way, while improving the recall rate (reducing omissions) through various methods, the semantic understanding ability of the LLM model is used to filter out entities that are matched by keywords but are semantically incorrect as much as possible, and to filter out entity words that do not exist in the original text that may be given by the LLM model in the previous link, thereby improving both precision and recall.

[0095] To sum up, the specific operation of labeling the content in the content database according to the structured tag tree to obtain content tags in this application is to construct a product knowledge graph, and use the entities and relationships in the product knowledge graph to construct a content category database from the media content database, so that the construction efficiency of the content category database is fast, and work efficiency is effectively improved; after the construction of the content category database is completed, information is extracted from it to construct a content tag tree, and when performing information extraction, the RaNER model is used to identify concrete entities, and then the large language model is used in combination with the information extraction prompt and the thinking chain summary prompt to identify abstract entities. Different types of entities are identified and extracted using different models, which effectively makes up for the problems of inaccurate extraction and recognition and incomplete entity recall caused by entity extraction by a single model; finally, the content text is labeled through the content tag tree; the whole process is simple and not cumbersome, the entity extraction accuracy is high, and the overall labeling accuracy is high; at the same time, it is efficient, reducing labor costs and time costs.

[0096] S2: Calculate the content interaction index of all content tags and the product sales index of all content tags; when calculating the index, for a certain product category, for a given time period (start date Ds, end date De), for the data in the content database and e-commerce product database in step S1, calculate; in a product category, for a content tag k, set X k is the sales index of the content tag k, Y k is the interaction index of the content tag k;

[0097] Specifically, the calculation of the content interaction index of the content tag in step S2 includes the following steps:

[0098] S21: Determine the product category in the content database and count the number n of all content with content tag k in the product category within a set time period (the set time period can be confirmed according to actual conditions, usually a start date and an end date are determined, and the days between the start date and the end date are the set time period), and obtain the content set C k ={content1, content2,…, contentn};

[0099] S22: Calculate the interaction value E of each content in the content set i :

[0100] E i = Number of likes for each content + Number of reposts for each content + Number of favorites for each content + Number of comments for each content; or

[0101] E i = Number of likes for each content, or number of reposts for each content, or number of favorites for each content, or number of reviews for each content; the specific calculation of the interaction value can be selected based on actual circumstances;

[0102] S23: Calculate the content interaction index Y of content tag k k :

[0103] or

[0104] Specifically, as shown in FIG2 , the step S2 of calculating the commodity sales index of the content tag includes the following steps:

[0105] S201: Determine the product category in the content database, and count the number of contents n of all content tags k in the product category within a set time period to obtain a collection C of content k ={content1, content2,…, contentn};

[0106] S202: For each content i in the content set, assuming that content i mentions m commodity SPUs, determine the m commodity SPUs corresponding to the single content i in the content set, and obtain the commodity SPU set P corresponding to the single content i. i ={SPU1, SPU2, ..., SPU m};

[0107] S203: Determine the sales of a single commodity SPU on a certain date d within a set time period; for example: for content i, starting from the release date of content i, set the sales of content i on commodity SPU jThe sales impact time window length is t = 14 days (it can also be adjusted to 30 days, etc., the time window setting here can be determined according to the actual situation). Taking 14 days as an example, during these 14 days, the content i has a significant impact on the product SPU. j The sales contribution values ​​are [S(j,0), S(j,1),…, S(j,13)]; for product SPU j For example, the total sales on a certain date d within the set time window is Sales(j,d), which can be queried in the database through SQL (SPU of date d). j The sum of sales of all corresponding product links);

[0108] S204: Determine the contribution of a single content to the SPU sales of a single product on a certain date d:

[0109] Among them, Sales(j,d) is the product SPU j Sales on date d; W i For a single content i on date d, the product SPU j The contribution weight of sales; u is the number of SPUs mentioned within the effective time window for date d j The total number of contents;

[0110] In step S204, SPU j Sales(j,d) on a certain date d is split into SPU j Content i, thus obtaining the SPU of content i on date d for product j Sales contribution value S i (j,d);

[0111] Assume that on a certain date d, there are u pieces of content {content 1, content 2, ..., content u} that are published within 14 days and all of them mention the product SPU j ; Let content i have SPU for product on date d j The contribution weight of sales is W i , W i You can choose different values, such as W i = the amount of interaction of content i, which can also be replaced by other values ​​(the specific replacement can be determined according to the actual situation); thus determining the SPU of a single content i on the product on date d j sales contribution value.

[0112] In order to further facilitate the understanding of step S204, the content i on date d is the product SPU j Sales contribution value S i(j, d), take the following example: As shown in Figure 3, there are 3 contents in Figure 3, content 1 and content 3 mention the products {SPU1, SPU2}, content 2 mentions the product {SPU1}, assuming that product SPU1 has sales on d9 (the sum of the corresponding product link sales), and SPU2 has sales on a certain day on d10, the sales of each product SPU on each date are split into all the content within the time window of that day according to the weight, and then the sales of all product SPUs assigned to each content are accumulated as the contribution value of each content to the product sales.

[0113] S205: Determine whether a single content i has a product P mentioned in the content i ={SPU1, SPU2, ..., SPU m Total sales contribution:

[0114] Where t is the time window length of the impact of the content release on the product sales calculation. When t is 14, then

[0115] S206: Determine the content set C corresponding to the content tag k k Cumulative value of impact on product sales:

[0116] S207: Calculate the sales index X of the content tag k : or Where p is the content set C corresponding to the content label k k The corresponding total SPU number of duplicated products can be obtained from the database through statistics.

[0117] Here we explain that the sales index X of content tag k k and interaction indicator Y k Log processing is not required, but since large extreme values ​​are likely to appear in big data, log processing can produce better visualization effects and facilitate user visualization analysis.

[0118] S3: Visualize and analyze content interaction and product sales metrics. This includes the following steps:

[0119] S31: Establish a two-dimensional coordinate system and assign the sales index X corresponding to the content label k k As the X-axis; the interaction indicator Y corresponding to the content label k As the Y axis;

[0120] S32: Determine the mean of the sales index and interaction index of all content tags, divide them into four areas with the mean of the sales index and interaction index as the boundary, and place the sales index and interaction index in the four areas respectively. Specifically, after calculating the mean of the sales index and the mean of the interaction index, the two means are used as dividing lines on the X-axis and the Y-axis, and two mutually perpendicular lines intersect to form four areas: the wait-and-see area (lower left), the opportunity area (upper left), the strong area (upper right), and the convinced area (lower right); then put the values ​​of the sales index and interaction index of all content tags into different areas: the strong area means that the sales index and the interaction index are both above their respective means; the opportunity area means that the sales index is below the mean and the interaction index is above the mean; the wait-and-see area means that the sales index and the interaction index are both below their respective means; the convinced area means that the sales index is above the mean and the interaction index is below the mean.

[0121] This application helps customers find content tags that perform better in content interaction, or better in product sales, or both, by calculating quantitative indicators of content tags in terms of content interaction and product sales; quantitative comparison of content tags makes the creation of marketing content simple, feasible and accurate, saving manpower and resource costs.

[0122] The examples described in the present invention are merely descriptions of the preferred embodiments of the present invention and are not intended to limit the concept and scope of the present invention. Without departing from the design concept of the present invention, various modifications and improvements made to the technical solutions of the present invention by engineers and technicians in this field should fall within the scope of protection of the present invention.

Claims

1. A method for quantitatively analyzing the interaction and sales indicators of content tags, characterized by: The following steps are involved: Obtain a content database of the social media platform, label the content in the content database to obtain content labels, and add a product SPU label to the content database; Obtain the e-commerce product database of the e-commerce platform within the social media platform and add product SPU tags to the product data; Calculate content engagement metrics for all content tags and calculate product sales metrics for all content tags; Conduct visual analysis of content interaction indicators and product sales indicators; The calculation of the content interaction index comprises the following steps: Determine the product category in the content database, and count the number of contents n of all content tags k in the product category within the set time period to obtain a collection C of content k ={content1, content2,…, contentn}; Calculate the interaction value E of each content in the content collection i : E i = the number of likes for each content + the number of reposts for each content + the number of favorites for each content + the number of comments for each content; or E i = the number of likes for each content, the number of reposts for each content, the number of collections for each content, or the number of reviews for each content; Calculate the content interaction index Y for content tag k k : or The calculation of the commodity sales index comprises the following steps: Determine the product category in the content database, and count the number of contents n of all content tags k in the product category within the set time period to obtain a collection C of content k ={content1, content2,…, contentn}; Determine the m commodity SPUs corresponding to a single content in the content set, and obtain the commodity SPU set P corresponding to a single content i = {SPU1, SPU2, …, SPU m }; Determine the sales of a single product SPU on each date within a set time period; Determine a date d, a single content i for a single product SPU j Sales contribution: Among them, Sales(j,d) is the product SPU j Sales volume on date d; W i For a single content i on date d, the product SPU j The contribution weight of sales; u is the SPU mentioned within the effective time window for date d j The total number of contents; Determine the relationship between a single content i and the product P mentioned in the content i = {SPU1, SPU2, …, SPU m Total sales contribution: Among them, t is the number of days in the time window from the release date that the calculated content affects the sales volume of the product; Determine the content set C corresponding to the content tag k k Cumulative value of impact on product sales: Calculate the sales index X of content tag k k : or Where p is the content set C corresponding to the content label k k The corresponding total SPU number of duplicated products; The tagging of the content in the content database to obtain the content tag specifically includes the following steps: Build a product knowledge graph; Build a content category database; obtain a media content database, use the product knowledge graph to filter out the content data related to the category of each product from the media content database, and build a content category database; Extract information from the content category database to construct a category content label tree, and label the data in the category database according to the category content label tree.

2. The method for quantitatively analyzing the interaction and sales indicators of content tags according to claim 1, characterized in that: The visual analysis of content interaction indicators and product sales indicators includes the following steps: Establish a two-dimensional coordinate system, with the sales index corresponding to the content tag as the X-axis and the interaction index corresponding to the content tag as the Y-axis; Determine the mean values ​​of the sales index and the interaction index of all content tags, divide them into four areas based on the mean values ​​of the sales index and the interaction index, and put the sales index and the interaction index into the four areas respectively.

3. The method for quantitatively analyzing the interaction and sales indicators of content tags according to claim 1, characterized in that: Extract information from the content category database to build a category content label tree, and label the data in the category database according to the category content label tree; this step specifically includes the following steps: First, the RaNER model is used to extract person entities, category entities, brand entities, and product attribute entities from the content category database. Then, the large language model is used in combination with the information extraction prompt and the thought chain summary prompt to perform semantic recognition on the category database and extract person entities, network hot word entities, user pain point entities, product feature entities, and applicable entities. Finally, the entity results are fused to obtain the final entity. The final entity is converted into a word vector through the text vectorization model; then several categories of word vectors are obtained through the clustering algorithm; then the words in each category are summarized into one or more tags through the large language model, and the keyword types output by the large language model are used to build a tree-structured content tag tree and the keywords of each tag; Label the content text in the category database according to the category content label tree.

4. The method for quantitatively analyzing the interaction and sales indicators of content tags according to claim 3 is characterized by: Tagging content text includes the following steps: When labeling the content text, determine whether the content text has been subjected to entity extraction; if entity extraction has been performed, the entity word becomes a candidate tag; if entity extraction has not been performed, after entity extraction, the tag corresponding to the identified entity word is added to the candidate tag set; And for the keywords or regular expressions corresponding to each tag in the tag tree, keyword matching and regular expression matching are used, and the matched tags are also added to the candidate tag set of the content text; Use the large language model to use the discriminant prompt to judge all the candidate tags that have been screened out for the content text to determine whether the candidate tags match the meaning of the corresponding content text; if they match, confirm the candidate tag, if not, modify it.

5. The method for quantitatively analyzing the interaction and sales indicators of content tags according to claim 3 is characterized by: Building a content category database includes the following steps: Collect content information from various social media to form a media content database; Use the product knowledge graph to match the text information in the media content database and establish a content screening database for product categories; Convert the image type content and video type content in the preliminary screening database into text content respectively; Perform fine screening and classification on the initial screening database to determine whether the text content is related to the product category.

6. The method for quantitatively analyzing the interaction and sales indicators of content tags according to claim 5, characterized in that: The media content database only stores original text description information: for graphic content, the content title, content text and picture link content are stored; for video content, the content title and video link content are stored.

Citation Information

Patent Citations

  • Video-centered convergence media content recommendation method and device

    CN113010701A

  • Data analysis method and system based on big data commodity precision marketing

    CN114971747A

  • Brand delivery effect evaluation method combining social media platform and e-commerce platform data and storage medium

    CN116091099A

  • Analysis method for quantitatively analyzing interaction and sales indexes of content labels

    CN117611243A

  • System and method for aggregating, classifying and enriching social media posts made by monitored author sources

    US20170193075A1

Cited By

  • Content extraction and delivery method based on multi-mode quotient single video

    CN120995408A

  • Marketing decision factor mining method and system fusing multi-modal data

    CN121456838A