Information content auditing management method and device
By extracting and matching metadata features of information content and using databases to predict user feedback behavior, the problem of insufficient utilization of channel data in power marketing has been solved, automated information content review has been achieved, and the accuracy of marketing strategies and user satisfaction have been improved.
Patent Information
- Application Number
- CN202510959963.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-11-21
AI Technical Summary
The effectiveness of electricity marketing and operation services has not been fully improved, channel data cannot be maximized, and there is a lack of comprehensive understanding of customer preferences and channel characteristics, resulting in insufficient optimization of marketing strategies and user satisfaction.
By extracting metadata features from news content, using a news comparison and review database for similarity matching, predicting user feedback behavior characteristics, quantifying the risks and benefits of publishing, and achieving automated review.
Accurately assess the risks and benefits of publishing information content, optimize marketing strategies, improve the effectiveness of electricity marketing and operation services, and enhance user satisfaction.
Smart Images

Figure CN120996728A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power marketing promotion, and in particular to an information content review management method and device. BACKGROUND
[0002] With the continuous operation and promotion of the company's "online State Grid" for many years and the rapid development of "Internet+" marketing business, the company has formed an online service channel matrix mainly based on online State Grid, supplemented by WeChat, Alipay, government affairs, and other services, and has accumulated a large amount of channel behavior data. At the same time, customers' experience requirements and service capabilities for Internet services are continuously improving, and the power company needs to further improve the fine operation and control capabilities of each service channel, continuously strengthen the precision service capabilities of the operation, and further improve the grid promotion management.
[0003] However, at present, there is a lack of comprehensive understanding of customer preferences, channel characteristics, and service effectiveness of the channel, and the channel data cannot maximize its value, resulting in the need for a technology for comprehensive and automatic review and risk effect evaluation of power marketing operation information content to reduce potential risks, optimize marketing strategies, and improve user satisfaction. SUMMARY
[0004] The embodiments of the present application provide an information content review management method and device to solve the problem of improving the service effect of power marketing operation.
[0005] In a first aspect, the embodiments of the present application provide an information content review management method, comprising:
[0006] Extracting metadata features of the information content to be reviewed, wherein the metadata features include promotion channels, titles, themes, keywords, content structures, language styles, hot time opportunities, and hot categories;
[0007] Based on the metadata features, matching in a preset information comparison and review database to obtain estimated user feedback behavior features of the information content to be reviewed, wherein the information comparison and review database contains metadata features and user feedback behavior features of a plurality of historical information contents;
[0008] Based on the estimated user feedback behavior features, evaluating the release risk and release benefit of the information content to be reviewed as the review result of the information content to be reviewed.
[0009] In a possible implementation manner, based on the metadata features, matching in a preset information comparison and review database to obtain estimated user feedback behavior features of the information content to be reviewed, comprising:
[0010] For each metadata category combination, the metadata similarity of the to-be-audited information content and each historical information content in the information comparison auditing database is calculated; wherein, each metadata category combination corresponds to one user feedback category combination, each metadata category combination includes at least one metadata feature category, and each user feedback category combination includes at least one user feedback behavior feature category;
[0011] According to the user feedback behavior feature category, the user feedback behavior features of each historical information content with the metadata similarity greater than the preset threshold are merged to obtain the estimated user feedback behavior features of the to-be-audited information content.
[0012] In a possible implementation, for each metadata category combination, the metadata similarity of the to-be-audited information content and each historical information content in the information comparison auditing database is calculated, including:
[0013] Based on each metadata feature category in the first metadata category combination, the metadata feature of the corresponding category of the to-be-audited information content is normalized and combined into a vector as the first metadata combination vector of the to-be-audited information content; wherein, the first metadata category combination is any metadata category combination;
[0014] Based on each metadata feature category in the first metadata category combination, the metadata feature of the corresponding category of each historical information content is normalized and combined into a vector as the first metadata combination vector of the historical information content;
[0015] The cosine similarity of the first metadata combination vector of each historical information content and the first metadata combination vector of the to-be-audited information content is calculated as the metadata similarity of the historical information content and the to-be-audited information content.
[0016] In a possible implementation, according to the user feedback behavior feature category, the user feedback behavior features of each historical information content with the metadata similarity greater than the preset threshold are merged to obtain the estimated user feedback behavior features of the to-be-audited information content, including:
[0017] For each metadata category combination, the user feedback behavior feature of the corresponding category of the historical information content with the metadata similarity greater than the preset threshold is taken as the candidate user feedback behavior feature of the to-be-audited information content;
[0018] For each user feedback behavior feature category, the values of each candidate user feedback behavior feature of the user feedback behavior feature category are merged to obtain the estimated range of the user feedback behavior feature category of the to-be-audited information content, and the estimated user feedback behavior features of the to-be-audited information content.
[0019] In a possible implementation, before calculating the metadata similarity between the to-be-audited information content and the metadata of each historical information content in the information comparison and auditing database for each metadata category combination, the method further includes the following steps:
[0020] combining each metadata feature category to obtain a plurality of metadata category combinations;
[0021] combining each user feedback behavior feature category to obtain a plurality of user feedback category combinations;
[0022] calculating a correlation index between each metadata category combination and each user feedback category combination based on the metadata features and the user feedback behavior features of the plurality of historical information contents;
[0023] ranking each metadata category combination and each user feedback category combination based on the correlation index, to determine the user feedback category combination corresponding to each metadata category combination.
[0024] In a possible implementation, the step of calculating a correlation index between each metadata category combination and each user feedback category combination based on the metadata features and the user feedback behavior features of the plurality of historical information contents includes the following steps:
[0025] if the first metadata category combination and the first user feedback category combination are both continuous variables, calculating a Pearson correlation coefficient between the first metadata category combination vector and the first user feedback category combination vector of each historical information content, and taking the average of the Pearson correlation coefficients as the correlation index between the first metadata category combination and the first user feedback category combination; wherein the first metadata category combination is any metadata category combination, and the first user feedback category combination is any user feedback category combination;
[0026] if the first metadata category combination and the first user feedback category combination are both ordered variables, calculating a Spearman rank correlation coefficient between the first metadata category combination vector and the first user feedback category combination vector of each historical information content, and taking the average of the Spearman rank correlation coefficients as the correlation index between the first metadata category combination and the first user feedback category combination; wherein the first metadata category combination is any metadata category combination, and the first user feedback category combination is any user feedback category combination.
[0027] In a possible implementation, the formula for calculating the publishing risk is as follows:
[0028] R = w1R negative + w2R complaint + w3R legal + w4R spread
[0029] Wherein, R is the release risk, w1, w2, w3, w4 are preset weights, R negative is the negative emotion proportion, R complaint is the number of reports and complaints, R legal is the legal compliance risk score, R spread is the risk score of uncontrolled spread;
[0030] The calculation formula of the release benefit is:
[0031] B = v1B engagement + v2B brand + v3B conversion
[0032] Wherein, B is the release benefit, v1, v2, v3 are preset weights, B engagement is the user participation score, B brand is the influence promotion score, B conversion is the conversion score.
[0033] In a possible implementation, before the matching based on the metadata features in the preset information comparison audit database to obtain the estimated user feedback behavior features of the to-be-audited information content, the method further includes:
[0034] Obtaining metadata features and user feedback behavior features of a plurality of historical information contents, and constructing an information comparison audit database.
[0035] In a second aspect, the embodiments of the present application provide an information content audit management device, which comprises:
[0036] An extraction module is configured to extract metadata features of to-be-audited information content, wherein the metadata features include a promotion channel, a title, a theme, a keyword, a content structure, a language style, a hot time and a hot category.
[0037] A matching module is configured to match based on the metadata features in a preset information comparison audit database to obtain estimated user feedback behavior features of the to-be-audited information content, wherein the information comparison audit database contains metadata features and user feedback behavior features of a plurality of historical information contents.
[0038] An evaluation module is configured to evaluate the release risk and the release benefit of the to-be-audited information content based on the estimated user feedback behavior features, as an audit result of the to-be-audited information content.
[0039] The embodiment of the present application provides an information content review management method and device, core attributes of information are reflected by extracting metadata features of information content, historical data in an information comparison review database is utilized, similarity matching is carried out based on the metadata features, user feedback behavior features of to-be-reviewed information are predicted, release risk and release benefit of the to-be-reviewed information are quantitatively evaluated accurately, automatic review is realized, and power marketing operation service effect is improved. BRIEF DESCRIPTION OF DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0041] Figure 1 is an implementation flowchart of the information content review management method provided by the embodiment of the present application;
[0042] Figure 2 is a structural schematic diagram of the information content review management device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0043] In the following description, specific details such as specific system structures, techniques, etc. are presented in order to thoroughly understand the embodiments of the present application. However, it should be clear to those skilled in the art that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits and methods are omitted to avoid unnecessary details that hinder the description of the present application.
[0044] In order to make the objects, technical solutions and advantages of the present application clearer, the following will be described by specific embodiments in conjunction with the drawings.
[0045] Figure 1 The implementation flowchart of the information content review management method provided by the embodiment of the present application is described in detail as follows:
[0046] Step 101, metadata features of to-be-reviewed information content are extracted; wherein, the metadata features include promotion channels, titles, themes, keywords, content structures, language styles, hot time opportunities and hot categories.
[0047] In the present embodiment, metadata refers to attribute information describing the content itself, and features can be extracted from the following aspects:
[0048] 1. Text features
[0049] Key / Theme: Extract the core keywords or themes of an article through Natural Language Processing (NLP) techniques.
[0050] Sensitive Word Detection: Identify the presence of rule-breaking words.
[0051] Sentiment Analysis: Determine the emotional orientation of the text (positive, negative, neutral) to identify potential risks.
[0052] Semantic Similarity: Compare the semantics with known risky content to determine if there are similar expressions.
[0053] 2. Author Features
[0054] Author Identity: Distinguish between official media, individual users, or anonymous sources.
[0055] Historical Record: Analyze the quality and risk level of the author's past published content.
[0056] Authentication Status: Whether there is authoritative authentication (such as Blue V authentication).
[0057] 3. Time Features
[0058] Publication Time: Determine whether the content is related to a hot event.
[0059] Timeliness: Whether the content is outdated or lagging, which may lead to misinformation.
[0060] 4. Source Features
[0061] Platform Source: The platform where the content comes from (such as Weibo, WeChat public account, etc. self-media, news website, radio and television, shopping platform, self-service app).
[0062] Geographical Information: The geographical area where the content is published, which may involve local policies or cultural differences.
[0063] 5. Transmission Features
[0064] Transmission Path: The forwarding link of the content, to determine whether it is a rumor spreading node.
[0065] Interaction Data: Comments, likes, and shares, reflecting the social impact of the content.
[0066] 6. Multimedia Features (if there are pictures or videos)
[0067] Image Recognition: Detect sensitive content (such as gore) in images.
[0068] OCR Text Extraction: Extract text from images and conduct audits.
[0069] Audio / Video Transcription: Convert speech content into text for analysis.
[0070] Step 102, based on the metadata features, match in the preset information comparison audit database to obtain the estimated user feedback behavior characteristics of the information content to be audited; wherein the information comparison audit database contains the metadata features and user feedback behavior characteristics of a plurality of historical information contents.
[0071] In this embodiment, the metadata features of the information affect user feedback from multiple dimensions, including the quality of the content itself, the credibility of the author, the reasonableness of the release time, and the design of the transmission path, etc. Through comprehensive analysis of these features, the behavior of the user can be better predicted, and the information release strategy can be optimized, so as to maximize the benefits and reduce the potential risks.
[0072] The user feedback behavior characteristics can include user reading, interaction, retransmission, conversion, emotion, etc. Specifically, it can include:
[0073] 1. Reading related data
[0074] Reading volume: Statistics of how many users opened the information content.
[0075] Stay time: The average time a user spends on a page, reflecting the attractiveness of the content.
[0076] Scroll depth: Whether the user has read the entire article.
[0077] Exit rate: The proportion of users who leave immediately after opening the page.
[0078] 2. Interaction related data
[0079] Like / favorite: The degree of user approval of the content.
[0080] Comment quantity and content: User feedback on the content, including positive and negative evaluations.
[0081] Sharing times: The frequency of users forwarding the content to social media or other platforms.
[0082] Report and complaint times: User complaints or questions about the content.
[0083] 3. Transmission related data
[0084] Transmission path: The transmission link of the information after being shared (such as from user A to user B).
[0085] Coverage: How many users were affected by the information, and which areas or groups were covered.
[0086] Secondary creation: Whether users have generated new content (such as articles, videos, etc.) based on the information content.
[0087] 4. Conversion related data
[0088] Click on links: Whether users click on external links in the information (such as policy interpretation, activity registration, etc.).
[0089] Form filling: Whether users participate in surveys or registration activities related to the information.
[0090] Purchase behavior (if there is an e-commerce association): Whether users have generated consumption behavior due to the information.
[0091] 5. Emotion-related data
[0092] Sentiment tendency: Analyze the sentiment (positive, negative, neutral) of user comments through natural language processing technology.
[0093] Keyword extraction: Identify high-frequency words in comments to determine user focus.
[0094] 6. Device and environment data
[0095] Access device: Whether the user uses a PC, mobile phone or tablet.
[0096] Access time: Time period when users read the information.
[0097] Geographical location: User's location, helping to analyze regional preferences.
[0098] Step 103, based on the estimated user feedback behavior characteristics, evaluate the release risk and release benefit of the to-be-audited information content as the audit result of the to-be-audited information content.
[0099] In this embodiment, the specific way to evaluate the release risk and release benefit of the to-be-audited information content can include:
[0100] 1. Risk assessment
[0101] Negative sentiment proportion: Calculate the proportion of negative comments in total comments, the higher the proportion, the greater the risk.
[0102] Report and complaint volume: If the number of reports exceeds a certain threshold, it needs to be paid special attention to.
[0103] Risk of uncontrolled spread: Analyze whether the transmission path exists malicious diffusion or misleading transmission.
[0104] Sensitive topic trigger: Combine public opinion monitoring tools to determine whether the information involves sensitive areas.
[0105] 2. Benefit assessment
[0106] User engagement: Comprehensive reading volume, number of likes, number of comments, etc. Measure the popularity of the information.
[0107] Conversion Effect: Calculate the actual behavior of users due to the information (such as clicking links, completing registration).
[0108] Brand Impact: Assess the impact of information on brand image through spread and coverage of the population.
[0109] Long-term Value: Analyze user retention rate and revisit rate to determine whether the information has sustained influence.
[0110] 3. Comprehensive Score
[0111] Risk Index: Calculate risk score based on negative sentiment proportion, number of reports, etc.
[0112] Profit Index: Calculate profit score based on user engagement, conversion effect, etc.
[0113] Balance Analysis: Compare risk index and profit index to draw final conclusion:
[0114] If the profit is much higher than the risk, continue to promote similar content.
[0115] If the risk is high, adjust the content strategy or strengthen the review.
[0116] For example, for information content with content category of power policy, if users misunderstand the policy interpretation, it may cause social public opinion pressure, and clear interpretation can help improve public understanding and support for the policy.
[0117] For information content with content category of power outage notice, if the notice is not timely or the wording is inappropriate, it may cause user dissatisfaction, and timely and accurate notice can reduce user complaints and improve service quality.
[0118] For information content with content category of energy-saving propaganda, if the content is too abstract or lacks practicality, it may reduce user interest, and effective propaganda can guide users to change their behavior habits and promote energy saving and emission reduction.
[0119] The embodiment of the present application reflects the core attributes of information by extracting the metadata features of information content, and uses the information comparison and review database to match the metadata features based on the historical data, to predict the user feedback behavior characteristics of the to-be-reviewed information, thereby accurately quantifying the release risk and release benefit of the to-be-reviewed information, realizing automatic review, and improving the effect of power marketing operation service.
[0120] In one possible implementation, based on the metadata features, the to-be-reviewed information content is matched in the preset information comparison and review database to obtain the estimated user feedback behavior characteristics of the to-be-reviewed information content, including:
[0121] For each metadata category combination, the metadata similarity of the to-be-audited information content and each historical information content in the information comparison audit database is calculated; wherein each metadata category combination corresponds to a user feedback category combination, each metadata category combination includes at least one metadata feature category, and each user feedback category combination includes at least one user feedback behavior feature category;
[0122] According to the user feedback behavior feature category, the user feedback behavior features of each historical information content with a metadata similarity greater than a preset threshold are merged to obtain the estimated user feedback behavior features of the to-be-audited information content.
[0123] In this embodiment, the metadata feature category is a single dimension attribute for describing information content, such as title, theme, keyword, language style, etc. The metadata category combination is a set composed of at least one metadata feature category, which is used to describe different dimensions of information content. For example, title + theme + keyword constitutes a metadata category combination.
[0124] The user feedback behavior feature category is a specific index category describing the user's reaction to the information content. For example, reading volume, like number, negative comment ratio, etc. The user feedback category combination is a set composed of at least one user feedback behavior feature category, which is used to comprehensively describe the multi-dimensional feedback mode of the user. For example, "reading volume + like number" or "negative comment ratio + report frequency".
[0125] The metadata similarity is used to measure the similarity of the to-be-audited information content and the historical information content in a certain metadata category combination. Based on the historical information content with a metadata similarity greater than a preset threshold, its corresponding user feedback behavior features are extracted as candidate features. These candidate features reflect the actual user feedback of the historical information similar to the to-be-audited information content.
[0126] For each user feedback behavior feature category, the values of the candidate features are merged to generate an estimated range. Common methods include:
[0127] Statistical summary: calculate the mean, median or standard deviation.
[0128] Frequency statistics: analyze high-frequency feature values.
[0129] Weighted average: calculate the weighted average value according to the similarity weight.
[0130] In this embodiment, the candidate user feedback behavior features are screened, and the values are merged according to the user feedback behavior feature category to generate the estimated user feedback behavior features of the to-be-audited information content. This method can effectively predict the user reaction after the information is published, and provide a scientific basis for automatic audit and decision-making in the power marketing scenario.
[0131] In a possible implementation, for each metadata category combination, the metadata similarity between the to-be-audited information content and the metadata of each historical information content in the information comparison and auditing database is calculated, including:
[0132] Based on each metadata feature category in the first metadata category combination, the metadata features of the corresponding category of the to-be-audited information content are normalized and combined into a vector as the first metadata combination vector of the to-be-audited information content; wherein the first metadata category combination is any metadata category combination;
[0133] Based on each metadata feature category in the first metadata category combination, the metadata features of the corresponding category of each historical information content are normalized and combined into a vector as the first metadata combination vector of the historical information content;
[0134] The cosine similarity between the first metadata combination vector of each historical information content and the first metadata combination vector of the to-be-audited information content is calculated as the metadata similarity between the historical information content and the to-be-audited information content.
[0135] In the embodiment, different metadata features can have different dimensions or value ranges (for example, the title length is an integer and the keyword heat is a decimal number). Through normalization processing, the dimension difference can be eliminated to ensure that the contribution weights of each feature to the similarity calculation are consistent. The normalized metadata features are combined into a vector according to the category to form a unified feature representation form, which is convenient for subsequent similarity calculation. The cosine similarity reflects the consistency of the direction of two vectors by calculating the cosine value of the included angle between the two vectors. The smaller the included angle, the higher the similarity.
[0136] In the embodiment, the similarity between the to-be-audited information content and the historical information content is evaluated by defining the first metadata category combination, extracting and normalizing the metadata features, constructing the metadata combination vector, and calculating the cosine similarity. This method can effectively quantify the similarity between information contents and provide a scientific basis for automatic auditing and user feedback prediction in the power marketing scenario.
[0137] In a possible implementation, according to the user feedback behavior feature category, the user feedback behavior features of each historical information content with a metadata similarity greater than a preset threshold are merged to obtain the estimated user feedback behavior features of the to-be-audited information content, including:
[0138] For each metadata category combination, the user feedback behavior features of the historical information content in the corresponding category with a metadata similarity greater than a preset threshold are used as the candidate user feedback behavior features of the to-be-audited information content;
[0139] For each user feedback behavior feature category, the values of the candidate user feedback behavior features of the user feedback behavior feature category are combined to obtain a predicted range of the user feedback behavior feature category of the information content to be audited, and the predicted range is used as the predicted user feedback behavior feature of the information content to be audited.
[0140] In this embodiment, based on the historical information content with a metadata similarity greater than a preset threshold, the corresponding user feedback behavior features are extracted as candidate features. These candidate features reflect the actual user feedback of the historical information similar to the information content to be audited.
[0141] For each user feedback behavior feature category, the values of the candidate features are combined to generate a predicted range. Common methods include:
[0142] Statistical summary: calculate the mean, median or standard deviation.
[0143] Frequency statistics: analyze high-frequency feature values.
[0144] Weighted average: calculate the weighted average value according to the similarity weight.
[0145] The predicted ranges of each user feedback behavior feature category are integrated to form a complete predicted user feedback behavior feature, which is used to evaluate the publishing effect of the information content to be audited.
[0146] In this embodiment, candidate user feedback behavior features are screened, and the values are combined according to the user feedback behavior feature categories to generate the predicted user feedback behavior features of the information content to be audited. This method can effectively predict the user reaction after the information is published, and provide a scientific basis for automatic auditing and decision-making in the power marketing scenario.
[0147] In one possible implementation, before calculating the metadata similarity between the information content to be audited and each historical information content in the information comparison and auditing database for each metadata category combination, the following steps are further included:
[0148] Combine each metadata feature category to obtain multiple metadata category combinations;
[0149] Combine each user feedback behavior feature category to obtain multiple user feedback category combinations;
[0150] Based on the metadata features and user feedback behavior features of the multiple historical information contents, calculate the correlation index between each metadata category combination and each user feedback category combination;
[0151] Sort each metadata category combination and each user feedback category combination based on the correlation index to determine the user feedback category combination corresponding to each metadata category combination.
[0152] In this embodiment, all possible metadata feature categories are extracted from the information content, such as title, theme, keyword, language style, hot opportunity, etc. Then the metadata feature categories are combined to generate multiple metadata category combinations. For example:
[0153] Combination 1: title + theme
[0154] Combination 2: keyword + content structure
[0155] Combination 3: language style + hot opportunity
[0156] From the user feedback behavior, all possible user feedback behavior feature categories are extracted, such as reading volume, like number, negative comment proportion, and report number. Then the user feedback behavior feature categories are combined to generate multiple user feedback category combinations. For example:
[0157] Combination 1: reading volume + sharing number
[0158] Combination 2: like number + comment number
[0159] Combination 3: negative comment proportion + report number
[0160] The correlation index is used to measure the degree of association between a certain metadata category combination and a certain user feedback category combination, which is usually calculated by statistical methods such as Pearson correlation coefficient and information gain. Evaluating the correlation between various metadata features of information and user feedback behavior can help understand which metadata features have the greatest impact on user reaction and optimize content strategy.
[0161] According to the correlation analysis results, the metadata features that have the greatest impact on user feedback can be found. For example, if "title attractiveness" is highly correlated with "reading volume", it can be set as the corresponding metadata category combination and user feedback category combination, and the title attractiveness of the information content to be reviewed is matched in the information comparison and review database to determine the estimated reading volume, a user feedback behavior feature of the information content to be reviewed.
[0162] Specifically, assuming that the correlation between "title attractiveness" and "reading volume" is to be analyzed, the following steps can be included:
[0163] 1. Extract "title length" and "keyword heat" in the data set as proxy variables for title attractiveness.
[0164] 2. Calculate the Pearson correlation coefficient of these variables and "reading volume".
[0165] 3. The results show that the correlation coefficient between "title length" and "reading volume" is 0.65, and the correlation coefficient between "keyword heat" and "reading volume" is 0.78.
[0166] 4. Conclusion: Including popular keywords in the title can improve the reading volume more than simply controlling the title length.
[0167] The single metadata feature category and the user feedback behavior feature category are combined to form multiple metadata category combinations and user feedback category combinations. This combination method can more comprehensively reflect the multi-dimensional characteristics of the information content and the multi-dimensional feedback mode of the user. According to the correlation index, the metadata category combinations and the user feedback category combinations are sorted, and the user feedback category combination with the highest correlation is selected as the corresponding combination of each metadata category combination. This matching method can ensure the accuracy of subsequent similarity calculation and user feedback prediction.
[0168] In one possible implementation, based on the metadata features and user feedback behavior features of multiple historical information contents, the correlation index between each metadata category combination and each user feedback category combination is calculated, including:
[0169] If the first metadata category combination and the first user feedback category combination are both continuous variables, the Pearson correlation coefficient between the first metadata category combination vector and the first user feedback category combination vector of each historical information content is calculated, and the average of each Pearson correlation coefficient is taken as the correlation index between the first metadata category combination and the first user feedback category combination; wherein the first metadata category combination is any metadata category combination, and the first user feedback category combination is any user feedback category combination.
[0170] If the first metadata category combination and the first user feedback category combination are both ordered variables, the Spearman rank correlation coefficient between the first metadata category combination vector and the first user feedback category combination vector of each historical information content is calculated, and the average of each Spearman rank correlation coefficient is taken as the correlation index between the first metadata category combination and the first user feedback category combination; wherein the first metadata category combination is any metadata category combination, and the first user feedback category combination is any user feedback category combination.
[0171] In this embodiment, continuous variables refer to variables that can take any real number value, such as reading volume (e.g. 500, 750) and number of likes (e.g. 20, 30). Ordered variables refer to variables that have a clear order relationship but the values between them are not necessarily equidistant, such as language style score (e.g. formal, neutral, casual) or user satisfaction level (e.g. very satisfied, satisfied, not satisfied). Pearson correlation coefficient is suitable for continuous variables (e.g. reading volume and publishing time). Spearman rank correlation coefficient is suitable for non-linear relationships or ordered variables (e.g. number of likes and author authority). The correlation is measured by calculating the difference in rank between the two variables.
[0172] In this embodiment, by judging the variable type (continuous type or ordered type), Pearson correlation coefficient or Spearman rank correlation coefficient calculation method is used respectively to realize the correlation evaluation between the metadata category combination and the user feedback category combination. This method can effectively quantify the degree of association between the two, providing a scientific basis for automated review and user feedback prediction in the power marketing scenario.
[0173] In a possible implementation, the calculation formula of the release risk is:
[0174] R = w1R negative + w2R complaint + w3R legal + w4R spread
[0175] Wherein, R is the release risk, w1, w2, w3, w4 are preset weights, R negative is the negative emotion proportion, R comolaint is the number of reports and complaints, R legal is the legal compliance risk score, and R spread is the risk score of out-of-control spread.
[0176] The calculation formula of the release benefit is:
[0177] B = v1B engagement + v2B brand + v3B conversion
[0178] Wherein, B is the release benefit, v1, v2, v3 are preset weights, B engagement is the user engagement score, B brand is the influence promotion score, and B conversion is the conversion score.
[0179] In this embodiment, w1, w2, w3, w4 are the weights of each risk factor, which can be adjusted according to business needs. The calculation formula of the negative emotion proportion is:
[0180] The legal compliance risk score can be calculated based on the negative emotion proportion, the number of reports and complaints, and the public opinion spread speed. The risk score of out-of-control spread can be calculated based on the spread path and coverage range through the gradient boosting decision tree algorithm.
[0181] v1, v2, v3 are the weights of each benefit factor, which are set according to business needs. The calculation formula of the user engagement score is: The influence promotion score can be calculated based on the covered population and the positive emotion proportion, and the conversion score can be the proportion of clicking links or completing purchases.
[0182] The number of negative comments, the total number of comments, the number of likes, the number of shares, the number of comments, the covered audience, and the proportion of positive sentiment used in the above calculation process are all estimated user feedback behavior characteristics obtained by matching the metadata characteristics of the information content to be audited in the preset information comparison and audit database.
[0183] Suppose a power company needs to audit an information about power outage notice, the estimated user feedback behavior characteristics are as follows:
[0184] The estimated negative comment proportion is Rnegative=0.15
[0185] The estimated number of complaints is Rcomplaint=3 (the normalized score is 0.2)
[0186] The legal compliance risk score is Rlegal=0.1
[0187] The risk of uncontrolled spread is Rspread=0.2
[0188] The risk weight is set as w1=0.4, w2=0.3, w3=0.2, w4=0.1, then:
[0189] R=0.4·0.15+0.3·0.2+0.2·0.1+0.1·0.2=0.18
[0190] The revenue prediction results are as follows:
[0191] The user engagement score is Bengagement=0.3
[0192] The brand influence promotion score is Bbrand=0.4
[0193] The business conversion score is Bconversion=0.2
[0194] The revenue weight is set as v1=0.5, v2=0.3, v3=0.2, then: B=0.5·0.3+0.3·0.4+0.2·0.2=0.31 To comprehensively evaluate the risk and revenue of publishing, the following formula can be used: S=α·(1-R)+β·B
[0195] Where: S: comprehensive score, used to measure the overall effect of information publishing. α and β: risk and revenue weight coefficients, reflecting the enterprise's preference for risk tolerance and revenue expectations. 1-R: convert the risk value to a positive indicator, indicating that the lower the risk, the higher the score.
[0196] Set α=0.6, β=0.4, then the comprehensive score is: S=0.6·(1-0.18)+0.4·0.31=0.67
[0197] If the threshold is set to S > 0.6, the information can be published.
[0198] In one possible implementation, before obtaining the estimated user feedback behavior features of the information content to be audited by matching the metadata features in the preset information comparison audit database based on the metadata features, the method further includes:
[0199] Obtaining metadata features and user feedback behavior features of a plurality of historical information contents, and constructing an information comparison audit database.
[0200] In this embodiment, by collecting and organizing the metadata features and user feedback behavior features of a plurality of historical information contents, a structured database is established, which can provide data support for automated auditing. The specific implementation steps can include:
[0201] 1. Determine the data source
[0202] Determine the data source of the historical information content, such as historical publication records within the enterprise, user interaction data on social media platforms, etc.
[0203] 2. Extract metadata features
[0204] For each historical information content, extract its metadata features. For example:
[0205] Title: Extract the title text and calculate the length, keywords, etc.
[0206] Topic: Identify the topic label through a text classification algorithm.
[0207] Keywords: Extract the core keywords in the article and their weights.
[0208] Language style: Analyze the degree of formality or emotional inclination.
[0209] Hot time: Determine whether the release time meets the time window of the current hot event.
[0210] Extract user feedback behavior features
[0211] 3. For each historical information content, extract its user feedback behavior features. For example:
[0212] Reading volume: Count the actual number of readers.
[0213] Number of likes: Count the number of likes.
[0214] Negative comment ratio: Calculate the proportion of negative comments in total comments.
[0215] Number of reports: Count the number of user reports.
[0216] 4. Then design the database table structure according to the requirements. For example:
[0217] Table 1: historical information content table, including fields: information ID, title, theme, keyword, language style, hot time opportunity, etc.
[0218] Table 2: user feedback behavior feature table, including fields: information ID, reading volume, like number, negative comment ratio, report number, etc.
[0219] 5. Finally, in order to improve the query efficiency, indexes are established on the key fields. For example:
[0220] A primary key index is established on the 'information ID' field.
[0221] Secondary indexes are established on the commonly used query fields (such as title, theme, keyword).
[0222] As can be seen from the above, the present application first collects information content on the network, and extracts corresponding user feedback, such as user comment information and user behavior information after reading the information content (including online operations such as purchasing related products and further consulting related materials). Then, metadata features of each information content are extracted, including information such as publishing channel, information category, promotion mode, language style, and an information comparison and review database is constructed. Finally, the information content to be reviewed is subjected to text vectorization processing and metadata feature extraction, and is matched with the content in the information comparison and review database. The user feedback of the matched information content is used to estimate the risk and effect after the publication of the information content to be reviewed, which is taken as the review result of the information content. The review personnel can review according to this automatic review result, thereby improving the publication effect of the operation information.
[0223] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.
[0224] The following is a device embodiment of the present application. For details not described in detail, reference can be made to the corresponding method embodiments described above.
[0225] Figure 2 The structure schematic diagram of the information content review management device provided by the embodiment of the present application is shown. In order to facilitate the description, only the parts related to the embodiment of the present application are shown, and the details are as follows:
[0226] As shown in Figure 2 The information content review management device 2 comprises:
[0227] The extraction module 21 is configured to extract metadata features of the information content to be reviewed, wherein the metadata features comprise a promotion channel, a title, a theme, a keyword, a content structure, a language style, a hot time opportunity and a hot category.
[0228] The matching module 22 is configured to match the metadata features in the preset information comparison and review database to obtain estimated user feedback behavior features of the information content to be reviewed; the information comparison and review database includes metadata features and user feedback behavior features of a plurality of historical information contents;
[0229] The evaluation module 23 is configured to evaluate the release risk and release benefit of the information content to be reviewed based on the estimated user feedback behavior features as the review result of the information content to be reviewed.
[0230] In a possible implementation, the matching module 22 is specifically configured to:
[0231] For each metadata category combination, the metadata similarity between the information content to be reviewed and each historical information content in the information comparison and review database is calculated; each metadata category combination corresponds to a user feedback category combination, each metadata category combination includes at least one metadata feature category, and each user feedback category combination includes at least one user feedback behavior feature category;
[0232] According to the user feedback behavior feature category, the user feedback behavior features of each historical information content with a metadata similarity greater than a preset threshold are merged to obtain the estimated user feedback behavior features of the information content to be reviewed.
[0233] In a possible implementation, the matching module 22 is specifically configured to:
[0234] Based on each metadata feature category in the first metadata category combination, the metadata features of the corresponding category of the information content to be reviewed are normalized and combined into a vector as a first metadata combination vector of the information content to be reviewed; the first metadata category combination is any metadata category combination;
[0235] Based on each metadata feature category in the first metadata category combination, the metadata features of the corresponding category of each historical information content are normalized and combined into a vector as a first metadata combination vector of the historical information content;
[0236] The cosine similarity between the first metadata combination vector of each historical information content and the first metadata combination vector of the information content to be reviewed is calculated as the metadata similarity between the historical information content and the information content to be reviewed.
[0237] In a possible implementation, the matching module 22 is specifically configured to:
[0238] For each metadata category combination, the user feedback behavior features of the corresponding category of the historical information content with a metadata similarity greater than a preset threshold are taken as candidate user feedback behavior features of the information content to be audited;
[0239] For each user feedback behavior feature category, the values of the candidate user feedback behavior features of the user feedback behavior feature category are combined to obtain a predicted range of the user feedback behavior feature category of the information content to be audited, and the predicted range is taken as a predicted user feedback behavior feature of the information content to be audited.
[0240] In a possible implementation, the matching module 22 is further configured to:
[0241] Before calculating the metadata similarity between the information content to be audited and each historical information content in the information comparison and auditing database for each metadata category combination, each metadata feature category is combined to obtain a plurality of metadata category combinations;
[0242] Each user feedback behavior feature category is combined to obtain a plurality of user feedback category combinations;
[0243] Based on the metadata features and the user feedback behavior features of the plurality of historical information contents, a correlation index between each metadata category combination and each user feedback category combination is calculated;
[0244] Each metadata category combination and each user feedback category combination are sorted based on the correlation index, and a user feedback category combination corresponding to each metadata category combination is determined.
[0245] In a possible implementation, the matching module 22 is specifically configured to:
[0246] If the first metadata category combination and the first user feedback category combination are both continuous variables, a Pearson correlation coefficient between the first metadata category combination vector and the first user feedback category combination vector of each historical information content is calculated, and an average value of each Pearson correlation coefficient is taken as a correlation index between the first metadata category combination and the first user feedback category combination; wherein the first metadata category combination is any metadata category combination, and the first user feedback category combination is any user feedback category combination;
[0247] If the first metadata category combination and the first user feedback category combination are both ordered variables, a Spearman rank correlation coefficient between the first metadata category combination vector and the first user feedback category combination vector of each historical information content is calculated, and an average value of each Spearman rank correlation coefficient is taken as a correlation index between the first metadata category combination and the first user feedback category combination; wherein the first metadata category combination is any metadata category combination, and the first user feedback category combination is any user feedback category combination.
[0248] In a possible implementation, a calculation formula of the publishing risk is:
[0249] R = w1R negative + w2R complaint + w3R legal + w4R spread
[0250] wherein R is the publishing risk, w1, w2, w3, and w4 are preset weights, R negative is the negative emotion proportion, R complaint is the number of reports and complaints, R legal is the legal compliance risk score, and R spread is the risk score of out-of-control spread.
[0251] A calculation formula of the publishing benefit is:
[0252] B = v1B engagement + v2B brand + v3B conversion
[0253] wherein B is the publishing benefit, v1, v2, and v3 are preset weights, B engagement is the user engagement score, B brand is the influence promotion score, and B conversion is the conversion score.
[0254] In a possible implementation, the matching module 22 is further configured to:
[0255] Before performing matching based on the metadata features in the preset information comparison and review database to obtain the estimated user feedback behavior features of the to-be-reviewed information content, metadata features and user feedback behavior features of a plurality of historical information contents are acquired to construct the information comparison and review database.
[0256] According to the embodiments of the present application, the metadata features of the information content are extracted to reflect the core attributes of the information, and the historical data in the information comparison and review database is used to perform similarity matching based on the metadata features to predict the user feedback behavior features of the to-be-reviewed information, so as to accurately quantitatively evaluate the publishing risk and the publishing benefit of the to-be-reviewed information, to realize automatic review, and to improve the effect of power marketing operation service.
[0257] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.
[0258] Those skilled in the art will recognize that the templates, units, and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0259] If the module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above-described information content review and management method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0260] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An information content review management method, characterized by, The application comprises the following steps: extracting metadata features of the information content to be audited; wherein the metadata features include promotion channels, titles, themes, keywords, content structures, language styles, hot time opportunities, and hot categories; matching the metadata features in a preset information comparison and audit database to obtain estimated user feedback behavior features of the information content to be audited; wherein the information comparison and audit database contains metadata features and user feedback behavior features of a plurality of historical information contents; evaluating the release risk and release benefit of the information content to be audited based on the estimated user feedback behavior features as the audit result of the information content to be audited.
2. The information content review management method of claim 1, wherein The matching of the metadata features in the preset information comparison and audit database to obtain the estimated user feedback behavior features of the information content to be audited comprises the following steps: for each metadata category combination, calculating the metadata similarity between the information content to be audited and each historical information content in the information comparison and audit database; wherein each metadata category combination corresponds to a user feedback category combination, each metadata category combination includes at least one metadata feature category, and each user feedback category combination includes at least one user feedback behavior feature category; combining the user feedback behavior features of each historical information content with a metadata similarity greater than a preset threshold according to the user feedback behavior feature category to obtain the estimated user feedback behavior features of the information content to be audited.
3. The information content review management method of claim 2, wherein The calculation of the metadata similarity between the information content to be audited and each historical information content in the information comparison and audit database for each metadata category combination comprises the following steps: based on each metadata feature category in the first metadata category combination, normalizing and combining the metadata features of the corresponding category of the information content to be audited into a vector as the first metadata combination vector of the information content to be audited; wherein the first metadata category combination is any metadata category combination; based on each metadata feature category in the first metadata category combination, normalizing and combining the metadata features of the corresponding category of each historical information content into a vector as the first metadata combination vector of the historical information content; calculating the cosine similarity between the first metadata combination vector of each historical information content and the first metadata combination vector of the information content to be audited as the metadata similarity between the historical information content and the information content to be audited.
4. The information content review management method of claim 2, wherein The combination of the user feedback behavior features of each historical information content with a metadata similarity greater than a preset threshold according to the user feedback behavior feature category to obtain the estimated user feedback behavior features of the information content to be audited comprises the following steps: for each metadata category combination, the user feedback behavior features of the historical information content with a metadata similarity greater than a preset threshold in the corresponding category are used as the candidate user feedback behavior features of the information content to be audited; merge the values of the candidate user feedback behavior features of the user feedback behavior feature category to obtain a predicted range of the user feedback behavior feature category of the to-be-audited information content, and use the predicted range as the predicted user feedback behavior feature of the to-be-audited information content.
5. The information content review management method of claim 2, wherein, Before the calculating the metadata similarity between the to-be-audited information content and each historical information content in the information comparison and audit database according to each metadata category combination, the method further comprises: combining each metadata feature category to obtain a plurality of metadata category combinations; combining each user feedback behavior feature category to obtain a plurality of user feedback category combinations; calculating a correlation index between each metadata category combination and each user feedback category combination based on the metadata features and the user feedback behavior features of the plurality of historical information contents; ranking each metadata category combination and each user feedback category combination based on the correlation index to determine the user feedback category combination corresponding to each metadata category combination.
6. The information content review management method of claim 5, wherein, The calculating a correlation index between each metadata category combination and each user feedback category combination based on the metadata features and the user feedback behavior features of the plurality of historical information contents comprises: if the first metadata category combination and the first user feedback category combination are both continuous variables, calculating a Pearson correlation coefficient between the first metadata category combination vector and the first user feedback category combination vector of each historical information content, and using the average of the Pearson correlation coefficients as the correlation index between the first metadata category combination and the first user feedback category combination; wherein the first metadata category combination is any metadata category combination, and the first user feedback category combination is any user feedback category combination; if the first metadata category combination and the first user feedback category combination are both ordered variables, calculating a Spearman rank correlation coefficient between the first metadata category combination vector and the first user feedback category combination vector of each historical information content, and using the average of the Spearman rank correlation coefficients as the correlation index between the first metadata category combination and the first user feedback category combination; wherein the first metadata category combination is any metadata category combination, and the first user feedback category combination is any user feedback category combination.
7. The information content review management method of claim 1, wherein The calculation formula of the publishing risk is: R = w1R negative + w2R complaint + w3R legal + w4R spread Wherein, R is the release risk, w1, w2, w3, w4 are preset weights, R negative is the negative emotion proportion, R complaint is the number of reports and complaints, R legal is the legal compliance risk score, R spread is the risk score of uncontrolled spread; The calculation formula of the publishing benefit is: B = v1B engagement + v2B brand + v3B conversion Wherein, B is the release income, v1, v2, v3 are preset weights, B engagement is the user participation score, B brand is the influence promotion score, B conversion is the conversion score.
8. The information content review management method of claim 1, wherein, Before the obtaining the predicted user feedback behavior feature of the to-be-audited information content based on the matching of the metadata features in the preset information comparison and audit database, the method further comprises: obtaining the metadata features and the user feedback behavior features of a plurality of historical information contents to construct the information comparison and audit database.
9. An information content review management apparatus characterized by comprising: comprises: an extraction module configured to extract metadata features of a to-be-audited information content; wherein the metadata features comprise a promotion channel, a title, a theme, a keyword, a content structure, a language style, a hot time, and a hot category; The matching module is configured to match, based on the metadata features, in a preset information comparison and review database to obtain estimated user feedback behavior features of the to-be-reviewed information content; the information comparison and review database includes metadata features and user feedback behavior features of a plurality of historical information contents; The evaluation module is configured to evaluate, based on the estimated user feedback behavior features, a release risk and a release benefit of the to-be-reviewed information content as a review result of the to-be-reviewed information content.
10. The apparatus according to claim 9, wherein The matching module is specifically configured to: For each metadata category combination, calculate metadata similarity of the to-be-reviewed information content and each historical information content in the information comparison and review database; each metadata category combination corresponds to a user feedback category combination, each metadata category combination includes at least one metadata feature category, and each user feedback category combination includes at least one user feedback behavior feature category; According to the user feedback behavior feature category, merge user feedback behavior features of each historical information content with metadata similarity greater than a preset threshold to obtain the estimated user feedback behavior features of the to-be-reviewed information content.