Social network influence prediction method and system
By comprehensively considering user characteristics, content characteristics and network structure, combined with machine learning and natural language processing technology, a multi-dimensional influence scoring model and a communication probability calculation model are built, the limitations of influence prediction in the existing technology are solved, high-precision and multi-dimensional influence prediction are achieved, and marketing effectiveness and data quality are improved.
Patent Information
- Application Number
- CN202510234491.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing social network impact prediction methods have limitations in many aspects, and it is difficult to meet the accurate prediction needs in the increasingly complex social network environment, including ignoring the dynamics of user behavior, the importance of content characteristics, the complexity of information dissemination, and the control of data quality.
A high-precision and multi-dimensional social network influence prediction method is proposed. By comprehensively considering user characteristics, content characteristics and network structure, combining advanced machine learning and natural language processing technology, a multi-dimensional influence scoring model and a communication probability calculation model are constructed to generate an impact prediction report, including KOL ranking, hot topic analysis and optimal intervention timing suggestions.
It significantly improves the accuracy and comprehensiveness of impact assessments, can effectively capture the trend of influence changes over time and topics, provide more insightful decision support, improves marketing effectiveness, and improves data quality and forecast reliability.
Smart Images

Figure CN120163675A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of social information technology, in particular to a method and system for predicting social network influence. Background Art
[0002] With the rapid development of social media, predicting social network influence has become an important topic in fields such as brand marketing, public opinion analysis, and user behavior research. In recent years, the academic and industrial communities have extensively explored this issue and proposed various prediction methods and models. However, existing technologies still have many limitations and deficiencies, making it difficult to meet the accurate prediction requirements in the increasingly complex social network environment.
[0003] Traditional methods for predicting social network influence mainly rely on the static attributes of users and simple network topologies. For example, early studies often used PageRank-like algorithms, mainly considering the follow relationships between users to evaluate influence. Although such methods have high computational efficiency, they ignore the dynamic nature of user behavior and the importance of content features, resulting in prediction results that often deviate significantly from the actual situation.
[0004] With the development of machine learning technology, some researchers have begun to attempt to use supervised learning methods to predict user influence. These methods usually train prediction models based on historical data of users, such as the number of followers, the number of posts, and the amount of interaction. Although they have improved compared to traditional methods, they are still difficult to effectively capture the complex dynamic processes and long-term evolution trends in social networks. In addition, these methods often regard user influence as a static indicator, ignoring the changes in influence across different topics and time scales.
[0005] Recently, some studies have started to focus on modeling the information dissemination process and attempt to predict user influence by simulating information diffusion. However, most existing dissemination models are too simplistic to accurately reflect the complex dissemination mechanisms in real social networks. For example, many models assume that all users have the same activation probability, or ignore the impact of user interests and content relevance on dissemination, resulting in prediction results lacking accuracy and interpretability.
[0006] Another important issue is that existing methods generally lack effective control over the quality of social network data. In practical applications, there are a large number of machine-generated fake accounts and spam information on social platforms, and these noisy data will seriously affect the accuracy of prediction. However, there are few studies that systematically address this problem, and most methods assume the reliability of input data, ignoring the importance of data cleaning and quality control.
[0007] In addition, existing technologies also have obvious deficiencies in early hot spot identification and prediction of the optimal intervention timing. Traditional methods often rely on ex post facto statistics and it is difficult to timely discover potential hot topics. This results in brands often missing the best marketing opportunities and being unable to effectively grasp the rapidly changing public opinion trends on social media.
[0008] In summary, the existing social network influence prediction methods have limitations in multiple aspects and are difficult to meet the accurate prediction requirements in the current rapidly changing social media environment. Therefore, there is an urgent need for a more comprehensive, accurate, and reliable social network influence prediction method to address the above challenges. Summary of the Invention
[0009] The present invention aims to solve the above problems existing in the prior art and provides a high-precision and multi-dimensional social network influence prediction method and system. This method comprehensively considers user characteristics, content characteristics, and network structure, and combines advanced machine learning and natural language processing technologies to achieve accurate prediction and analysis of social network influence.
[0010] The present invention proposes a social network influence prediction method and system, including:
[0011] An acquisition step, including:
[0012] Obtaining user data, content data, and interaction data related to a specified brand from multiple social media platforms;
[0013] A processing step, including:
[0014] Based on the user data, content data, and interaction data, constructing a multi-dimensional influence scoring model;
[0015] According to the multi-dimensional influence scoring model, calculating the initial influence score of a user;
[0016] Based on the social network cascade propagation theory, constructing a propagation probability calculation model;
[0017] According to the propagation probability calculation model, predicting the information dissemination range and the change of user influence;
[0018] An output step, including:
[0019] Generating an influence prediction report, including the ranking of key opinion leaders (KOLs), hot topic analysis, and suggestions on the optimal intervention timing.
[0020] Preferably, the acquisition step specifically includes:
[0021] Using web crawler technology to scrape data from social media platforms;
[0022] Preprocess the captured data, including information verification, information cleaning, and word segmentation;
[0023] Among them, the information verification includes deleting duplicate, expired, or false information based on the authority and timeliness of the information source; the information cleaning includes filtering out punctuation marks, stop words, numbers, and emojis; the word segmentation process performs word segmentation on information, topics, and user comments based on semantic analysis.
[0024] Preferably, the multi-dimensional influence scoring model includes:
[0025] A user quality scoring sub-model that calculates the user quality score based on the number of followers, likes, reposts, and comments of the user;
[0026] A content quality scoring sub-model that calculates the content quality score based on the word frequency distribution related to the specified brand in the content published by the user;
[0027] Among them, the calculation results of the user quality scoring sub-model and the content quality scoring sub-model are weighted by weight coefficients to obtain the initial influence score of the user.
[0028] Preferably, the construction process of the propagation probability calculation model includes:
[0029] Construct a network cascade model based on the activation state and activation probability of user nodes;
[0030] Use sentiment analysis methods to calculate the relationship strength between users;
[0031] Calculate the user activation probability based on the activation degree, interest degree, relationship type, and initial influence of the user;
[0032] Adjust the propagation probability according to the degree of association between the user and the specified brand.
[0033] Preferably, it also includes an early hot spot identification step:
[0034] Extract sentences containing the specified brand from the acquired content data;
[0035] Use the BERT sentiment analysis model to perform sentiment classification on the sentences;
[0036] Input the sentiment classification results into the topic model to identify early hot spots related to the specified brand;
[0037] Predict the optimal intervention timing based on the development trend of early hot spots.
[0038] Preferably, it also includes a machine-generated fake account identification step:
[0039] Train a spam comment detection model based on the support vector machine algorithm;
[0040] Calculate the proportion of spam comments of users;
[0041] Use the BERT sentiment analysis model to conduct sentiment classification on spam comments and analyze the proportion of positive sentiment comments;
[0042] When the proportion of spam comment users exceeds the preset threshold, mark this user as a potential machine-forged account.
[0043] Preferably, it further includes a user label feature learning step:
[0044] Use user labels as nodes of graph data and the relationships between users as edges to construct an undirected network graph;
[0045] Adopt a graph enhancement network based on the attention mechanism to perform feature learning on the undirected network graph;
[0046] Obtain the feature vector representation of user labels.
[0047] Preferably, the construction process of the multi-dimensional influence scoring model further includes an influence index weight calculation step:
[0048] Construct a multi-level index system including interaction volume, dissemination volume, conversion volume, and user comprehensive score;
[0049] Adopt the analytic hierarchy process and the expert investigation method to calculate the subjective weight;
[0050] Adopt the entropy method and the coefficient of variation method to calculate the objective weight;
[0051] Perform weighted averaging on the subjective weight and the objective weight to obtain the comprehensive weight.
[0052] Preferably, it further includes a user-content-KOL preference analysis step based on singular value decomposition (SVD):
[0053] Construct a data matrix including content exposure volume, user fan volume, and interaction behavior;
[0054] Use the SVD method to decompose the data matrix into a user preference matrix for brand content, a user preference matrix for KOL, and a KOL preference matrix for users;
[0055] Based on the decomposed matrices, conduct KOL recommendation and influence evaluation.
[0056] A social network influence prediction system, including:
[0057] A data collection module, used to obtain user data, content data, and interaction data related to a specified brand from multiple social media platforms;
[0058] A data preprocessing module for performing information verification, information cleaning, and word segmentation on the data obtained by the data collection module;
[0059] An influence scoring module for constructing a multi-dimensional influence scoring model and calculating the initial influence score of a user;
[0060] A propagation model construction module for constructing a propagation probability calculation model based on the social network cascade propagation theory;
[0061] An influence prediction module for predicting the information dissemination range and the change of user influence according to the propagation probability calculation model;
[0062] A hot spot identification module for identifying early hot spots related to a specified brand and predicting the optimal intervention time;
[0063] An account authenticity verification module for identifying potential machine-forged accounts;
[0064] A user portrait construction module for constructing a user portrait based on user label feature learning;
[0065] A KOL recommendation module for performing KOL recommendations based on user-content-KOL preference analysis;
[0066] A report generation module for generating an influence prediction report including KOL rankings, hot topic analysis, and suggestions for the optimal intervention time.
[0067] The beneficial effects of the present invention are mainly reflected in the following aspects:
[0068] First of all, the multi-dimensional influence scoring model proposed by the present invention can comprehensively capture all aspects of user influence. By comprehensively considering user quality, content quality, and interaction quality, the model significantly improves the accuracy and comprehensiveness of influence evaluation. In particular, this method introduces a content quality evaluation based on sentiment analysis, effectively quantifying the actual influence of the content and avoiding biases that may be brought about by relying solely on surface indicators (such as the number of likes).
[0069] Secondly, the propagation probability calculation model constructed by the present invention based on the social network cascade propagation theory can accurately simulate the diffusion process of information in the social network. The model not only considers the relationship strength between users but also introduces factors such as user interests and content relevance, making the prediction results closer to the actual propagation situation. This dynamic propagation model enables this method to effectively capture the trend of influence changing with time and topics, providing more insightful decision-making support for brands.
[0070] Thirdly, the early hot spot identification and optimal intervention timing prediction functions of the present invention provide a powerful tool for brands to seize marketing opportunities in a timely manner. By combining BERT sentiment analysis and LDA topic model, this method can accurately identify potential hot spots when the topic first emerges and predict its development trend. This enables brands to intervene in the topic discussion at the best time and significantly improve the marketing effect.
[0071] In addition, the machine-generated fake account identification function of the present invention effectively improves the data quality and prediction reliability. By comprehensively analyzing user behavior patterns and content features, this method can accurately identify and filter machine-generated fake accounts, ensuring the data quality for subsequent analysis. This function not only improves the prediction accuracy but also provides brands with more real and reliable social network insights.
[0072] Finally, the KOL recommendation function of the present invention, based on innovative user-content-KOL preference analysis, can provide brands with more accurate advice on selecting opinion leaders. By deeply exploring the potential relationships among users, content, and KOLs, this method can identify the most influential KOLs that are most compatible with the brand, thereby effectively improving the accuracy and effectiveness of brand cooperation.
[0073] In summary, the social network influence prediction method and system provided by the present invention, through the organic combination and synergistic effect of multiple innovative points, significantly improve the accuracy, comprehensiveness, and practicality of prediction. This not only solves many problems existing in the prior art but also provides a powerful support tool for brands to formulate accurate marketing strategies in the complex and ever-changing social media environment. The method of the present invention is expected to play an important role in multiple fields such as social media marketing, public opinion analysis, and user behavior research, promoting the further development of related technologies and applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 is the overall flowchart of the social network influence prediction method of the present invention;
[0075] Figure 2 is the detailed flowchart of the data preprocessing module of the present invention;
[0076] Figure 3 is the working flowchart of the influence scoring module of the present invention;
[0077] Figure 4 is the flowchart of the propagation model construction module of the present invention;
[0078] Figure 5 is the working flowchart of the hot spot identification module of the present invention;
[0079] Figure 6 is the flowchart of the account authenticity verification module of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0080] Please refer to the appendix Figure 1-6 For this invention, a social network influence prediction method and its system are provided. Through multi-dimensional data collection, in-depth analysis, and modeling, this method can accurately predict the influence of users in the social network and its changing trends, providing strong support for brand marketing decisions. The following will elaborate on this invention in combination with specific implementation manners.
[0081] The social network influence prediction method of this invention includes an acquisition step, a processing step, and an output step.
[0082] In the acquisition step, this method obtains user data, content data, and interaction data related to a specified brand from multiple social media platforms. Preferably, this invention uses distributed crawler technology to simultaneously collect data from multiple mainstream social platforms (such as Weibo, WeChat, Douyin, etc.) to ensure the comprehensiveness and real-time nature of the data. For example, for a certain clothing brand, the following data can be collected, including but not limited to: user basic information (such as the number of followers, the number of people followed), content published by users (text, pictures, videos), interaction behaviors between users (such as likes, comments, forwards), etc.
[0083] In the processing step, first, a multi-dimensional influence scoring model is constructed based on the acquired data. In one embodiment of this invention, this model comprehensively considers three dimensions: user quality, content quality, and interaction quality. Specifically, the following formula can be used to calculate the initial influence score of a user:
[0084] I0 = w u ·Q u + w c ·Q c + w i ·Q i .
[0085] Wherein, I0 is the initial influence score of the user, Q u , Q c and Q i are the user quality score, content quality score, and interaction quality score respectively, and w u , w c and w i are the corresponding weight coefficients. Empirically, w u can be set to 0.4, w c to 0.3, w i to 0.3, but the specific weights can be adjusted according to the characteristics of different brands and industries.
[0086] Next, this method constructs a propagation probability calculation model based on the social network cascade propagation theory. In a preferred embodiment, the Independent Cascade model is adopted and improved by combining the relationship strength between users. The propagation probability can be expressed as:
[0087]
[0088] where P ij is the probability that user i influences user j, R ij is the relationship strength between users i and j, I i is the influence score of user i, and α and β are adjustable parameters. Through a large number of experiments, it is found that when α = 0.6 and β = 0.4, the model has the best effect.
[0089] Based on the above model, this method can predict the information dissemination range and the change of user influence. For example, the average dissemination range and the change amplitude of the influence of each user can be calculated by simulating the propagation process multiple times through Monte Carlo.
[0090] In the output step, this method generates an influence prediction report, including the ranking of key opinion leaders (KOLs), the analysis of hot topics, and the suggestion of the optimal intervention time. Preferably, the KOL ranking is dynamically adjusted based on the predicted influence change, the hot topic analysis uses the LDA topic model to extract keywords and calculate the topic popularity, and the optimal intervention time is determined according to the predicted propagation speed and range.
[0091] The acquisition step of the present invention also includes data preprocessing. Specifically, first, web crawler technology is used to crawl data from social media platforms. In an embodiment of the present invention, the Scrapy framework is used to construct a distributed crawler system, and by setting multiple proxy IPs and simulating user behaviors, the interference of the anti-crawler mechanism is effectively avoided.
[0092] Subsequently, the crawled data is preprocessed, including information verification, information cleaning, and word segmentation. In the information verification link, this method deletes duplicate, expired, or false information based on the authority and timeliness of the information source. For example, the time threshold can be set to 30 days to delete expired information that exceeds 30 days; by comparing the information consistency of multiple sources, false information is identified and deleted.
[0093] In the information cleaning link, this method filters out punctuation marks, stop words, numbers, and emojis. Preferably, regular expressions are used for matching and replacement to improve the processing efficiency. For example, the following regular expression can be used to remove emoji expressions.
[0094] In the word segmentation process, this method performs word segmentation on information, topics, and user comments based on semantic analysis. Preferably, for Chinese texts, the jieba word segmentation library can be used, and a custom dictionary can be combined to improve the accuracy of word segmentation.
[0095] The multi-dimensional influence scoring model of the present invention includes a user quality evaluation sub-model and a content quality evaluation sub-model.
[0096] The user quality evaluation sub-model calculates the user quality score based on the number of fans, the number of likes, the number of reposts, and the number of comments of the user. In an embodiment of the present invention, the following formula is used to calculate the user quality score:
[0097] Q u =w f ·log(1 + F)+w l ·log(1 + L)+w r ·log(1 + R)+w c ·log(1 + C),
[0098] where F, L, R, and C are the number of fans, the number of likes, the number of reposts, and the number of comments of the user respectively, and w f , w l , w r , w c are the corresponding weight coefficients. Using the logarithmic function can effectively handle the problem of uneven data distribution. Empirically, it can be set that w f =0.3, w l =0.2, w r =0.3, w c =0.2, but the specific weights can be adjusted according to the characteristics of different platforms.
[0099] The content quality evaluation sub-model calculates the content quality score based on the word frequency distribution related to the specified brand in the content published by the user. Preferably, the TF-IDF algorithm is used to calculate the importance of brand-related words, and sentiment analysis technology is combined to evaluate the content quality. The specific formula is as follows:
[0100]
[0101] where w i is the keyword related to the brand, TF-IDF(w i ) is the TF-IDF value of this word, and S(w i ) is the sentiment score of this word (1 for positive, 0 for neutral, -1 for negative). n is the total number of keywords, which can be set according to actual needs, for example, taking the top 10 keywords.
[0102] Finally, the user quality score and the content quality score are weighted by the weight coefficients to obtain the initial influence score of the user. Preferably, the user quality weight can be set to 0.6 and the content quality weight can be set to 0.4, that is:
[0103] I0 = 0.6·Q u + 0.4·Q c ,
[0104] Through the above method, the present invention can comprehensively and accurately evaluate the initial influence of social network users, laying a foundation for subsequent propagation prediction and KOL identification.
[0105] The construction process of the propagation probability calculation model of the present invention includes multiple steps, aiming to accurately simulate the information propagation process in the social network.
[0106] First, the method constructs a network cascade model based on the activation state and activation probability of user nodes. In a preferred embodiment, an improved version of the Independent Cascade (IC) model is adopted. Specifically, a social network G=(V, E) is defined, where V is the set of user nodes and E is the set of user relationship edges. For each edge (u, v)∈E, a propagation probability p(u, v) is assigned. The propagation process can be described as follows:
[0107] When node u is activated at time step t, it will attempt to activate all its unactivated neighbor nodes v. The probability of successful activation is p(u, v). Whether successful or not, u will not attempt to activate v again in subsequent time steps. This process continues until no new nodes are activated.
[0108] Next, the method of the present invention uses sentiment analysis technology to calculate the relationship strength between users. Preferably, the BERT (Bidirectional Encoder Representations from Transformers) model is used to perform sentiment analysis on the interaction content between users. The relationship strength can be expressed as:
[0109]
[0110] where, R ij is the relationship strength between users i and j, S k is the sentiment score of the kth interaction content (range: [-1, 1]), and n is the total number of interaction contents.
[0111] In an embodiment of the present invention, the user activation probability is calculated based on the user's activation degree, interest degree, relationship type, and initial influence. The activation probability can be expressed as:
[0112] P a(u) = w1·A(u) + w2·I(u) + w3·T(u) + w4·I0(u),
[0113] where P a (u) is the activation probability of user u, A(u) is the activation degree of the user (which can be calculated from the recent activity frequency), I(u) is the user's interest in the relevant topic, T(u) is the weight of the user's relationship type (such as friend, fan, etc.), and I0(u) is the user's initial influence score. w1, w2, w3, and w4 are weight coefficients, and preferably can be set to 0.3, 0.2, 0.2, and 0.3.
[0114] Finally, the method adjusts the propagation probability according to the degree of association between the user and the specified brand. In a preferred embodiment of the present invention, the following formula is used for adjustment:
[0115] p′(u,v) = p(u,v)·(1 + β·B(u)),
[0116] where p′(u,v) is the adjusted propagation probability, p(u,v) is the original propagation probability, B(u) is the degree of association between user u and the brand (ranging from [0,1]), and β is an adjustment coefficient, which can be set according to the actual situation, and the preferred value is 0.5.
[0117] Through the above steps, the present invention constructs a propagation probability calculation model that comprehensively considers various factors, and can more accurately simulate the information propagation process in the social network.
[0118] The present invention also includes an early hotspot identification step. This step is of great significance for timely discovering brand-related hot topics and grasping marketing opportunities.
[0119] First, the method extracts sentences containing the specified brand from the obtained content data. In one embodiment, regular expressions can be used to match sentences containing the brand name and its variants.
[0120] Next, the method of the present invention uses the BERT sentiment analysis model to perform sentiment classification on the extracted sentences. Preferably, a fine-tuned BERT model is used to classify the sentiment into three categories: positive, neutral, and negative. The sentiment classification results can be used for subsequent hotspot analysis to help the brand understand the user's attitude towards different topics.
[0121] Subsequently, the method inputs the sentiment classification results into a topic model to identify early hotspots related to the specified brand. In a preferred embodiment of the present invention, the LDA (Latent Dirichlet Allocation) topic model is used. The LDA model can be expressed as:
[0122] P(w|d) = ∑ zP(w|z)P(z|d),
[0123] where w represents a word, d represents a document, and z represents a topic. P(w|z) represents the probability that word w appears in topic z, and P(z|d) represents the probability that topic z appears in document d.
[0124] Through the LDA model, multiple latent topics can be extracted, and each topic is represented by a set of keywords. Combining the word frequency and the sentiment analysis results, the popularity score of each topic can be calculated:
[0125] H t = ∑ w∈t TF(w)·(1 + α·S(w)),
[0126] where H t is the popularity score of topic t, TF(w) is the word frequency of word w, S(w) is the average sentiment score of word w (ranging from [-1, 1]), α is a regulation coefficient, and the preferred value is 0.5.
[0127] Finally, the method of the present invention predicts the optimal intervention timing based on the development trend of early hotspots. In one embodiment, a time series analysis method, such as the ARIMA (AutoRegressive Integrated Moving Average) model, can be used to predict the future popularity change of hot topics. The optimal intervention timing can be defined as the time point with the highest popularity growth rate:
[0128]
[0129] where H(t) is the function of the popularity changing with time t, represents the growth rate of the popularity.
[0130] Through the above steps, the present invention can timely identify the early hotspots related to the brand and provide suggestions on the optimal intervention timing for the brand, which helps to improve the marketing effect.
[0131] The present invention further includes the step of identifying machine - forged accounts. This step is of great significance for ensuring the accuracy of social network data analysis.
[0132] First, this method trains a spam comment detection model based on the support vector machine (SVM) algorithm. In a preferred embodiment of the present invention, an SVM model with a radial basis function (RBF) kernel is adopted. The decision function of the model can be expressed as:
[0133]
[0134] where x is the input feature vector, x i is the support vector, and y iis the class label, α i is the Lagrange multiplier, b is the bias term, and K(x i , x) is the kernel function. For the RBF kernel, K(x i , x) = exp(-γ|x i - x| 2 ), where γ is the kernel parameter.
[0135] During model training, the following features can be used: comment length, keyword frequency, number of URLs, proportion of special characters, comment posting time interval, etc. The optimal hyperparameters, such as C (penalty parameter) and γ, are selected through cross-validation.
[0136] Next, this method calculates the proportion of spam comments of the user. For each user u, the proportion of its spam comments can be expressed as:
[0137]
[0138] where, N spam (u) is the number of spam comments posted by user u, and N total (u) is the total number of comments posted by user u.
[0139] Subsequently, the method of the present invention uses the BERT sentiment analysis model to perform sentiment classification on spam comments and analyze the proportion of positive sentiment comments. The purpose of this step is to identify those machine accounts that may improve the brand reputation by posting a large number of positive comments. The proportion of positive sentiment comments can be expressed as:
[0140]
[0141] where, N pos (u) is the number of positive sentiment spam comments posted by user u.
[0142] Finally, when the proportion of spam comment users exceeds the preset threshold, this method marks this user as a potential machine forgery account. Specifically, two thresholds can be set: the spam comment proportion threshold T s pam and the positive sentiment comment proportion threshold T pos . The judgment criterion can be expressed as:
[0143]
[0144] where, IsFake(u) = 1 indicates that user u is determined to be a machine forgery account. According to experience, T spam = 0.7, T pos = 0.8, but the specific thresholds can be adjusted according to the actual situation.
[0145] Through the above steps, the present invention can effectively identify machine - forged accounts in social networks, improving the accuracy and reliability of data analysis.
[0146] The present invention also includes a user - tag feature learning step. This step aims to better understand and represent the interests and characteristics of users, providing support for subsequent influence prediction and content recommendation.
[0147] First, the method constructs an undirected network graph with user tags as nodes of the graph data and user - to - user relationships as edges. In a preferred embodiment, the NetworkX library can be used to construct and operate the graph structure. For example:
[0148]
[0149]
[0150] Among them, `user_tags` is a dictionary that stores users and their corresponding tag sets; `user_relations` is a list that stores user - to - user relationship pairs.
[0151] Next, the method of the present invention uses a graph enhancement network based on the attention mechanism to perform feature learning on the constructed undirected network graph. In an embodiment of the present invention, the Graph Attention Network (GAT) model is used. The core of GAT is to assign different weights to neighbor nodes through the attention mechanism, thereby aggregating neighbor information. For node i, its representation update can be described as:
[0152]
[0153] Among them, represents the feature representation of node i in the I - th layer, is the set of neighbors of node i, W (l) is the weight matrix of the I - th layer, and σ is an activation function (such as ReLU). The attention coefficient α ij is calculated in the following way:
[0154]
[0155] Among them, a is the attention vector, and | represents the concatenation operation.
[0156] During model training, a node classification task can be used as a supervision signal. For example, users can be classified into different interest categories, and then the cross - entropy loss is minimized:
[0157]
[0158] Among them, γ is the set of labeled nodes, C is the number of categories, and Yic is the true label, P ic is the predicted probability.
[0159] Finally, the method obtains the feature vector representation of user labels. After the model training is completed, the output of the last layer can be directly used as the feature vector of user labels. These feature vectors capture the semantic information and correlation of user labels in the social network structure.
[0160] Through the above steps, the present invention can learn rich feature representations of user labels, and these features can be used for subsequent tasks such as user portrait construction, similar user discovery, and personalized recommendation, thereby improving the accuracy and application value of social network influence prediction.
[0161] The construction process of the multi-dimensional influence scoring model of the present invention further includes an influence index weight calculation step. This step aims to scientifically and reasonably determine the importance of each influence index, thereby improving the accuracy and interpretability of influence scoring.
[0162] First, the method constructs a multi-level index system including interaction volume, dissemination volume, conversion volume, and user comprehensive score. In a preferred embodiment of the present invention, the index system can be expressed as follows:
[0163] 1. Interaction volume; number of likes; number of comments; number of shares;
[0164] 2. Dissemination volume; number of forwards; number of citations;
[0165] 3. Conversion volume; click-through rate; conversion rate;
[0166] 4. User comprehensive score; activity; influence persistence;
[0167] Next, the method of the present invention uses the Analytic Hierarchy Process (AHP) and the expert survey method to calculate the subjective weight. In the AHP method, first construct the judgment matrix A:
[0168]
[0169] where a ij represents the importance degree of index i relative to index j, and usually the 1-9 scale method is adopted. Then calculate the eigenvector w:
[0170] Aw = λ max w,
[0171] where λ max is the maximum eigenvalue. Normalize w to obtain the subjective weight vector.
[0172] Subsequently, the method uses the entropy value method and the coefficient of variation method to calculate the objective weight. For the entropy value method, first calculate the entropy value of the j-th index:
[0173]
[0174] Among them, x ij is the j-th index value of the i-th sample. Then calculate the weights:
[0175]
[0176] For the coefficient of variation method, calculate the coefficient of variation of each index:
[0177]
[0178] Among them, σ j and are the standard deviation and average value of the j-th index respectively. Then calculate the weights:
[0179]
[0180] Finally, the method of the present invention performs a weighted average on the subjective weight and the objective weight to obtain the comprehensive weight. Preferably, the following formula is used:
[0181] W = αW subjective +(1 - α)W objective
[0182] Among them, α is the importance coefficient of the subjective weight and can be adjusted according to the actual situation. In an embodiment of the present invention, α can be set to 0.6, that is, slightly emphasizing the subjective weight to make full use of expert experience.
[0183] Through the above steps, the present invention can scientifically and reasonably determine the weights of each influence index, providing a reliable basis for the multi-dimensional influence scoring model.
[0184] The present invention further includes a user-content-KOL preference analysis step based on singular value decomposition (SVD). This step aims to deeply explore the potential relationship between users, content, and KOLs, providing support for precision marketing and personalized recommendation.
[0185] First of all, this method constructs a data matrix R containing content exposure, user fan volume, and interaction behavior. In a preferred embodiment, this matrix can be expressed as:
[0186]
[0187] Among them, r ij can be the interaction intensity of user i with content j, or the attention degree of user i to KOL j, etc. Next, the method of the present invention uses the SVD method to decompose the data matrix R into the product of three matrices:
[0188] R = UΣV T ,
[0189] where U is an m×m orthogonal matrix representing the user feature space; V is an n×n orthogonal matrix representing the content / KOL feature space; and Z is an m×n diagonal matrix, and the elements on the diagonal, σ i are called singular values and represent the importance of the features.
[0190] In practical applications, usually only the first k largest singular values are retained to obtain the matrix after dimensionality reduction:
[0191]
[0192] where the selection of k can be based on the cumulative variance contribution rate. For example, select the value of k that makes the cumulative variance contribution rate reach 90%.
[0193] Based on the decomposed matrix, this method performs KOL recommendation and influence evaluation. For KOL recommendation, the cosine similarity between the user vector and the KOL vector can be calculated:
[0194]
[0195] where u i is the feature vector of user i, and v j is the feature vector of KOL j. Select the TopN KOLs with the highest similarity for recommendation.
[0196] For influence evaluation, it can be calculated based on the position and importance of the KOL in the feature space. For example, the following formula can be used:
[0197]
[0198] where σ i is the i-th singular value, and v ji is the value of KOL j in the i-th feature dimension.
[0199] Through the above steps, the present invention can deeply analyze the potential relationship between users - content - KOLs, providing strong support for precision marketing and personalized recommendation.
[0200] The present invention also provides a social network influence prediction system. The system includes multiple functional modules, and each module works in cooperation to implement the functions of the method of the present invention. The following is a detailed description of each module:
[0201] The data collection module 1 is used to obtain user data, content data, and interaction data related to a specified brand from multiple social media platforms. This module can adopt distributed crawler technology, support concurrent collection across multiple platforms, and ensure the comprehensiveness and real-time nature of the data.
[0202] The data preprocessing module 2 is used to perform information verification, information cleaning, and word segmentation on the data obtained by the data collection module 1. This module can integrate various preprocessing algorithms, such as regular expression matching, stop word filtering, Chinese word segmentation, etc., to improve the accuracy of subsequent analysis.
[0203] The influence scoring module 3 is used to construct a multi-dimensional influence scoring model and calculate the initial influence score of users. This module implements the multi-dimensional scoring mechanism of the present invention, comprehensively considering user quality, content quality, and interaction quality.
[0204] The propagation model construction module 4 is used to construct a propagation probability calculation model based on the social network cascade propagation theory. This module implements an improved Independent Cascade model, considering factors such as user relationship strength and activation probability.
[0205] The influence prediction module 5 is used to predict the information dissemination range and the change of user influence according to the propagation probability calculation model. This module can adopt methods such as Monte Carlo simulation to achieve dynamic influence prediction.
[0206] The hot spot identification module 6 is used to identify early hot spots related to a specified brand and predict the optimal intervention time. This module integrates BERT sentiment analysis and LDA topic model, and can timely discover potential hot topics.
[0207] The account authenticity verification module 7 is used to identify potential machine-forged accounts. This module implements spam comment detection based on SVM and BERT sentiment analysis, effectively improving data quality.
[0208] The user portrait construction module 8 is used to construct a user portrait based on user label feature learning. This module uses a graph attention network (GAT) for feature learning to capture the semantic information and relevance of user labels.
[0209] The KOL recommendation module 9 is used to recommend KOLs based on user-content-KOL preference analysis. This module implements a matrix factorization algorithm based on SVD to deeply explore potential relationships.
[0210] The report generation module 10 is used to generate an influence prediction report containing KOL rankings, hot topic analysis, and suggestions for the optimal intervention time. This module integrates the results of each analysis module to provide intuitive and actionable decision support.
[0211] The above-mentioned modules exchange information and cooperate with each other through a data bus. The system may also include a user interface module, which provides interactive data visualization and parameter configuration functions.
[0212] Through the collaborative work of the above-mentioned modules, the social network influence prediction system of the present invention can comprehensively and accurately analyze and predict the influence spread in the social network, providing strong support for brand marketing decisions.
[0213] In order to verify the superiority of the social network influence prediction method and system of the present invention, a series of simulation experiments were conducted. The following will introduce the embodiments, comparative examples, and related test results and analyses in detail.
[0214] Embodiment 1: Social network influence prediction method of the present invention
[0215] In this embodiment, the marketing campaign of a well-known sports brand on the Weibo platform was selected as the research object. The brand planned to launch a new type of running shoes and hoped to optimize its marketing strategy through social media influence prediction. Using the method of the present invention, the data of 100,000 users were analyzed for 30 days.
[0216] Comparative Example 1: Traditional PageRank algorithm
[0217] In this comparative example, the classic PageRank algorithm was used to evaluate user influence. This algorithm mainly calculates the influence score based on the follow relationship between users.
[0218] Comparative Example 2: Machine learning method based on user attributes
[0219] This comparative example adopted a machine learning method based on user attributes (such as the number of fans, the number of posts, etc.) and used the random forest algorithm for influence prediction.
[0220] Test metrics and methods:
[0221] 1. Prediction accuracy: The root mean square error (RMSE) was used to measure the difference between the prediction result and the actual spread result.
[0222] 2. Early hot topic recognition rate: Calculate the proportion of successfully recognized early hot topics among all hot topics.
[0223] 3. KOL recommendation accuracy: Calculate the proportion of KOLs recommended that actually have a significant impact.
[0224] 4. Computational efficiency: Record the time required to complete the entire prediction process.
[0225] 5. Machine-generated fake account recognition rate: Calculate the proportion of successfully recognized machine-generated fake accounts among all fake accounts.
[0226] Test results:
[0227] Index Example 1 Comparative Example 1 Comparative Example 2 Prediction accuracy (RMSE) 0.082 0.215 0.163 Early hot spot recognition rate 87.5% Not applicable 62.3% KOL recommendation accuracy 92.1 78.6 83.4 Computing efficiency (hours) 2.5 1.8 3.2 Machine forged account recognition rate 94.7% Not applicable 81.2%
[0228] Analysis and discussion are as follows:
[0229] 1. Prediction accuracy: The method of the present invention is significantly superior to the two comparative methods in terms of prediction accuracy. This is mainly due to the multi-dimensional influence scoring model and the calculation of propagation probability based on the social network cascade propagation theory. The lower RMSE value (0.082) indicates that this method can more accurately predict the information dissemination range and the change of user influence.
[0230] 2. Early hot spot recognition rate: The method of the present invention performs excellently in the recognition of early hot spots and successfully identifies 87.5% of the hot topics. This benefits from the combined use of BERT sentiment analysis and LDA topic model. Comparative example 1 does not have this function, while the recognition rate of comparative example 2 is relatively low (62.3%), indicating that this method has obvious advantages in timely discovering potential hot spots.
[0231] 3. KOL recommendation accuracy: The method of the present invention leads the comparative methods in terms of KOL recommendation accuracy. The high accuracy rate of 92.1% indicates that the user-content-KOL preference analysis based on SVD can effectively capture complex social relationships and provide more accurate KOL selection for the brand.
[0232] 4. Computational efficiency: Although the computational time of this method (2.5 hours) is slightly longer than that of the traditional PageRank algorithm (1.8 hours), considering the comprehensiveness of the function and the accuracy of the prediction, this computational time is acceptable. Compared with the machine learning method based on user attributes (3.2 hours), this method also has a slight advantage.
[0233] 5. Machine-generated account recognition rate: The method of the present invention performs excellently in the recognition of machine-generated accounts, reaching a recognition rate of 94.7%. This is significantly higher than 81.2% of comparative example 2, while comparative example 1 does not have this function. The high recognition rate helps to improve the data quality and thus enhance the prediction accuracy.
[0234] In practical applications, the method of the present invention has successfully helped the sports brand optimize its marketing strategy for the new running shoes. By accurately predicting the influence spread and timely identifying hot topics, the brand intervened in several high-potential topic discussions at the best time. At the same time, based on accurate KOL recommendations, the brand selected 5 of the most influential sports bloggers for cooperation, significantly enhancing the product exposure and user engagement.
[0235] Based on the above test results, the optimal implementation parameters of the method of the present invention are determined:
[0236] 1. In the multi-dimensional influence scoring model, the weight ratios of user quality, content quality, and interaction quality are set to 4:3:3.
[0237] 2. In the propagation probability calculation model, the relationship strength weight α is set to 0.6, and the initial influence weight β is set to 0.4.
[0238] 3. For early hot spot identification, 10 LDA topics are used, and the sentiment analysis threshold is set to ±0.5.
[0239] 4. For the identification of machine-forged accounts, the threshold of the spam comment ratio is set to 0.7, and the threshold of the positive sentiment comment ratio is set to 0.8.
[0240] Using these parameters, the method of the present invention has achieved optimal performance in various indicators, especially in terms of prediction accuracy and KOL recommendation accuracy, with significant improvements.
[0241] In summary, the social network influence prediction method and system of the present invention are significantly superior to traditional methods in multiple key indicators. It can not only accurately predict the spread of influence, but also identify early hot spots, recommend accurate KOLs, and effectively filter machine-forged accounts. These advantages enable this method to have a wide range of application prospects in actual marketing scenarios and can provide more comprehensive and accurate decision-making support for brands.
[0242] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A social network influence prediction method, characterized in that: include: The acquisition steps include: Obtain user data, content data, and interaction data related to a given brand from multiple social media platforms; Processing steps include: Building a multi-dimensional influence scoring model based on the user data, content data and interaction data; Calculate the user's initial influence score according to the multi-dimensional influence scoring model; Based on the social network cascade propagation theory, a propagation probability calculation model is constructed; According to the propagation probability calculation model, predict the information dissemination scope and user influence changes; Output steps include: Generate influence prediction reports, including key opinion leaders (KOL) rankings, hot topic analysis, and recommendations on optimal timing for intervention.
2. The social network influence prediction method according to claim 1, characterized in that: The acquisition step specifically includes: Use web crawler technology to extract data from social media platforms; Preprocessing the captured data, including information verification, information cleaning and word segmentation processing; Among them, the information verification includes deleting duplicate, expired or false information based on the authority and timeliness of the information source; the information cleaning includes filtering out punctuation marks, stop words, numbers and emoticons; the word segmentation processing segments information, topics and user comments based on semantic analysis.
3. The social network influence prediction method according to claim 1, characterized in that: The multi-dimensional influence scoring model includes: The user quality scoring sub-model calculates the user quality score based on the number of followers, likes, reposts, and comments of the user; The content quality scoring sub-model calculates the content quality score based on the frequency distribution of words related to the specified brand in the content posted by users; The calculation results of the user quality scoring sub-model and the content quality scoring sub-model are weighted by a weight coefficient to obtain the user's initial influence score.
4. The social network influence prediction method according to claim 1, characterized in that: The construction process of the propagation probability calculation model includes: Construct a network cascade model based on the activation status and activation probability of user nodes; Use sentiment analysis methods to calculate the strength of relationships between users; Calculate user activation probability based on user activation, interest, relationship type, and initial influence; Adjust the probability of communication based on the user's degree of association with a given brand.
5. The social network influence prediction method according to claim 1, characterized in that: Also included are early hotspot identification steps: extracting sentences containing a specified brand from the acquired content data; Use the BERT sentiment analysis model to classify the sentiment of the sentence; The sentiment classification results are fed into a topic model to identify early hot spots related to a given brand; Predict the optimal time for intervention based on the development trend of early hot spots.
6. The social network influence prediction method according to claim 1, characterized in that: It also includes machine-forged account identification steps: Training spam detection model based on support vector machine algorithm; Calculate the percentage of spam comments by users; Use the BERT sentiment analysis model to classify spam comments and analyze the proportion of positive sentiment comments; When the proportion of spam comment users exceeds a preset threshold, the user will be marked as a potential machine-forged account.
7. The social network influence prediction method according to claim 1, characterized in that: It also includes the user label feature learning step: Use user labels as nodes of graph data and relationships between users as edges to construct an undirected network graph; Using a graph enhancement network based on an attention mechanism to perform feature learning on the undirected network graph; Get the feature vector representation of the user tag.
8. The social network influence prediction method according to claim 1, characterized in that: The construction process of the multi-dimensional influence scoring model also includes the step of calculating the influence indicator weights: Build a multi-level indicator system including interaction volume, dissemination volume, conversion volume and user comprehensive score; The subjective weights were calculated using the analytic hierarchy process and expert survey method; The objective weights were calculated using the entropy method and the coefficient of variation method; The subjective weight and the objective weight are weighted averaged to obtain a comprehensive weight.
9. The social network influence prediction method according to claim 1, characterized in that: It also includes the user-content-KOL preference analysis steps based on singular value decomposition (SVD): Construct a data matrix including content exposure, user fans and interactive behaviors; Decomposing the data matrix into a user-to-brand content preference matrix, a user-to-KOL preference matrix, and a KOL-to-user preference matrix using an SVD method; KOL recommendations and influence assessment are performed based on the decomposed matrix.
10. A social network influence prediction system for executing the method according to any one of claims 1 to 9, characterized in that: include: A data collection module for acquiring user data, content data, and interaction data related to a specified brand from multiple social media platforms; A data preprocessing module, used to perform information verification, information cleaning and word segmentation processing on the data acquired by the data acquisition module; The influence scoring module is used to build a multi-dimensional influence scoring model and calculate the user's initial influence score; The propagation model building module is used to build a propagation probability calculation model based on the social network cascade propagation theory; An influence prediction module, used to predict the information dissemination scope and user influence changes according to the dissemination probability calculation model; Hotspot identification module, used to identify early hotspots related to a specified brand and predict the best time to intervene; Account authenticity verification module, used to identify potential machine-forged accounts; User portrait construction module, used to build user portraits based on user label feature learning; KOL recommendation module, used to make KOL recommendations based on user-content-KOL preference analysis; The report generation module is used to generate influence prediction reports including KOL rankings, hot topic analysis and optimal intervention timing recommendations.
Citation Information
Cited By
High-quality user mining determination method and system based on automobile brand
CN120409971A
Internet information propagation effect analysis method based on multi-dimensional data
CN120492864A
Multi-dimensional feature fusion dispute detection method and system based on attention mechanism
CN121093290A
AI intelligent marketing person-reaching recommendation method based on brand demands
CN121167325A
Social network organization influence assessment method and assessment system thereof
CN121391242A