Marketing risk early warning method and system based on multi-channel public opinion perception
Patent Information
- Application Number
- CN202611339617.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-09-01
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明提供了一种基于多渠道舆情感知的营销风险预警方法及系统,以解决无法动态追踪跨平台舆情传播路径并识别关键扩散节点的问题
[0016](1)本发明通过获取多渠道的负面信息文本,提取负面信息文本中的跨平台传播特征并按发布时间戳排序,构建初始传播轨迹序列;再根据该序列提取不同平台用户之间的转发交互关系,构建跨平台的节点关联图谱,获得结构化的表达矩阵。该方案通过将各渠道非结构化的舆情文本转换为统一的图结构表达形式,从而实现了从碎片化、孤岛化的单平台数据到全局关联性数据的有效整合,为跨平台舆情传播路径的动态追踪提供了基础支撑。
Smart Images

Figure CN122838641A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of public opinion monitoring technology, and in particular to a marketing risk early warning method and system based on multi-channel public opinion perception. Background Technology
[0002] In the current era of digital marketing, corporate brand reputation and market performance are highly dependent on public opinion feedback from multiple channels, including social media, short video platforms, and news portals. Negative information in marketing campaigns often spreads across platforms in a short period, creating a media storm and causing serious brand damage to companies. How to perceive public opinion across multiple channels in real time and provide early warnings of marketing risks has become a key issue in corporate marketing management.
[0003] In existing technologies, current public opinion monitoring methods typically employ techniques such as keyword matching and sentiment analysis to identify negative information. Taking the monitoring of a brand's marketing campaign as an example, the system first collects text data from various platforms. It then uses a pre-trained text classification model based on deep learning to independently perform semantic analysis and polarity classification on the text from each platform, marking negative content. Subsequently, based on the keyword matching results, it statistically analyzes the frequency and trends of negative information mentions within each platform, generating independent public opinion monitoring reports for each platform. The monitoring results from each channel are presented in independent reports or dashboards, with only statistical-level summaries between them. Although some existing solutions have attempted to cluster topics through cross-platform text semantic similarity calculations, these solutions can only form flat, static topic groups and cannot construct a temporal propagation chain, thus failing to track the propagation trajectory and evolution of the same negative topic across different platforms. Furthermore, existing solutions cannot effectively identify key influencing nodes in the cross-platform propagation chain.
[0004] In summary, existing technologies have the problem of being unable to dynamically track the cross-platform spread of public opinion and identify key diffusion nodes. Summary of the Invention
[0005] This invention provides a marketing risk early warning method and system based on multi-channel public opinion perception to solve the problem of being unable to dynamically track the cross-platform public opinion propagation path and identify key diffusion nodes.
[0006] Firstly, in order to solve the above-mentioned technical problems, the present invention provides a marketing risk early warning method based on multi-channel public opinion perception, including:
[0007] Text data from different platforms is acquired, the text data is preprocessed, and an initial propagation trajectory sequence is generated.
[0008] Based on the initial propagation trajectory sequence, the forwarding interaction relationship between users on different platforms is extracted. Based on the forwarding interaction relationship, the source node and secondary node are determined and node connections are established. Based on each node and the node connection, a cross-platform node association graph is constructed. The node association graph is subjected to graph structure embedding processing to obtain a structure expression matrix.
[0009] Based on the structural representation matrix, a graph neural network is used to generate deep topological features. The deep topological features are then clustered and deduplicated to generate a multi-level diffusion path set.
[0010] The centrality value of each node is calculated based on the multi-level diffusion path set, and the nodes whose centrality value is greater than the preset centrality threshold are identified as key nodes.
[0011] Obtain the text sequence generated by the key node, perform sentiment polarity analysis on the text sequence to obtain the sentiment polarity change, and concatenate the sentiment polarity change to obtain the evolutionary feature vector;
[0012] The evolutionary feature vector is input into a pre-trained random forest model, which outputs a critical probability value. If the critical probability value is greater than a preset warning trigger threshold, a warning instruction is generated for the key node. Based on the warning instruction, a risk prediction result for cross-platform propagation is generated.
[0013] Secondly, the present invention provides a marketing risk early warning system based on multi-channel public opinion perception, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described above.
[0014] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described above.
[0015] Compared with the prior art, the present invention has the following beneficial effects:
[0016] (1) This invention acquires negative information text from multiple channels, extracts cross-platform propagation features from the negative information text, sorts them by publication timestamp, and constructs an initial propagation trajectory sequence; then, based on this sequence, it extracts the forwarding interaction relationships between users on different platforms, constructs a cross-platform node association graph, and obtains a structured expression matrix. This scheme converts unstructured public opinion text from various channels into a unified graph structure expression form, thereby achieving effective integration from fragmented and isolated single-platform data to globally correlated data, providing basic support for the dynamic tracking of cross-platform public opinion propagation paths.
[0017] (2) This invention extracts deep topological features from the node association graph to determine a multi-level set of diffusion paths and calculates the centrality of each node to screen key nodes. Simultaneously, it obtains the dynamic evolution characteristics of key nodes within a continuous time window and extracts the change in emotional polarity to track their emotional fluctuation trajectory. This scheme reveals the hierarchical progression and structural importance of nodes in the propagation network, combined with the emotional evolution trend of nodes over time, thus achieving a leap from static single-point-of-time observation to dynamic continuous tracking. This provides an analytical basis for accurately locating the core hubs of public opinion dissemination and predicting the direction of risk evolution.
[0018] (3) This invention quantifies the confidence level of the emotional evolution characteristics of key nodes belonging to the high-risk category to obtain a critical probability value. When this value exceeds a preset threshold, it automatically generates a public opinion storm warning instruction for the key node. Subsequently, it extracts the dissemination source characteristics and cross-platform mapping relationship corresponding to the warning instruction to generate a risk prediction result for cross-platform dissemination. This solution realizes the transformation from passive response to proactive prediction by constructing a complete early warning chain from risk quantification assessment to source tracing, providing a decision-making basis for enterprises to intervene in the early stage of public opinion dissemination. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of a marketing risk early warning method based on multi-channel public opinion perception provided in the first embodiment of the present invention;
[0020] Figure 2 This is a schematic diagram of the structural representation matrix provided in the first embodiment of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Reference Figure 1 The first embodiment of the present invention provides a marketing risk early warning method based on multi-channel public opinion perception, including the following steps:
[0023] S11, acquire text data from different platforms, preprocess the text data, and generate an initial propagation trajectory sequence;
[0024] S12, extract the forwarding interaction relationship between users on different platforms based on the initial propagation trajectory sequence, determine the source node and secondary node based on the forwarding interaction relationship and establish node connections, construct a cross-platform node association graph based on each node and the node connections, and perform graph structure embedding processing on the node association graph to obtain a structure expression matrix;
[0025] S13, perform graph neural network processing on the structure expression matrix to generate deep topological features, and perform clustering and deduplication processing on the deep topological features to generate a multi-level diffusion path set;
[0026] S14, calculate the centrality value of each node according to the multi-level diffusion path set, and determine the nodes whose centrality value is greater than the preset centrality threshold as key nodes.
[0027] S15, obtain the text sequence generated by the key node, perform sentiment polarity analysis on the text sequence to obtain the sentiment polarity change, and concatenate the sentiment polarity change to obtain the evolutionary feature vector;
[0028] S16, input the evolutionary feature vector into the pre-trained random forest model, output the critical probability value, if the critical probability value is greater than the preset warning trigger threshold, generate a warning instruction for the key node, and generate a risk prediction result for cross-platform propagation based on the warning instruction.
[0029] In step S11, text data from different platforms is acquired, the text data is preprocessed, and an initial propagation trajectory sequence is generated.
[0030] The text data is preprocessed to generate an initial propagation trajectory sequence, including:
[0031] The text data is classified by sentiment polarity, and texts with negative sentiment polarity are selected as negative information texts.
[0032] Identify the core entities and corresponding publication timestamps in the negative information text, and concatenate the core entities with the publication timestamps to generate entity-related text;
[0033] Calculate the text similarity between the entity-related texts on different platforms, and determine the entity-related texts with text similarity greater than a preset similarity threshold as cross-platform propagation features;
[0034] Based on the cross-platform propagation characteristics, the entity-related texts are arranged chronologically according to the publication timestamp to generate an initial propagation trajectory sequence.
[0035] Specifically, this embodiment collects text data from different online platforms using application programming interfaces (APIs) or web crawlers provided by each platform. These platforms include Weibo, WeChat official accounts, Douyin, Kuaishou, Xiaohongshu, Toutiao, and news portals. During the collection process, each platform pushes text data conforming to the collection rules to the data collection module of this embodiment at fixed time intervals. The time interval is determined based on the push frequency of each platform's data interface, such as incremental collection every five minutes. The text data includes text content, a publishing user identifier, a publishing timestamp, and a source platform identifier. The publishing timestamp is in the format of year, month, day, hour, minute, and second, and the source platform identifier is a unique code for each platform.
[0036] After acquiring the text data, it is preprocessed. In this embodiment, the text data is classified by sentiment polarity, and texts with negative sentiment polarity are selected as negative information text. This embodiment uses SnowNLP to classify the sentiment polarity of the text data. This tool outputs a sentiment polarity score for each text, with scores ranging from 0 to 1; the smaller the score, the more negative the sentiment. A pre-set sentiment polarity threshold is determined by randomly selecting 10,000 text samples labeled with sentiment polarity from historical data of various platforms, statistically analyzing the distribution of sentiment polarity scores for negative samples, and taking the upper quartile of the negative sample score distribution as the threshold. The specific value of this threshold is 0.4. Texts with sentiment polarity scores less than this threshold are judged as negative text, marked as negative information text, and retained; texts with sentiment polarity scores greater than or equal to this threshold are judged as positive or neutral text and discarded.
[0037] After filtering out the negative information text, this embodiment uses a named entity recognition model based on conditional random fields to identify the core entities of the negative information text. The model outputs entity category labels corresponding to each character, including brand labels, product model labels, event keyword labels, and irrelevant labels. The continuous character sequence marked as brand label, product model label, or event keyword label is extracted as the core entity. At the same time, the publication timestamp is read from the metadata of the negative information text. The extracted core entity and the corresponding publication timestamp are concatenated in the format of "publication timestamp_core entity" to generate the entity-related text. The training process of this named entity recognition model is as follows: 50,000 text samples with labeled core entities are collected from historical text data of various platforms and divided into training set, validation set and test set in a ratio of 7:1:2. Character-level features are used as input and entity category labels are used as supervision signals. The L-BFGS optimization algorithm is used to solve the conditional random field model parameters. The F1 score changes are monitored on the validation set. Training stops when the F1 score on the validation set no longer increases for three consecutive rounds. The model parameters with the highest F1 score on the validation set are selected as the final model parameters.
[0038] After generating the entity-related text, this embodiment uses a term frequency-inverse document frequency (TNF) algorithm combined with cosine similarity to calculate the text similarity between entity-related texts from different platforms. For any two entity-related texts from different platforms, TNF feature vectors are constructed, and the cosine similarity between the two feature vectors is calculated as the text similarity. A similarity threshold is pre-set. This threshold is determined by collecting a set of positive samples known to have cross-platform propagation relationships and a set of negative samples known not to have cross-platform propagation relationships from historical cross-platform propagation data to form a validation dataset. Candidate thresholds are iterated within the range of 0.5 to 0.95 with a step size of 0.05. For each candidate threshold, entity-related text pairs with a similarity greater than the threshold are judged to have cross-platform propagation characteristics. The precision and recall rates between the judgment results and the annotation results of the validation dataset are statistically analyzed, and the F1 score is calculated. The candidate threshold that maximizes the F1 score is selected as the current threshold. If the text similarity is greater than the threshold, the associated text of the entity is determined to have cross-platform propagation characteristics; if the text similarity is less than or equal to the threshold, the associated text of the entity is determined not to have cross-platform propagation characteristics and is not included as a component of the subsequent trajectory sequence.
[0039] After determining the cross-platform propagation characteristics, this embodiment arranges the entity-related texts chronologically according to the publication timestamps based on the cross-platform propagation characteristics to generate an initial propagation trajectory sequence. Specifically, all entity-related texts determined to have cross-platform propagation characteristics are selected, the publication timestamp field is extracted from each entity-related text, and all entity-related texts with cross-platform propagation characteristics are arranged in ascending order according to the chronological order of the publication timestamps. The arranged entity-related texts are then connected sequentially to generate an initial propagation trajectory sequence. The initial propagation trajectory sequence retains the original text content, publishing user identifier, publication timestamp, and source platform identifier corresponding to each entity-related text to support forwarding behavior identification and node association analysis in subsequent steps.
[0040] In step S12, the forwarding interaction relationship between users on different platforms is extracted based on the initial propagation trajectory sequence. The source node and secondary node are determined based on the forwarding interaction relationship and node connections are established. A cross-platform node association graph is constructed based on each node and the node connections. The node association graph is then subjected to graph structure embedding processing to obtain a structure representation matrix.
[0041] The process includes extracting forwarding interaction relationships between users on different platforms based on the initial propagation trajectory sequence, determining the source node and secondary nodes based on the forwarding interaction relationships, and establishing node connections, including:
[0042] Identify action keywords representing forwarding behavior from the initial propagation trajectory sequence, classify and statistically analyze the action keywords, and obtain forwarding interaction relationships;
[0043] Based on the aforementioned forwarding interaction relationship, the platform user with the earliest publishing time and no upstream forwarding record is identified as the source node, and the platform user with forwarding and receiving records is identified as the secondary node.
[0044] Determine whether there is a target link or direct quote in the text data of the secondary node that points to the content published by the source node. If so, establish a directed node connection between the source node and the secondary node.
[0045] In this embodiment, the entity-related texts in the initial propagation trajectory sequence originate from text data from different platforms. Each text data includes text content and a publishing user identifier. This embodiment performs keyword matching on each text data to identify action keywords representing forwarding behavior. These action keywords include "forwarded from," "reprinted," "quoted," "reposted," "via," and "source." If the text content of a text data contains any one or more of these action keywords, the text data is marked as having forwarding behavior, and a correspondence is established between the original publishing user identifier pointed to by the action keyword and the current text data's publishing user identifier, recorded as a forwarding interaction record. If the text content of a text data does not contain any of the action keywords, it is marked as original publication. After traversing all entity-related texts in the initial propagation trajectory sequence, all forwarding interaction records are categorized and statistically analyzed. Specifically, they are grouped according to the combination of the original publishing user identifier and the forwarding user identifier, and the frequency of each combination is counted. The statistical results are summarized into the forwarding interaction relationship. The forwarding interaction relationship includes the original publishing user identifier, the forwarding user identifier, and the number of forwards for each forwarding behavior.
[0046] Specifically, each forwarding action in the forwarding interaction relationship is associated with a corresponding publication timestamp. All forwarding actions in the forwarding interaction relationship are traversed. For each forwarding interaction record, it is checked whether an upstream forwarding record exists. An upstream forwarding record refers to a record where the same original posting user ID has been forwarded by other users before the publication timestamp of this forwarding action. If no upstream forwarding record exists, the platform user corresponding to the original posting user ID in this forwarding action is determined as the source node. If an upstream forwarding record exists, the platform user corresponding to the forwarding user ID in this forwarding action is determined as a secondary node. It should be noted that if the same platform user simultaneously meets the criteria for both a source node and a secondary node—that is, the user has both published original content and been forwarded by other users, and has also forwarded content from other users—then the user is simultaneously marked as both a source node and a secondary node, and corresponding node identifiers are established for each.
[0047] After identifying the source node and the secondary node, this embodiment determines whether the text data of the secondary node contains a target link or direct quote pointing to the content published by the source node. If so, a directed node connection is established between the source node and the secondary node. Specifically, for each pair of users identified as the source node and the secondary node, the hyperlink field and text quote field in the text data of the secondary node are obtained. The Uniform Resource Locator (URL) in the hyperlink field is matched with the URL of the content published by the source node. If the match is successful, it is determined that a target link pointing to the content published by the source node exists. When the content published by the source node does not have a URL, the combination of the core entity of the content published by the source node and its publication timestamp is used as a substitute matching basis. This is compared with the text similarity description in the text data of the secondary node. If the similarity is greater than a preset threshold, it is determined that a target link exists.
[0048] If hyperlink matching fails, further text matching is performed on the text citation field, which includes direct and indirect citations. Direct citations include text snippets enclosed in quotation marks, citations prefixed with "forwarded:" or "via," and citations in the platform's native forwarding format that use the "@" symbol to point to the original publisher. The content of these citations is consistent with the content published by the original node. Indirect citations refer to text snippets beginning with words like "as," "for example," or "as shown," citing content published by the original node, as well as content containing the core entities of the original text and expressed in paraphrased form. If the text data of the secondary node contains either a direct or indirect citation, it is determined that a direct citation pointing to the content published by the original node exists. If a target link or a direct citation exists, a directed node connection is established between the original node and the secondary node. The direction of the directed node connection is from the original node to the secondary node, i.e., from the original publisher of the forwarded content to the user who forwarded the content.
[0049] Specifically, a cross-platform node association graph is constructed based on each node and the connections between them. The node association graph is then subjected to graph structure embedding processing to obtain a structure representation matrix, including:
[0050] The interaction frequency is obtained by counting the number of times each node connection appears within a preset time window.
[0051] Based on the nodes, the connections between the nodes, and the interaction frequency, a cross-platform node association graph is constructed.
[0052] The node association graph is subjected to node embedding processing to obtain a topological feature vector;
[0053] An adjacency matrix is constructed based on the topological feature vectors, and the adjacency matrix is then subjected to dimensionality reduction to obtain the structural representation matrix.
[0054] After establishing the node connections, this embodiment constructs a cross-platform node association graph based on the nodes and their connections. Nodes include all platform users identified as source or secondary nodes. This embodiment uses each node as a vertex and the node connections as directed edges to construct a directed graph structure. The direction and platform origin of each node connection are statistically analyzed, and each node connection is associated with a corresponding source platform identifier. After integrating all nodes and node connections, the cross-platform node association graph is formed.
[0055] Subsequently, graph structure embedding processing is performed on the node association graph to obtain the structure representation matrix. In this embodiment, the original adjacency matrix of the node association graph is extracted, and principal component analysis (PCA) is used to reduce its dimensionality. Principal components are selected based on a cumulative variance contribution rate of 95%, resulting in the dimensionality-reduced structure representation matrix. It should be noted that 95% is a commonly used threshold in this field for preserving data integrity during PCA, and this embodiment adopts this general standard. In practical applications, if the dimensionality reduction exceeds 64 dimensions, the actual dimension corresponding to 95% should be used, or an adaptive adjustment can be made by setting the upper limit of the dimensionality reduction to 64 dimensions while ensuring an information retention rate of no less than 95%.
[0056] like Figure 2 As shown in the figure, this diagram presents a visualization of the structural representation matrix. The vertical axis represents user nodes in the cross-platform node association graph, and the horizontal axis represents the five principal component dimensions (PC1 to PC5) extracted after dimensionality reduction of the adjacency matrix. The color intensity of each cell indicates the magnitude of the feature value of the corresponding node in the corresponding principal component dimension, with the color gradually changing from blue (negative values) to red (positive values). This diagram intuitively demonstrates the differentiated feature distribution patterns of each user node in the principal component space after graph structure embedding and dimensionality reduction. Different nodes have different feature values in different principal component dimensions, thereby compressing the high-dimensional topological associations in the original graph into a low-dimensional and compact structural representation. This provides a numerical structural input basis for subsequent steps such as graph neural network processing based on this matrix, generating deep topological features, and clustering deduplication.
[0057] In step S13, deep topological features are generated by graph neural network processing based on the structure representation matrix. These deep topological features are then clustered and deduplicated to generate a multi-level diffusion path set, including:
[0058] The structure representation matrix is subjected to graph convolution to obtain the hidden state vector of each node;
[0059] Based on the hidden state vector, connected subgraphs are identified, the propagation depth of each connected subgraph is extracted, and deep topological features are generated.
[0060] Extract the branch evolution trajectory from the connected subgraph where the propagation depth is greater than a preset depth threshold;
[0061] Clustering the branch evolution trajectories yields multi-level diffusion paths;
[0062] The multi-level diffusion paths are deduplicated to obtain a set of multi-level diffusion paths.
[0063] This embodiment performs graph convolution processing on the structure representation matrix to obtain the hidden state vectors of each node. The structure representation matrix has a dimension of 64. Each row vector in the structure representation matrix is used as the initial feature of each node. The structure representation matrix and the adjacency matrix of the node association graph are input together into the graph convolutional network. The graph convolutional network contains two graph convolutional layers: the first layer has an input dimension of 64 and an output dimension of 128, and the second layer has an input dimension of 128 and an output dimension of 32. The first and second layers are followed by a ReLU activation function. After processing by the graph convolutional network, the hidden state vectors of each node are output, and these hidden state vectors have a dimension of 32.
[0064] It should be noted that the training process of this graph convolutional network is as follows: five thousand graph samples with labeled node classification labels are collected from historical propagation data and divided into training set, validation set and test set in a 7:1:2 ratio. The cross-entropy loss function is used as the supervision signal, and the Adam optimization algorithm is used for iterative training on the training set. The classification accuracy is monitored on the validation set. Training stops when the accuracy of the validation set no longer increases for ten consecutive rounds. The model parameters with the highest accuracy on the validation set are selected as the final parameters.
[0065] After obtaining the hidden state vectors of each node, this embodiment identifies connected subgraphs based on the hidden state vectors, extracts the propagation depth of each connected subgraph, and generates deep topological features. Specifically, the cosine similarity between nodes is calculated based on the hidden state vectors. Node pairs with a cosine similarity greater than a preset similarity threshold are considered to have potential connections, and a weighted undirected graph is constructed. A depth-first search algorithm is used to traverse all nodes in the weighted undirected graph, dividing interconnected nodes into the same connected component, with each connected component serving as a connected subgraph. The preset similarity threshold is determined by randomly selecting a sample set containing node pairs from historical propagation data, calculating the cosine similarity of each node pair and marking whether there is an actual propagation relationship, traversing candidate thresholds in the range of 0.5 to 0.95 with a step size of 0.05, and selecting the threshold that maximizes the F1 score as the threshold. For each connected subgraph, the maximum value of the shortest path length between any two nodes in the connected subgraph is taken as the propagation depth of the connected subgraph, i.e., the propagation depth is equal to the diameter of the connected subgraph. The number of nodes, edge density, and propagation depth of each connected subgraph are extracted, and the extraction results are summarized into the deep topological features.
[0066] After extracting the propagation depth of each connected subgraph, this embodiment extracts branch evolution trajectories from the connected subgraphs whose propagation depth is greater than a preset depth threshold. The depth threshold is determined by collecting a sample set containing complete propagation chains from historical propagation data. On the validation set, candidate thresholds are traversed in the range of 3 to 10 with a step size of 1. For each candidate threshold, the proportion of connected subgraphs with a propagation depth greater than that threshold from which effective branch evolution trajectories can be extracted in subsequent steps is statistically analyzed. A comprehensive score is calculated based on the number of nodes covered by the trajectory, and the candidate threshold with the highest comprehensive score is selected as the preset depth threshold. In this specific implementation, the threshold is set to 5. From each connected subgraph with a propagation depth greater than this threshold, a depth-first search algorithm is used to traverse all propagation paths within the connected subgraph along the directed edges. Each complete path from the starting node to the ending node is taken as a branch evolution trajectory, and the branch evolution trajectory corresponding to each connected subgraph is output.
[0067] After obtaining the branch evolution trajectory, this embodiment performs clustering processing on the branch evolution trajectory to obtain a multi-level diffusion path. This embodiment uses the K-means clustering algorithm to cluster the branch evolution trajectory. The node identifier sequence of each branch evolution trajectory is used as the original feature, and the longest common subsequence distance is used as the similarity measure between two trajectories. The K value is determined by the elbow rule. Candidate K values are traversed within the integer range of 2 to 15. K-means clustering is performed on each candidate K value, and the sum of squared distances from all sample points to the center of their respective clusters is calculated. A curve showing the sum of squared distances changing with the K value is plotted, and the slope change of the curve at each candidate K value is calculated. The K value corresponding to the inflection point of the curve is selected as the preset number of clusters. All the branch evolution trajectories are input into the K-means clustering algorithm, and after iterative calculation, they are divided into multiple clusters, each cluster corresponding to a multi-level diffusion path.
[0068] Finally, this embodiment employs a hash comparison algorithm to deduplicate the multi-level diffusion paths. For each multi-level diffusion path, its node identifier sequence is concatenated into a string in sequence, and the SHA-256 hash function is used to perform hash calculation on this string to generate a fixed-length hash value. If the hash value of a multi-level diffusion path already exists, the path is determined to be a redundant path and is removed; otherwise, the path is retained. After traversing all multi-level diffusion paths, all retained paths are summarized into the multi-level diffusion path set.
[0069] In step S14, the centrality value of each node is calculated based on the multi-level diffusion path set, and the nodes with the centrality value greater than the preset centrality threshold are identified as key nodes.
[0070] This embodiment extracts the connection relationships between nodes in the multi-level diffusion path set and constructs an initial adjacency matrix. The multi-level diffusion path set contains multiple diffusion paths, each consisting of several nodes and directed edges between them. All directed edges in the multi-level diffusion path set are traversed, and the number of times a direct connection exists between any two nodes is counted. A square matrix is constructed with the total number of nodes as the dimension. The element in the i-th row and j-th column of the square matrix indicates whether a direct connection exists between the i-th node and the j-th node; if it exists, it is marked as 1, and if it does not exist, it is marked as 0. This constructed square matrix is used as the initial adjacency matrix.
[0071] Singular value decomposition (SVD) is performed on the initial adjacency matrix to obtain structural feature vectors. This embodiment uses SVD to decompose the initial adjacency matrix. The initial adjacency matrix is a square matrix of dimension N x N, where N is the total number of nodes. After SVD, a left singular matrix, a singular value diagonal matrix, and a right singular matrix are obtained. The matrix containing the principal singular values is selected, and the left singular vectors corresponding to the first 16 largest singular values are extracted as the principal feature components. The extracted left singular vectors are concatenated column-wise to form a matrix of dimension N x 16. The mean feature value of each node is calculated along the row direction of the matrix to obtain a structural feature vector of dimension 16. The structural feature vector is used to characterize the structural position of each node in the entire propagation network. The 16 is obtained by traversing the number of candidate feature dimensions in the historical propagation network sample set with a step size of 2 in the range of 4 to 32. For each candidate dimension, the extracted singular vectors are concatenated column by column to calculate the silhouette coefficient of each node in the dimension reduction space. The number that maximizes the silhouette coefficient is selected as this number. The silhouette coefficient is determined by calculating the average distance between each node and other nodes in its own cluster and the average distance with the nearest other cluster. The value ranges from -1 to +1. The larger the value, the better the clustering effect of the node in the dimension reduction space. The above structural feature vector is used to characterize the structural position of each node in the entire propagation network and serves as an auxiliary feature for subsequent node influence analysis. It does not participate in the calculation of degree distribution values.
[0072] Calculate the degree distribution value of each node; by traversing the initial adjacency matrix, count the number of valid edges directly connected to each node, and use the statistical result as the degree distribution value of each node, wherein the degree distribution value is an integer greater than or equal to 0; the degree distribution value is directly obtained from the sum of the rows of the initial adjacency matrix, that is, the degree distribution value of the i-th node is equal to the sum of all elements in the i-th row of the initial adjacency matrix.
[0073] The degree distribution value of each node is multiplied by a preset weight allocation ratio to obtain the centrality value of each node. In this embodiment, the weight allocation ratio is preset. The method for determining the weight allocation ratio is as follows: a set of propagation network samples containing labeled key nodes is extracted from historical propagation data. Candidate weight values are traversed in the range of 0.5 to 5.0 with a step size of 0.1. For each candidate weight value, the degree distribution value of each node is multiplied by the weight value to obtain the centrality value. The top M nodes are selected as candidate key nodes after sorting by centrality value, where M is 20% of the total number of nodes N, rounded up. The F1 score between the candidate key nodes and the labeled key nodes is calculated, and the weight value that maximizes the F1 score is selected as the preset weight allocation ratio. In a specific implementation, the weight allocation ratio is set to 2.
[0074] The centrality values of each node are sorted in descending order. The number of key nodes K is determined according to a preset percentage threshold, where K is equal to the total number of nodes multiplied by the percentage threshold and rounded up. The nodes ranked in the top K positions in the sort are selected as key nodes. The remaining nodes are marked as non-key nodes and are not processed in subsequent steps. The percentage threshold is determined by iterating through candidate percentage values in the historical propagation network sample set in a step size of 1% within the range of 5% to 20%. For each candidate value, the nodes whose centrality values rank in that percentage are marked as key nodes. The F1 score between the marking result and the annotation result is calculated, and the percentage value that maximizes the F1 score is selected as the preset percentage threshold. In a specific implementation, the percentage threshold is set to 10%.
[0075] In step S15, the text sequence generated by the key node is obtained, the emotional polarity analysis of the text sequence is performed to obtain the emotional polarity change, and the emotional polarity change is concatenated to obtain the evolutionary feature vector.
[0076] Specifically, sentiment polarity analysis is performed on the text sequence to obtain a change in sentiment polarity. These changes in sentiment polarity are then concatenated to obtain an evolutionary feature vector, which includes:
[0077] The text sequence is subjected to term frequency-inverse document frequency feature extraction to obtain a semantic feature set;
[0078] The semantic feature set is matched with a preset sentiment dictionary to obtain a polarity score sequence;
[0079] Perform a difference operation on the polarity score sequence to obtain the changing gradient value;
[0080] Extract the fluctuation amplitude data where the absolute value of the changing gradient value is greater than a preset fluctuation threshold, and construct dynamic evolution features;
[0081] The change in emotional polarity is extracted from the dynamic evolutionary features and concatenated in chronological order to obtain the evolutionary feature vector.
[0082] First, this embodiment acquires the text sequence generated by the key node within a preset continuous time window. Specifically, taking the time point when the key node is determined as the starting time, a preset duration is extracted as the continuous time window. The preset duration is determined by collecting a sample of the time interval from the key node's first statement to the peak of public opinion from historical public opinion event data, and taking the 75th percentile of the sample as the preset duration; in a specific implementation, the preset duration is 72 hours. The preset duration is divided into multiple sub-windows, and the time step of each sub-window is 12 hours, that is, the duration of each sub-window is 12 hours, thus dividing the 72-hour continuous time window into 6 sub-windows. For each sub-window, all text data published by the key node within that sub-window is collected to form the text sequence corresponding to that sub-window. If the key node does not publish any text within a certain sub-window, the text sequence of that sub-window is marked as empty.
[0083] After obtaining the text sequence corresponding to each sub-window, this embodiment performs term frequency-inverse document frequency (IF-IVF) feature extraction on each text sequence. For the text sequence of each sub-window, word segmentation is first performed to obtain a word sequence; the frequency of each word appearing in the sub-window is counted and divided by the total number of words in the sub-window to obtain the term frequency value; the number of sub-windows in which each word appears in the text sequences of all sub-windows is counted, and the logarithm (base 10) of the total number of sub-windows divided by the number of sub-windows containing the word is taken to obtain the IVF value; the term frequency value and the IVF value are multiplied to obtain the term frequency-inverse document frequency weight value of each word. The weight values of all words are arranged in lexicographical order to obtain the semantic feature vector corresponding to the sub-window. After traversing all 6 sub-windows, the semantic feature vectors of each sub-window are summarized into the semantic feature set, which contains 6 semantic feature vectors.
[0084] After obtaining the semantic feature set, this embodiment matches each semantic feature vector in the semantic feature set with a preset sentiment dictionary to obtain the polarity score corresponding to each sub-window. The sentiment dictionary is a pre-constructed dictionary containing words and their sentiment polarity scores. The sentiment polarity scores range from -1 to +1, with negative values indicating negative sentiment and positive values indicating positive sentiment. The sentiment polarity scoring system used in this step is an independent scoring standard from the SnowNLP scoring system in step S11, and each maintains logical consistency within this step. The negative judgment in step S11 is based on the score of 0 to 1 output by SnowNLP that is less than the threshold, while the negative judgment in this step is based on the negative value of the sentiment dictionary's score of -1 to +1. The construction of this sentiment dictionary involves collecting 50,000 text samples with labeled sentiment polarity scores from historical text data of various platforms. After segmenting each sample, the frequency of occurrence of each word under different sentiment polarities is counted, and the sentiment polarity score of each word is calculated, with a value range from -1 to +1. The words and their scores are stored in the dictionary in the form of key-value pairs, and the score of words not included in the dictionary is recorded as zero. For each sub-window's semantic feature vector, each word it contains is matched against the sentiment dictionary. If a match is found, the corresponding sentiment polarity score is taken; otherwise, the score is recorded as zero. The sentiment polarity scores of all words are weighted and summed, then divided by the sum of the weighted values to obtain the polarity score of that sub-window. This polarity score is a floating-point number between -1 and +1. After traversing all sub-windows, the polarity scores of each sub-window are arranged in chronological order to obtain a polarity score sequence.
[0085] After obtaining the polarity score sequence, this embodiment performs a difference operation on the polarity score sequence to obtain the change gradient values between each adjacent sub-window. Specifically, the polarity score sequence contains 6 polarity score values arranged in chronological order. For two adjacent sub-windows, the difference between the polarity score value of the later sub-window and the polarity score value of the earlier sub-window is the change gradient value corresponding to that adjacent interval. By traversing all pairs of adjacent sub-windows in the above manner, a total of 5 change gradient values are obtained.
[0086] After obtaining the change gradient values corresponding to each adjacent sub-window, this embodiment extracts the fluctuation amplitude data where the absolute value of the change gradient value is greater than a preset fluctuation threshold, and constructs dynamic evolution features. The preset fluctuation threshold is determined by collecting polarity score change sequence samples of key nodes in public opinion events from historical public opinion data, calculating the absolute value of the change gradient at adjacent time points for each sample, and taking the 75th percentile of all change gradient absolute value data as candidate thresholds; on the validation set, the candidate thresholds are traversed in the interval from 0.1 to 1.5 with a step size of 0.05. For each candidate threshold, the interval where the absolute value of the change gradient is greater than the threshold is marked as a significant fluctuation interval. The F1 score is calculated based on whether a public opinion risk escalation event occurs within the significant fluctuation interval as the evaluation criterion, and the threshold that maximizes the F1 score is selected as the preset fluctuation threshold; in specific implementation, the fluctuation threshold is set to 0.5. The absolute value of each gradient value is compared with the fluctuation threshold. If the absolute value of a gradient value is greater than the fluctuation threshold, the fluctuation amplitude data corresponding to that gradient value is extracted. The fluctuation amplitude data includes the gradient value itself and its corresponding time interval information. If the absolute value of a gradient value is less than or equal to the fluctuation threshold, that gradient value is discarded and not included in the dynamic evolution feature. All extracted fluctuation amplitude data are arranged in chronological order to construct the dynamic evolution feature.
[0087] After obtaining the dynamic evolution features, this embodiment extracts the emotional polarity change corresponding to each time window from the dynamic evolution features and concatenates them in chronological order to obtain an evolution feature vector. The dynamic evolution features include the extracted fluctuation amplitude data and their corresponding time interval information. For each fluctuation amplitude data in the dynamic evolution features, its change gradient value is used as the emotional polarity change of that time window. The emotional polarity changes of each time window are arranged sequentially according to time order to form a one-dimensional vector, the dimension of which is equal to the number of time windows judged to have significant fluctuations. The emotional polarity change corresponding to the time windows where no significant fluctuations occur is recorded as zero, thus obtaining a numerical sequence with a fixed length of 5 (the number of adjacent sub-window pairs), which is used as the evolution feature vector. The evolution feature vector has a dimension of 5.
[0088] In step S16, the evolutionary feature vector is input into a pre-trained random forest model, and the critical probability value is output. If the critical probability value is greater than the preset warning trigger threshold, a warning instruction is generated for the key node, and a risk prediction result for cross-platform propagation is generated based on the warning instruction.
[0089] The evolutionary feature vector is input into a pre-trained random forest model, and the output of the critical probability value includes:
[0090] The evolutionary feature vector is input into a pre-trained random forest model, and feature splitting is performed in each decision tree of the random forest model to obtain the set of leaf node weights.
[0091] The Gini impurity value of each branch in each decision tree is calculated based on the set of leaf node weights, and decision tree branches with Gini impurity values less than a preset impurity threshold are selected.
[0092] Extract the classification voting matrix corresponding to the selected decision tree branches, and perform aggregation operation on the classification voting matrix to obtain the confidence distribution sequence;
[0093] Logistic regression mapping is performed on the confidence distribution sequence to obtain the critical probability value.
[0094] In this embodiment, the evolutionary feature vector is input into a pre-trained random forest model, and the output is a critical probability value. The training process of the random forest model is as follows: sample data is collected from a historical public opinion event database. The historical public opinion event database records the evolutionary feature vectors of key nodes in each historical public opinion event and the labeling results of whether the event eventually evolved into a public opinion storm. Specifically, for each historical public opinion event, the evolutionary feature vectors of key nodes in the event are extracted in the same way as in the previous steps S11 to S15. The dimension of the evolutionary feature vectors is 5-dimensional. At the same time, whether the event is eventually manually labeled as a public opinion storm event is used as a label. If the event eventually triggers public opinion attention across the entire network and lasts for more than 48 hours, the label is 1; otherwise, it is 0. A total of 5,000 samples are collected and divided into training set, validation set and test set in a ratio of 7:1:2. The training set contains 3,500 samples, the validation set contains 500 samples, and the test set contains 1,000 samples. The random forest model contains 100 decision trees. The classification accuracy is monitored on the validation set, and the model parameters with the highest F1 score on the validation set are selected as the final model parameters. The model is put into use when the classification accuracy on the test set reaches more than 85%.
[0095] In this embodiment, the evolutionary feature vector is input into the random forest model. The random forest model receives an evolutionary feature vector with a dimension of 5 as input, and performs classification reasoning independently through 100 decision trees. Each decision tree outputs a high-risk or low-risk category prediction result. The prediction results of all decision trees are aggregated, and the proportion of decision trees predicting the high-risk category is used as the critical probability value. The critical probability refers to the probability that a public opinion event exceeds a preset risk threshold and evolves into a public opinion storm event, and its value ranges from 0 to 1.
[0096] It is worth noting that the Gini impurity threshold is a preset impurity threshold selected during the training of the random forest model, based on the distribution of Gini impurity values of each decision tree branch and the selection of the branch with the largest F1 score on the validation set as the screening effect. In practice, this impurity threshold is set to 0.1.
[0097] The process of generating a risk prediction result for cross-platform propagation based on the aforementioned warning instruction includes:
[0098] Based on the warning instructions, the characteristics of the propagation source are extracted, and platform accounts are matched based on the characteristics of the propagation source to obtain cross-platform mapping relationships;
[0099] A propagation evolution trajectory is generated based on the cross-platform mapping relationship and the publication timestamp;
[0100] Calculate the risk spillover probability based on the propagation evolution trajectory;
[0101] The risk spillover probability and the propagation source characteristics are used to construct a prediction feature vector. The prediction feature vector is then classified to generate a risk prediction result for cross-platform propagation.
[0102] After obtaining the critical probability value, this embodiment compares the critical probability value with a preset warning trigger threshold. The warning trigger threshold is determined by iterating through candidate thresholds in the range of 0.5 to 0.95 with a step size of 0.05 on the validation set. For each candidate threshold, samples with a critical probability value greater than the threshold are judged as triggering a warning. The precision and recall rates between the judgment results and the actual labels are statistically analyzed, and the F1 score is calculated. The candidate threshold that maximizes the F1 score is selected as the preset warning trigger threshold. In a specific implementation, the warning trigger threshold is set to 0.8. If the critical probability value is greater than the preset warning trigger threshold, it is determined that there is a risk of a public opinion storm, and the warning instruction generation process is triggered. If the critical probability value is less than or equal to the preset warning trigger threshold, it is determined that there is no significant risk of a public opinion storm, and the subsequent warning process is not triggered.
[0103] When the critical probability value exceeds the preset warning trigger threshold, this embodiment generates a warning instruction for the key node. Specifically, the node identifier of the key node and the critical probability value are used as the instruction payload, encapsulated according to a preset instruction format to generate the warning instruction. The warning instruction format includes a key node identifier code, a critical probability value, and a trigger timestamp field, with each field connected by a delimiter to form a structured string. After obtaining the warning instruction, this embodiment extracts the propagation source features based on the warning instruction. The node identifier code of the key node is parsed from the warning instruction, and the corresponding propagation source features are read from the propagation source information database based on the node identifier code. The propagation source features include account registration duration, average historical post views, number of followers, and account authentication type. The propagation source information database collects the registration duration, average historical post views, number of followers, and account authentication type of authenticated accounts through public interfaces of various platforms, stores them with the account identifier as the key, and updates them incrementally on a daily basis.
[0104] After extracting the source features of the dissemination, this embodiment performs platform account matching based on the source features to obtain cross-platform mapping relationships. Specifically, the account nickname and authentication entity information in the source features of the dissemination are used as the matching basis, and compared with a preset multi-platform account database. The multi-platform account database contains the nicknames, authentication entities, and platform identification information of verified accounts on various mainstream platforms. By collecting the nicknames, authentication entity names, platform identifications, account registration durations, and number of followers of verified enterprise accounts and personal accounts from the public interfaces or official authentication pages of various mainstream platforms (including Weibo, WeChat Official Accounts, Douyin, Kuaishou, Xiaohongshu, and Toutiao), the system performs cross-platform mapping. For the same authentication entity (based on the enterprise business license registration number or personal account registration number), the system performs cross-platform mapping. For accounts opened on multiple platforms using real-name authentication information as a unique identifier, the account information from each platform is linked and stored as a single cross-platform mapping record, using the authentication entity as the association key. For accounts where official authentication information cannot be obtained, supplementary data is collected from the cross-platform account information declared by the account in the search results pages of each platform. The association relationship is confirmed after fuzzy string matching (edit distance less than or equal to 2). All data is stored in the database after deduplication and is updated incrementally daily. Newly collected account information is compared with existing records. If the authentication entity identifier matches, the records are merged and updated; otherwise, they are inserted as new records. If accounts corresponding to the same authentication entity are matched across multiple platforms, a mapping relationship is established between these accounts to obtain the cross-platform mapping relationship.
[0105] After obtaining the cross-platform mapping relationship, this embodiment generates a propagation evolution trajectory based on the cross-platform mapping relationship and the release time information corresponding to the warning instruction. Specifically, it obtains the text data related to the same core entity released by each platform account before and after the warning instruction trigger time in the cross-platform mapping relationship, along with their release timestamps, and arranges the propagation events of each platform in chronological order of release time to form the propagation evolution trajectory.
[0106] After obtaining the propagation trajectory, this embodiment calculates the risk spillover probability based on the propagation trajectory. This embodiment uses the Cox proportional hazards model to calculate the risk spillover probability. The training process of the Cox proportional hazards model is as follows: event samples containing complete propagation trajectories are collected from a historical public opinion event database. The order of appearance and time intervals on each platform are used as input features, and whether the event eventually spreads to mainstream authoritative news media is used as the event label. Partial likelihood estimation is used to fit the model parameters. The training sample size of the Cox proportional hazards model is two thousand historical public opinion event samples containing complete propagation trajectories. The current propagation trajectory is input into the Cox proportional hazards model, and the risk spillover probability is output.
[0107] The risk spillover probability and the propagation source characteristics are combined to construct a predictive feature vector. In this embodiment, the risk spillover probability is used as one dimension, and the account registration time, average historical post views, number of followers, and account authentication type in the propagation source characteristics are encoded into numerical vectors. The risk spillover probability is then appended to the end of the numerical vectors to form the predictive feature vector.
[0108] The predicted feature vectors are classified to generate risk prediction results for cross-platform propagation. In this embodiment, a support vector machine (SVM) classifier is used to classify the predicted feature vectors. The SVM classifier is pre-trained on a historical public opinion event sample set, using the predicted feature vectors of historical public opinion events as input and the final risk level label of the event as the supervision signal, trained using a radial basis function kernel. The training sample size for the SVM classifier is three thousand historical public opinion event samples, divided into training, validation, and test sets in a 7:1:2 ratio. The predicted feature vectors are input into the SVM classifier, which outputs the corresponding risk level label. This risk level label is used as the risk prediction result for cross-platform propagation.
[0109] In summary, this invention tracks the path of public opinion dissemination by constructing a cross-platform node association graph, combines key node identification and sentiment evolution trend analysis to conduct risk quantification assessment, and automatically triggers early warning instructions and traces the source of dissemination when a preset threshold is exceeded, thus realizing dynamic tracking and proactive prediction of cross-platform marketing risks.
[0110] The second embodiment of the present invention provides a marketing risk early warning system based on multi-channel public opinion perception, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the method described above.
[0111] It should be noted that the marketing risk early warning system based on multi-channel public opinion perception provided in this embodiment of the invention is used to execute all the process steps of the marketing risk early warning method based on multi-channel public opinion perception in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.
[0112] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0113] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A marketing risk early warning method based on multi-channel public opinion perception, characterized in that, include: Text data from different platforms is acquired, the text data is preprocessed, and an initial propagation trajectory sequence is generated. Based on the initial propagation trajectory sequence, the forwarding interaction relationship between users on different platforms is extracted. Based on the forwarding interaction relationship, the source node and secondary node are determined and node connections are established. Based on each node and the node connection, a cross-platform node association graph is constructed. The node association graph is subjected to graph structure embedding processing to obtain a structure expression matrix. Based on the structural representation matrix, a graph neural network is used to generate deep topological features. The deep topological features are then clustered and deduplicated to generate a multi-level diffusion path set. The centrality value of each node is calculated based on the multi-level diffusion path set, and the nodes whose centrality value is greater than the preset centrality threshold are identified as key nodes. Obtain the text sequence generated by the key node, perform sentiment polarity analysis on the text sequence to obtain the sentiment polarity change, and concatenate the sentiment polarity change to obtain the evolutionary feature vector; The evolutionary feature vector is input into a pre-trained random forest model, which outputs a critical probability value. If the critical probability value is greater than a preset warning trigger threshold, a warning instruction is generated for the key node. Based on the warning instruction, a risk prediction result for cross-platform propagation is generated.
2. The marketing risk early warning method based on multi-channel public opinion perception as described in claim 1, characterized in that, The preprocessing of the text data to generate an initial propagation trajectory sequence includes: The text data is classified by sentiment polarity, and texts with negative sentiment polarity are selected as negative information texts. Identify the core entities and corresponding publication timestamps in the negative information text, and concatenate the core entities with the publication timestamps to generate entity-related text; Calculate the text similarity between the entity-related texts on different platforms, and determine the entity-related texts with text similarity greater than a preset similarity threshold as cross-platform propagation features; Based on the cross-platform propagation characteristics, the entity-related texts are arranged chronologically according to the publication timestamp to generate an initial propagation trajectory sequence.
3. The marketing risk early warning method based on multi-channel public opinion perception as described in claim 1, characterized in that, The step of extracting the forwarding interaction relationship between users on different platforms based on the initial propagation trajectory sequence, and determining the source node and secondary node based on the forwarding interaction relationship and establishing node connections includes: Identify action keywords representing forwarding behavior from the initial propagation trajectory sequence, classify and statistically analyze the action keywords, and obtain forwarding interaction relationships; Based on the aforementioned forwarding interaction relationship, the platform user with the earliest publishing time and no upstream forwarding record is identified as the source node, and the platform user with forwarding and receiving records is identified as the secondary node. Determine whether there is a target link or direct quote in the text data of the secondary node that points to the content published by the source node. If so, establish a directed node connection between the source node and the secondary node.
4. The marketing risk early warning method based on multi-channel public opinion perception as described in claim 1, characterized in that, The process involves constructing a cross-platform node association graph based on each node and its connections, and then performing graph structure embedding processing on the node association graph to obtain a structure representation matrix, including: The interaction frequency is obtained by counting the number of times each node connection appears within a preset time window. Based on the nodes, the connections between the nodes, and the interaction frequency, a cross-platform node association graph is constructed. The node association graph is subjected to node embedding processing to obtain a topological feature vector; An adjacency matrix is constructed based on the topological feature vectors, and the adjacency matrix is then subjected to dimensionality reduction to obtain the structural representation matrix.
5. The marketing risk early warning method based on multi-channel public opinion perception as described in claim 1, characterized in that, The step of generating deep topological features by performing graph neural network processing based on the structure representation matrix, and then performing clustering and deduplication processing on the deep topological features to generate a multi-level diffusion path set includes: The structure representation matrix is subjected to graph convolution to obtain the hidden state vector of each node; Based on the hidden state vector, connected subgraphs are identified, the propagation depth of each connected subgraph is extracted, and deep topological features are generated. Extract the branch evolution trajectory from the connected subgraph where the propagation depth is greater than a preset depth threshold; Clustering the branch evolution trajectories yields multi-level diffusion paths; The multi-level diffusion paths are deduplicated to obtain a set of multi-level diffusion paths.
6. The marketing risk early warning method based on multi-channel public opinion perception according to claim 1, characterized in that, The emotional polarity analysis of the text sequence yields the change in emotional polarity, and the resulting evolutionary feature vector is obtained by concatenating these changes in emotional polarity. This includes: The text sequence is subjected to term frequency-inverse document frequency feature extraction to obtain a semantic feature set; The semantic feature set is matched with a preset sentiment dictionary to obtain a polarity score sequence; Perform a difference operation on the polarity score sequence to obtain the changing gradient value; Extract the fluctuation amplitude data where the absolute value of the changing gradient value is greater than a preset fluctuation threshold, and construct dynamic evolution features; The change in emotional polarity is extracted from the dynamic evolutionary features and concatenated in chronological order to obtain the evolutionary feature vector.
7. The marketing risk early warning method based on multi-channel public opinion perception according to claim 1, characterized in that, The step of inputting the evolutionary feature vector into a pre-trained random forest model and outputting a critical probability value includes: The evolutionary feature vector is input into a pre-trained random forest model, and feature splitting is performed in each decision tree of the random forest model to obtain the set of leaf node weights. The Gini impurity value of each branch in each decision tree is calculated based on the set of leaf node weights, and decision tree branches with Gini impurity values less than a preset impurity threshold are selected. Extract the classification voting matrix corresponding to the selected decision tree branches, and perform aggregation operation on the classification voting matrix to obtain the confidence distribution sequence; Logistic regression mapping is performed on the confidence distribution sequence to obtain the critical probability value.
8. The marketing risk early warning method based on multi-channel public opinion perception according to claim 2, characterized in that, The step of generating a risk prediction result for cross-platform propagation based on the early warning instruction includes: Based on the warning instructions, the characteristics of the propagation source are extracted, and platform accounts are matched based on the characteristics of the propagation source to obtain cross-platform mapping relationships; A propagation evolution trajectory is generated based on the cross-platform mapping relationship and the publication timestamp; Calculate the risk spillover probability based on the propagation evolution trajectory; The risk spillover probability and the propagation source characteristics are used to construct a prediction feature vector. The prediction feature vector is then classified to generate a risk prediction result for cross-platform propagation.
9. A marketing risk early warning system based on multi-channel public opinion perception, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the method described in any one of claims 1 to 8.