An information popularity prediction method and system based on trend factors and mutation factors
By introducing trend factors and mutation factors into information popularity prediction, and using emotion classification model and Transformer architecture for feature extraction and integration, the problems of insufficient factor consideration and limited prediction accuracy in the existing technology are solved, and more efficient and accurate information popularity prediction is achieved.
Patent Information
- Application Number
- CN202411071031.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-08-06
AI Technical Summary
The prior art considers few factors in information popularity prediction, has low efficiency and limited prediction accuracy, especially in emergencies or hot topics, and it is difficult to quickly and accurately predict the trend of information popularity change.
The information heat prediction method based on trend factors and mutation factors is adopted, historical information data is obtained through the emotional classification model, trend factors and mutation factors are extracted, and the Transformer architecture and local information feature extraction model are used to perform feature extraction and integration to obtain the final prediction results.
The efficiency and accuracy of information popularity prediction are improved, and the trend of information popularity can be reflected more comprehensively, especially in emergencies or hot topics, which can quickly and accurately predict the changes in information popularity.
Smart Images

Figure CN118861300B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis, and in particular, to a method and system for predicting the information popularity of trend factors and mutation factors. Background Art
[0002] In the era of information explosion, information popularity, as an important indicator to measure the degree of attention of a certain event, topic or specific object in the cyber space, has inestimable value for government decision-making, corporate public relations, marketing, and the prediction and management of social hot events; after the outbreak of an information event, the real-time monitoring and accurate prediction of information popularity can provide decision-making support for quickly evaluating the severity of the situation, timely intervention and effective response.
[0003] Currently, for the prediction of information popularity, a variety of technologies and methods have been proposed and applied, such as keyword retrieval based on search engines, information monitoring and analysis systems based on big data, sentiment analysis based on natural language processing (NLP), etc.; however, these methods still have some deficiencies in practical applications.
[0004] In the current data analysis process, fewer factors are considered. Although the number of forwards, comments, likes, reads, entity numbers, and collections of related topics are considered, these factors can reflect the popularity trend of information to a certain extent, but it is still not comprehensive enough. For example, the emotional trend of the public is ignored. In addition, there are situations of low efficiency and limited prediction accuracy in the data processing process. In particular, when considering the information popularity propagation law, in the face of emergencies or hot topics, how to quickly and accurately predict the change trend of information popularity is still an urgent problem to be solved currently. Summary of the Invention
[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a method for predicting information popularity based on trend factors and mutation factors, which solves the deficiencies existing in the prior art.
[0006] The purpose of the present invention is achieved through the following technical solutions: A method for predicting information popularity based on trend factors and mutation factors, the prediction method includes:
[0007] Step 1: Obtain the historical information data of a certain topic through a sentiment classification model;
[0008] Step 2: For each topic data, use the first feature data within a unit time obtained as the feature dimension for predicting the historical trend, denoted as the trend influence factor, use the second feature data within a unit time as the feature dimension of the mutation factor, denoted as the mutation influence factor, and organize the obtained information popularity trend influence factor and mutation influence factor into different data structures;
[0009] Step 3: According to the characteristics of the data, use the Transformer architecture to extract the features of the trend sequence, use a model with good performance in dealing with local features to process the mutation influencing factors, extract the mutation characteristics, and combine the features of the trend influencing factors and the mutation influencing factors to obtain the final prediction result.
[0010] The specific content of the above Step 1 includes the following:
[0011] Collect a part of the information data of various topics and perform positive, negative, and neutral sentiment annotations;
[0012] Fine-tune the pre-trained model with the annotated data to obtain an information sentiment classification model;
[0013] Collect according to historical hot topics, use the sentiment classification model to obtain the sentiment type of each piece of information, and obtain the positive, negative, and neutral sentiment numbers of the information within a unit time as feature dimensions for storage.
[0014] Regarding the first feature data within a unit time obtained as the feature dimension for predicting the historical trend, denoted as the trend influencing factor, and the second feature data within a unit time as the feature dimension of the mutation factor, denoted as the mutation influencing factor, it includes the following:
[0015] Divide the number of forwards, comments, likes, reads, entity numbers, collections, and sentiment numbers of topic articles showing the sequence trend development law of information heat within a unit time into the first feature data, and name it the trend influencing factor;
[0016] Further organize the number of official media and personal media to obtain the number of communication media, the number of authoritative media, the number of user participations above the upper threshold of the number of fans, and the proportions of participating users and new participating users within a unit time as the second feature data, and name it the mutation influencing factor.
[0017] The organization of the obtained information heat trend influencing factor and mutation influencing factor into different data structures includes:
[0018] Organize the trend influencing factor data in the way of time window. For each topic, select a time window size n, which means selecting n time-length data as feature data, and the heat ranking at the n+1th time as the label data of this feature;
[0019] Two-dimensionally structure the mutation influencing factor data, put the feature data of each time period into a column, each column represents all the feature data of each time step, and each row represents the same feature data of different time steps.
[0020] The above Step 3 includes:
[0021] Capture long - distance dependencies through the self - attention mechanism in the Transformer architecture to extract features of the trend influencing factors of the topic;
[0022] Since mutation factors have strong representations in the local information of the data, the local feature information of the mutation - influencing factor data is fully extracted through the local information feature extraction model;
[0023] Integrate the features of the extracted trend - influencing factors and mutation - influencing factors through the normalization layer and the feed - forward neural network of the Transformer architecture to obtain the final heat prediction value.
[0024] An information heat prediction system based on trend factors and mutation factors, the system includes: a data acquisition module, a data sorting and partitioning module, and a model prediction module;
[0025] The data acquisition module: is used to obtain the historical information data of a certain topic through the sentiment classification model;
[0026] The data sorting and partitioning module: for each topic data, take the first - feature data within the acquired unit time as the feature dimension for predicting the historical trend, denoted as the trend - influencing factor, take the second - feature data within the unit time as the feature dimension of the mutation factor, denoted as the mutation - influencing factor, and organize the obtained information heat trend - influencing factors and mutation - influencing factors into different data structures;
[0027] The model prediction module: is used to extract the trend sequence features using the Transformer architecture according to the characteristics of the data, use a model with good performance in processing local features to process the mutation - influencing factors to extract the mutation characteristics, and combine the features of the trend - influencing factors and the mutation - influencing factors to obtain the final prediction result.
[0028] The data sorting and partitioning module includes a trend - influencing factor partitioning unit, a mutation - influencing factor partitioning unit, and an organizing unit;
[0029] The trend - influencing factor partitioning unit: is used to divide the number of forwards, comments, likes, reads, entity numbers, collections, and sentiment numbers of topic articles showing the sequence trend development law of information heat within the unit time into the first - feature data and name it the trend - influencing factor;
[0030] The mutation - influencing factor partitioning unit: is used to further organize the number of official media and personal media to obtain the number of communication media, the number of authoritative media, the number of user participations above the upper threshold of the number of fans, the proportion of participating users and new participating users within the unit time as the second - feature data and name it the mutation - influencing factor;
[0031] The sorting unit: sorts the trend influencing factor data in the way of using a time window. For each topic, a time window size n is selected, indicating that n data with a time length are selected as feature data, and the popularity ranking at the (n + 1)-th time is used as the label data for this feature; two-dimensionally structures the mutation influencing factor data, puts the feature data of each time period into a column, each column represents all the feature data of each time step, and each row represents the same feature data of different time steps.
[0032] The model prediction module specifically includes:
[0033] Captures long-distance dependencies through the self-attention mechanism in the Transformer architecture to extract features of the trend influencing factors of the topic;
[0034] Since the mutation factors have strong representations in the local information of the data, the local feature information of the mutation influencing factor data is fully extracted through the local information feature extraction model;
[0035] Integrates the features of the extracted trend influencing factors and mutation influencing factors through the normalization layer and the feedforward neural network of the Transformer architecture to obtain the final popularity prediction value.
[0036] The present invention has the following advantages:
[0037] 1. For the problem of less consideration of emotional factors in popularity prediction, a pre-trained model is used for fine-tuning to obtain an emotion classification model, and then the emotion classification model is directly used to obtain data when collecting data, which improves the efficiency of obtaining data features.
[0038] 2. For the problems of unclear division of popularity influencing factors and less consideration of mutation factors, each dimension of the popularity influencing factors is analyzed, and the data is divided into two categories according to the influencing characteristics of each dimension: trend influencing factors and sudden influencing factors.
[0039] 3. For the different data characteristics of different influencing factor data, different application models are built to process the feature extraction of the data. For the trend influencing factors, considering that long time series features need to be extracted and parallel training can be carried out to improve efficiency, the present invention selects to build a Transformer framework model for processing. For the problem of mutation data feature extraction, a model with advantages in extracting local features is selected for processing, and then the trend factor features and mutation factor features are integrated to obtain the final prediction result. The built model can realize end-to-end training and improve the efficiency of model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a framework schematic diagram of the present invention;
[0041] Figure 2 Schematic diagram for data collection and analysis;
[0042] Figure 3 Schematic diagram for data sorting and partitioning;
[0043] Figure 4 Schematic diagram of the model structure. Detailed implementation manners
[0044] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. Usually, the components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the protection scope of the claimed present application, but merely represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application. The present invention will be further described below with reference to the accompanying drawings.
[0045] The present invention specifically relates to a method for predicting information popularity based on trend factors and mutation factors. Aiming at the problem of fewer factors considered in the current information popularity prediction, the present invention makes various selections of influencing factors, and the selected influencing factors can reflect the change trend of information popularity from multiple aspects; the pre-trained model is fine-tuned using sentiment data so that in the subsequent data collection stage, this fine-tuned model can be used to more efficiently obtain sentiment attributes, thereby improving the overall data processing efficiency.
[0046] Aiming at the problem of limited accuracy of popularity prediction, the present invention takes into account mutation influencing factors while considering trend influencing factors. First, the influencing factors are clearly divided into trend influencing factors and mutation influencing factors; subsequently, different prediction models are built to process these two types of factors respectively, and the two are effectively integrated through the models, and the final prediction result is obtained by synthesizing the trend law and mutation of the development of information popularity.
[0047] As Figure 1 shown, the present invention specifically includes three parts: data collection and analysis, data sorting and partitioning, and model building and prediction.
[0048] First, in data collection and analysis, before officially collecting data, a part of information data on different topics needs to be collected first and sentiment is labeled. This part of data is used for fine-tuning training of the pre-trained model to obtain a sentiment classification model.
[0049] Secondly, collect historical information data of a certain topic on each information platform in chronological order. When collecting data, use a sentiment classification model to obtain sentiment information for each piece of information. The sentiment classification results include: positive, negative, and neutral. The sentiment attribute of the information will be used as one of the feature dimensions for topic popularity prediction; perform data sorting and division. For each topic data, use data such as the number of forwards, comments, likes, reads, entities, collections, positive sentiment numbers, negative sentiment numbers, and neutral sentiment numbers obtained within a unit time as the feature dimensions for predicting historical trends; consider the factors that affect the sudden change in the popularity of each topic, and classify factors such as the number of media in a unit time, the number of authoritative media, the number of users with more than 100,000 fans participating, and the ratio of existing participants to new participants as the feature dimensions of the mutation factors; organize the two parts of the information heat trend influencing factors and mutation influencing factors into different data structures.
[0050] Finally, build a model and make predictions. According to the characteristics of the data, first use the Transformer architecture to extract the trend sequence features, then use a model with better performance in processing local features to process the mutation influencing factors to extract the mutation characteristics, and finally combine the trend influencing factors and mutation influencing factor features to obtain the final prediction result.
[0051] Furthermore, as Figure 2 shown, the data collection and analysis include the following:
[0052] Select information collection platforms, including but not limited to mainstream social platforms such as Weibo, WeChat, Twitter, and Facebook;
[0053] Before officially starting data collection, first collect a part of the information data of various topics, and perform sentiment annotation during collection. The sentiment annotated here is mainly divided into: positive, negative, and neutral;
[0054] Use the annotated data for fine-tuning pre-trained models including but not limited to BERT, GPT, XLNet, RoBERTa, ALBERT, etc. to obtain an information sentiment classification model;
[0055] Officially collect information data, collect data on the selected social platforms, mainly collect data according to historical hot topics, and store the collected data in chronological order. The stored data content includes: the number of forwards, comments, likes, reads, entities, collections, media quantity, user participation, heat ranking, etc. of topic articles within a unit time;
[0056] When collecting data, use the sentiment classification model to obtain the sentiment type of each piece of information, and then obtain the positive, negative, and neutral sentiment numbers of the information within a unit time, and store them as feature dimensions.
[0057] Further, as Figure 3 shown, the data sorting and partitioning specifically include the following:
[0058] First, in the process of predicting the popularity of information topics, the popularity ranking is used as the evaluation criterion for information popularity, and the feature dimensions other than the evaluation criterion are used as the basis for evaluation. Here, the evaluation basis needs to be further divided.
[0059] Data such as the number of forwards, comments, likes, reads, entities, collections, and sentiment of topic articles within a unit time show the development law of the sequence trend of information popularity. Therefore, these data are divided into the same type of data and named trend influencing factors.
[0060] The official media and personal media numbers are further sorted to obtain data such as the number of communication media, authoritative media, the number of users with more than 100,000 fans participating, and the ratio of participating users to new participating users within a unit time. The changes in this type of data have a significant impact on the mutation of information popularity. Therefore, this type of data is divided into the same category and named mutation influencing factors.
[0061] Since the processing models of the two factors are different, data structure adjustment is required; the trend influencing factor data is sorted using the time window method. For each topic, first select a time window size n, indicating that n time-length data are selected as feature data, and the popularity ranking at the (n + 1)-th time is used as the label data for this feature; the mutation influencing factor data is two-dimensionally structured, and the feature data within each time period is placed in a column. Each column represents all the feature data at each time step, and each row represents the same feature data at different time steps.
[0062] Further, as Figure 4 shown, in the process of building the model, the trend influencing factors and mutation influencing factors of information popularity need to be considered. Since there are differences in the data structures of different factors, when building the model, it is necessary to fully consider how to extract data features. Therefore, the model building and prediction specifically include the following:
[0063] The Transformer architecture has made remarkable achievements in the field of natural language processing. This framework can capture long-distance dependencies through the self-attention mechanism and has great advantages in extracting long-term dependencies in sequence data. In the model built in the present invention, the Transformer architecture is used to extract features of the trend influencing factors of the topic.
[0064] Further, first convert the trend influencing factors into embedding vectors that can be processed by the Transformer architecture, and then pass through Figure 4Process the data for the leftmost part of the encoder section, and then pass the features obtained from the encoder section processing into Figure 4 the middle process of the decoder section. When the decoder processes data, it also needs to pass in the embedding vector, and then obtain the features of the trend factor after passing through the last multi-head attention layer of the decoder section.
[0065] Since the mutation factor has a strong representation in the local information of the data, a model that can fully extract local information can be used to obtain features. The present invention uses convolutional neural network architecture models that are advantageous in extracting local features from two-dimensional data, including but not limited to SENet, Adaptive Convolutional Neural Network (ACNN), etc. In the local information feature extraction model, the steps of the final classification calculation of models such as SENet and ACN are removed, and the results before the classification calculation are retained as feature data. SENet is a special convolutional neural network architecture. By introducing the Squeeze and Excitation modules, it enables the model to adaptively adjust the importance of each channel in the convolutional neural network, thereby enhancing the model's ability to model the relationship between feature channels. This mechanism allows the SENet model to learn which feature channels are more important during the training process and accordingly adjust their contributions to the final prediction, thus enhancing the model's ability to extract mutation information. Compared with traditional CNN, ACN can dynamically increase, delete, or adjust its neurons, layers, or connection methods according to the characteristics of the data and the requirements of the task, and automatically adjust its parameters to adapt to different data distributions. Its design includes some special mechanisms, such as Dynamic Convolution, which can adaptively adjust the convolution parameters according to the input image, and dynamically aggregate multiple parallel convolution kernels through the attention mechanism, thereby significantly enhancing the model's expressive ability while keeping the computational cost low. Therefore, it can further strengthen the local features with strong influencing factors during the two-dimensional data processing process to achieve the extraction of mutation information features.
[0066] After the features of the trend influencing factor and the mutation influencing factor are respectively extracted, the two parts of the features are integrated through the normalization layer of the Transformer and the network weighted calculation of the feed-forward neural network to obtain the final heat prediction value, so as to achieve end-to-end training and improve the efficiency of model training.
[0067] During the training process of the present invention, methods including but not limited to network search, genetic algorithm, Bayesian optimization, etc. are used to automatically optimize the hyperparameters of the Transformer model and the local information feature extraction model to improve the tuning efficiency of the model.
[0068] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and improvements, and can be changed within the scope of the concept described herein through the above teachings or the techniques or knowledge in related fields. Any changes and variations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A method for predicting information heat based on trend factors and mutation factors, characterized by: The prediction method comprises: Step 1: Obtain historical information data on a topic through a sentiment classification model; Step 2: For each topic data, the first characteristic data obtained within a unit time is used as the characteristic dimension for predicting historical trends, recorded as trend influencing factors, and the second characteristic data within a unit time is used as the characteristic dimension of mutation factors, recorded as mutation influencing factors, and the obtained information heat trend influencing factors and mutation influencing factors are sorted into different data structures; Step 3: According to the characteristics of the data, the Transformer architecture is used to extract the trend sequence features, and the model with good performance in processing local features is used to process the mutation influencing factors to extract the mutation characteristics. The trend influencing factors and mutation influencing factors are combined to obtain the final prediction results. The method of using the first characteristic data acquired within a unit time as the characteristic dimension for predicting the historical trend, recorded as the trend influencing factor, and using the second characteristic data acquired within a unit time as the characteristic dimension for the mutation factor, recorded as the mutation influencing factor, includes the following contents: The forwarding number, comment number, like number, reading number, entity number, collection number and sentiment number of topic articles that show the sequence trend of information popularity within a unit time are divided into the first feature data and named as trend influencing factors; The number of official media and personal media is further sorted out to obtain the number of dissemination media per unit time, the number of authoritative media, the number of users participating above the threshold of the number of fans, and the proportion of participating users and new participating users, which are divided into the second characteristic data and named as mutation influencing factors; The step three comprises: The self-attention mechanism in the Transformer architecture is used to capture long-distance dependencies and extract features of factors affecting topic trends. Since mutation factors are strongly represented in the local information of the data, the local feature information of the mutation influencing factor data is fully extracted through the local information feature extraction model; The extracted features of trend influencing factors and mutation influencing factors are integrated through the normalization layer and feedforward neural network of the Transformer architecture to obtain the final heat prediction value.
2. The information heat prediction method based on trend factors and mutation factors according to claim 1 is characterized by: The step 1 specifically includes the following contents: Collect some information data on various topics and annotate them with positive, negative and neutral sentiments; Fine-tune the pre-trained model with the labeled data to obtain an information sentiment classification model; The information is collected according to historical hot topics, and the sentiment classification model is used to obtain the sentiment type of each piece of information. The number of positive, negative and neutral sentiments of the information per unit time is obtained and stored as the feature dimension.
3. The information heat prediction method based on trend factors and mutation factors according to claim 1 is characterized by: The information heat trend influencing factors and mutation influencing factors obtained are sorted into different data structures, including: Use the time window method to organize the trend influencing factor data. For each topic, select a time window size n, which means that n time length data are selected as feature data, and the popularity ranking of the n+1th time is used as the label data of this feature; The mutation influencing factor data is structured in two dimensions, and the characteristic data of each time period is placed in a column. Each column represents all the characteristic data of each time step, and each row represents the same characteristic data of different time steps.
4. An information heat prediction system based on trend factors and mutation factors, characterized by: The system comprises: a data acquisition module, a data sorting and division module and a model prediction module; The data acquisition module is used to acquire historical information data of a topic through a sentiment classification model; The data sorting and division module is used to, for each topic data, use the first characteristic data obtained within a unit time as the characteristic dimension for predicting the historical trend, recorded as the trend influencing factor, use the second characteristic data within a unit time as the characteristic dimension of the mutation factor, recorded as the mutation influencing factor, and sort the obtained information heat trend influencing factors and mutation influencing factors into different data structures; The model prediction module is used to extract trend sequence features using the Transformer architecture according to the characteristics of the data, process mutation influencing factors using a model with good performance in processing local features, extract mutation characteristics, and combine the characteristics of trend influencing factors and mutation influencing factors to obtain the final prediction result; The data sorting and division module includes a trend influencing factor division unit, a mutation influencing factor division unit and a sorting unit; The trend influencing factor division unit is used to divide the forwarding number, comment number, like number, reading number, entity number, collection number and sentiment number of the topic article showing the sequence trend development law of information popularity within a unit time into the first feature data, and name it as the trend influencing factor; The mutation influencing factor classification unit is used to further sort out the number of official media and the number of personal media, and obtain the number of dissemination media, the number of authoritative media, the number of users participating in the threshold of the number of fans per unit time, and the proportion of participating users and new participating users, which are classified into the second characteristic data and named as mutation influencing factors; The collating unit: collates the trend influencing factor data by using a time window method. For each topic, a time window size n is selected, which means that n time length data are selected as feature data, and the popularity ranking of the n+1th time is used as the label data of this feature; the mutation influencing factor data is two-dimensionally structured, and the feature data of each time period is put into a column, each column represents all feature data of each time step, and each row represents the same feature data of different time steps; The model prediction module specifically includes: The self-attention mechanism in the Transformer architecture is used to capture long-distance dependencies and extract features of factors affecting topic trends. Since mutation factors are strongly represented in the local information of the data, the local feature information of the mutation influencing factor data is fully extracted through the local information feature extraction model; The extracted features of trend influencing factors and mutation influencing factors are integrated through the normalization layer and feedforward neural network of the Transformer architecture to obtain the final heat prediction value.
Citation Information
Patent Citations
Social media content popularity prediction method fusing multi-scale cascade and time sequence characteristics
CN115495669A
Public opinion popularity prediction method
CN118193824A