Social media hot list topic popularity prediction method and system
The integration of time series and non-time series data with semantic and lifecycle features in the SELF-TF model enhances social media topic popularity prediction accuracy and adaptability, addressing existing challenges in model obsolescence and early prediction.
Patent Information
- Application Number
- CN202510454239.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-15
AI Technical Summary
The existing technology has low accuracy in predicting social media topic popularity, and cannot conduct early predictions. The model is prone to obsolete over time and is not adaptable to the list topic scenarios, especially the cold start of new topics and the life cycle characteristics of topics are not fully utilized.
The SELF-TF model is constructed, and by fusing time-series data and non-time-series data, integrating topic semantic information and adding early life cycle characteristics of topics, learning non-time semantic information by BERT, LSTM performs local time processing and multi-head attention, and decoding with the decoder structure of TFT.
It improves the accuracy of predicting hot topics on the list, especially in the early prediction of new topics and the timeliness of models, and improves the adaptability and interpretability of models.
Smart Images

Figure CN120318002A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of social media analysis and topic popularity prediction, and particularly to an optimized method and device for constructing a prediction model to estimate the future popularity potential of a topic based on the historical performance data of the topic and the data of other topics simultaneously listed on the list. Background Art
[0002] Social media lists are a key battlefield in the competition for user attention. With the rise of the attention economy, the demand for in-depth and effective analysis of trending topics is increasing. Accurately predicting the future popularity of a topic is crucial for public opinion monitoring, content distribution, topic recommendation, and online marketing. Popular topics are ranked by popularity scores, which are also displayed in the rankings. On Twitter, this score is defined by the number of tweets discussing the topic, while on Weibo, it is calculated based on multiple factors such as search volume, discussion, and real-time activities, providing a comprehensive popularity index. Predicting the popularity of list topics is a multivariate time series regression task that uses historical trend data and external relationship data. This process analyzes the semantics, features, and historically observable inputs of the topic to predict future popularity metrics.
[0003] In the prior art, predicting the popularity of list topics is a specific application of topic popularity prediction. One of the key focuses is to apply topic popularity prediction methods to the list scenario while emphasizing list topic ranking and competition. Another approach is to use time series prediction as a technical method, which mainly focuses on popularity trends rather than dissemination paths, making it very suitable for the characteristics of list topic data.
[0004] 1. Prediction of Social Media Topic Popularity
[0005] Predicting the popularity of social media topics involves predicting the future or final popularity at a specific point in time. This task generally follows two main methods: point process modeling and feature-based modeling.
[0006] For the popularity prediction method based on point process modeling, the information dissemination process is mainly regarded as an arrival process of user forwarding behaviors. The core of this type of method is to model the rate function in the arrival process. For the popularity prediction type based on feature modeling, four types of information are mainly used, namely message content, publisher characteristics, time dynamics, and dissemination structure, etc., for representation modeling, among which the temporal characteristics are the most crucial. The modeling method based on point process is suitable for describing discrete events (such as likes, forwards), but insufficient data will limit the model performance and its applicability in the list scenario is poor.
[0007] 2. Prediction of Time Series Data
[0008] Time series prediction is to learn and analyze the historical data of the previous t-1 moments to predict the data values in the future time period. Time series data prediction is generally divided into three categories: statistical learning, machine learning, and deep learning methods.
[0009] In terms of deep learning methods, they perform excellently in complex, non-linear, and long-term dependent time series prediction tasks and are the most widely used type of method. They mainly include prediction models with five structural designs: Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Self-Attention Mechanism (Transformer), and Graph Networks (GNNs). Among them, TFT is a hybrid deep learning time series prediction method. The architecture of the TFT model combines several main components, including the input layer and embedding layer, Variable Selection Network, LSTM encoder / decoder, self-attention mechanism, and Gated Residual Network. One of its advantages is to improve the prediction accuracy by combining multiple data sources and features, including static variables, known future inputs, and historically observed exogenous time series.
[0010] To sum up, although the existing time series data prediction methods can predict the popularity of list topics, the effect is not good, mainly manifested in three aspects:
[0011] First, the accuracy is not high. Traditional time series prediction models are difficult to capture semantic information, while traditional Transformer models can capture semantics but cannot utilize time series data and perform poorly in topic prediction. Therefore, applying the existing univariate time series prediction methods to the current scenario and using the historical popularity of the collected topics as input to predict the future popularity of the topics results in poor prediction effects.
[0012] Second, it is difficult to predict the cold start of new topics. The list topic prediction method based on time series data only uses the historical records of list topics as input. For newly listed topics, there is not enough historical data to support their prediction, and the prediction effect is extremely poor.
[0013] Third, the model is prone to becoming outdated or obsolete over time. It can predict topics that have appeared in the training samples, but if the training topics and test topics are not in the same distribution, such as using historical list data with a long interval to predict the current or future list topics, the prediction accuracy will drop rapidly. Since topic features are closely related to time and the topic types may be completely different at different times, it is required that the model be updated frequently. If the update iteration is slow, the prediction accuracy will be greatly reduced.
[0014] There are three main reasons for the low accuracy, inability to carry out early prediction, and the model being prone to becoming outdated in the existing technology:
[0015] First, only time series is used for prediction, and the non-time series information of the topic is not fully utilized. The non-time series data of the topic includes the type, semantics, contributors, sources, etc. of the topic. These data do not change with time. By using these data to judge the "attention-grabbing" degree of the topic and the reliability of the source, based on the performance of similar topics in history, the early prediction and future impact of the target topic are important inputs that cannot be ignored;
[0016] Second, only the fixed time window information is used, and the early data of the topic is not considered. From the moment a topic first appears in the trend list, it can start to be observed. Traditional time series prediction tasks only use the observable original input (a time series data with a fixed time length), ignoring the early change information of the topic, which is not conducive to the accuracy of topic prediction. At the same time, the life cycle of the topics on the list is of unequal length and has obvious periodic characteristics. In order to better utilize the early life characteristics of the topic, the model structure proposed by the present invention needs to be able to integrate the early information of the topic to ensure that the data from the first appearance of the topic in the trend list to the prediction time is not utilized.
[0017] Third, the existing social media topic popularity prediction methods are not very adaptable to the prediction of the popularity of the topics on the list. In the scenario of the topics on the list, the topics on the list have short topic texts and it is difficult to obtain detailed topic data such as the comment and repost process. At the same time, the hot topics on the list also have the characteristics of rapid generation, heat explosion and rapid disappearance, resulting in the topic disappearing within a few hours or even minutes, greatly reducing the amount of available data. These unique characteristics hinder the traditional methods and lead to poor performance in predicting trends.
[0018] Therefore, it is urgent to study a new topic popularity prediction method that integrates time series data and non-time series data, integrates topic semantic information and adds the early life cycle characteristics of the topic. Summary of the Invention
[0019] In order to solve the three challenges of low accuracy, inability to conduct early prediction, and the model being prone to becoming outdated over time in the existing topic heat prediction technology, by analyzing the data characteristics of the list scenario, this application proposes a list topic heat prediction model with the integration of time series data and non-data data as input. This model improves the effect of topic prediction in the list scenario.
[0020] In the first aspect, an embodiment of this application provides a method for predicting the popularity of hot topics on social media. The method includes:
[0021] Steps for constructing the prediction model: Construct the SELF-TF model for predicting the topic popularity of the list. The SELF-TF model is used to fuse time-series data and non-time-series data, integrate topic semantic information, and add periodic features of the topic. The SELF-TF model includes: a semantic time fusion encoder and a time-series fusion decoder. The semantic time fusion encoder further includes: a topic semantic encoder and a time-series feature encoder.
[0022] Steps for predicting topic popularity: Input the integrated topic semantic information and the early life cycle features of the topic into the semantic time fusion encoder. After learning the non-temporal semantic information and integrating the semantic and time-varying information of different time steps, enter the decoder for decoding to complete and output the predicted topic popularity value.
[0023] In the embodiments of the present application, the above steps for predicting topic popularity include:
[0024] Semantic encoding step: Receive the non-time-series data and time-series data of the list topic through the topic semantic encoder to form a semantic-based prediction and output the learned semantic vector. Among them, the non-time-series data includes: the semantics, sponsors, types, and introductions of the list topic; the time-series data includes: the discussion volume, like volume, popularity value, and time of the list topic.
[0025] Time-series feature encoding step: Input the learned semantic vector, historical features, and early cycle features into the time-series feature encoder to integrate the time dynamic change features of the topic. By normalizing and encoding the topic time-series features, learn the topic information of the list topic changing dynamically over time and output the learned list topic time-series vector. Among them, the topic information includes: the encoding information of the past known time series and the time series information of the future list topic.
[0026] In the embodiments of the present application, the above steps for predicting topic popularity further include:
[0027] Time-series fusion decoding step: Input the semantic vector learned by the topic semantic encoder and the list topic time-series vector learned by the time-series feature encoder into the time-series fusion decoder for decoding and fusion. Learn the semantic and time-series relationships among the semantic input of the entire life cycle of the list topic, the observable long-term input, and the short-term input changing with the list topic at any time. Obtain the predicted topic popularity value according to the learning result.
[0028] In the embodiments of the present application, the above topic semantic encoder uses BERT to learn non-temporal semantic information; the time-series feature encoder uses LSTM for local time processing and multi-head attention to integrate the semantic and time-varying information of different time steps; the time-series fusion decoder uses the decoder structure of TFT.
[0029] Second aspect, an embodiment of the present application provides a social media hot list topic popularity prediction system, which adopts the above-mentioned social media hot list topic popularity prediction method. The system includes:
[0030] Prediction model construction module: Construct a list topic popularity prediction SELF-TF model. The SELF-TF model is used to fuse time series data and non-time series data, integrate topic semantic information and add periodic features of the topic. The SELF-TF model includes: a semantic time fusion encoder and a time series fusion decoder; the semantic time fusion encoder further includes: a topic semantic encoder and a time series feature encoder;
[0031] Topic popularity prediction module: Integrate topic semantic information and early life cycle characteristics of the topic, input them into the semantic time fusion encoder, learn non-temporal semantic information, integrate semantic and time-varying information of different time steps, and then enter the decoder for decoding to complete, and output the to-be-predicted topic popularity value.
[0032] In the embodiment of the present application, the above-mentioned semantic time fusion encoder further includes: a topic semantic encoder and a time series feature encoder;
[0033] The topic semantic encoder receives non-time series data and time series data of the list topic, forms a semantic-based prediction, and outputs the learned semantic vector; among them, the non-time series data includes: the semantics, sponsors, types and introductions of the list topic; the time series data includes: the discussion volume, like volume, popularity value and time of the list topic;
[0034] Input the learned semantic vector, historical features and early cycle features into the time series feature encoder, integrate the time dynamic change features of the topic, learn the topic information of the list topic changing dynamically over time by normalizing and encoding the topic time series features, and output the learned list topic time series vector, where the topic information includes: the time series information coding information of the past known time series and the future list topic.
[0035] In the embodiment of the present application, input the semantic vector learned by the topic semantic encoder and the list topic time series vector learned by the time series feature encoder into the time series fusion decoder for decoding and fusion, learn the semantic and time series relationship between the semantic input of the entire life cycle of the list topic, the observable long-term input and the short-term input changing between list topics at any time, and obtain the to-be-predicted topic popularity value according to the learning result.
[0036] In the embodiment of the present application, the above-mentioned topic semantic encoder uses BERT to learn non-temporal semantic information; the time series feature encoder uses LSTM for local time processing and multi-head attention to integrate semantic and time-varying information of different time steps; the time series fusion decoder uses the decoder structure of TFT.
[0037] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method for predicting the topic popularity of a social media hot list are implemented.
[0038] In a fourth aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for predicting the topic popularity of a social media hot list as described above are implemented.
[0039] Compared with the related prior art, it has the following outstanding beneficial effects:
[0040] 1) The method of the present invention proposes a prediction model structure for the popularity of list topics that fuses time-series data and non-data data as inputs. This structure receives four types of non-time-series data, namely the semantics, originator, type, and introduction of the list topic, and six types of time-series data, such as the discussion volume, like volume, popularity value, and time of the list topic, as inputs. By effectively fusing time-series and non-time-series data, compared with traditional time-series prediction models, it solves the problem of insufficient utilization of non-time-series information.
[0041] 2) The method of the present invention proposes a module for representing and integrating the early life characteristics of list topics, using a total of 10 features in 3 categories, namely time, achievement, and fluctuation, to describe the historical global information of list topics outside the current slice. The representation of the features is as Figure 3 shown. Fusing the features into the prediction model better adapts to the characteristics of the variable length and periodicity of list topics, helps the model better capture the trend of topics and locate the topic stage where the prediction interval is located, and also has better interpretability.
[0042] 3) The method of the present invention proposes a pre-trained language model to carry out the semantic feature representation of list topics. The topic semantic encoder integrates the semantic information of the list topic, enabling the model to more deeply understand the name, type, and details of the topic. The model organically combines time-series data and static semantic data, and this method solves the challenge of early prediction of new trend topics lacking historical data. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:
[0044] Figure 1 It is a schematic diagram of the method for predicting the topic popularity of a social media hot list of the present invention;
[0045] Figure 2 It is a schematic diagram of the structure of the prediction model of an embodiment of the present invention;
[0046] Figure 3 Schematic diagram of the experimental results of the model of this application in the embodiments of the present invention;
[0047] Figure 4 Schematic diagram of representing the early life characteristics of the topic used by the model in the embodiments of the present invention;
[0048] Figure 5 Schematic diagram of the topic popularity prediction system for the popular list on social media of the present invention;
[0049] Figure 6 Schematic diagram of the computer hardware of the present invention. Detailed implementation manners
[0050] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one (piece) of the following" or its similar expression refers to any combination of these items, including any combination of single item (piece) or plural items (pieces). For example, at least one (piece) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0051] It should also be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.
[0052] It should also be understood that in various embodiments of the present invention, the sequence numbers of the above processes do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0053] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.
[0054] The unit described as a separation component may or may not be physically separated. The component presented as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0055] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0056] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store program codes.
[0057] To make the above features and effects of the present invention more clearly understandable, specific embodiments are hereinafter given and described in detail in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are only for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.
[0058] The following is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment, and in order to avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.
[0059] Through sufficient experimental comparison and data analysis, the present invention finds that most current time series data prediction models only accept one type of input of time series data, fail to effectively integrate time series and non-time series data structurally, and at the same time, the influence relationship between different list topics is restricted by time conditions and topic semantic similarity.
[0060] Therefore, after comparing the effects of different models, the present invention proposes a trending topic popularity prediction model SELF-TF (Semantic Fusion and Early Lifecycle Feature-Enhanced Time Series Model for Trending topic Popularity Forecasting) that takes the fusion of time series data and non-time series data as input, integrates topic semantic information and incorporates the early lifecycle features of the topic. The model structure is as shown in Figure 2 and takes static semantic data, time-varying past inputs, and time-derived future known inputs as input.
[0061] The non-temporal semantic information, including topic name, type, and introduction, is learned through the BERT (Bidirectional Encoder Representations from Transformers) mechanism. For the handling of time dependencies, the architecture uses LSTM for local time processing and multi-head attention to integrate semantic and time-varying information at different time steps. Finally, various comparison and ablation experiments are conducted on the collected Weibo trending topic dataset, indicating that the method proposed by the present invention can effectively improve the prediction accuracy of trending topic popularity, the prediction effect of early topics, and alleviate the problem of model failure over time.
[0062] The method of the embodiments of the present application will be described in detail below in conjunction with specific embodiments:
[0063] Embodiment 1
[0064] As shown in Figure 1 shown, Figure 1 is a schematic diagram of the method for predicting the popularity of trending topics on social media according to the present invention; the present invention provides a method for predicting the popularity of trending topics on social media, and the method includes:
[0065] Step 101 of constructing a prediction model: Construct a trending topic popularity prediction SELF-TF model, where the SEL-TF model is used to fuse time series data and non-time series data, integrate topic semantic information, and incorporate the periodic features of the topic; the SEL-TF model includes: a semantic time fusion encoder and a time series fusion decoder; the semantic time fusion encoder further includes: a topic semantic encoder and a time series feature encoder.
[0066] Specifically, the structure of the prediction model in the specific embodiments of the present application is as shown in Figure 2As shown in the figure, the model encoding part consists of semantic encoding based on the pre-trained language model and traditional list topic features and time feature encoding based on LSTM list topics. The decoding part refers to the TFT (Temporal Fusion Transformer) list topic Decoder to predict the topic heat value.
[0067] The specific embodiment of the present application proposes a list topic popularity prediction model structure that uses a fusion of time series data and non-data data as input. The structure receives four types of non-time series data, namely the semantics, initiator, type, and introduction of the list topic, and six types of time series data, namely the discussion volume, like volume, popularity value, and time of the list topic as input. Compared with traditional time series prediction models, this structure solves the problem of insufficient utilization of non-time series information by effectively fusing time series and non-time series data.
[0068] In the specific embodiment of this application, the topic popularity prediction technology proposed can be divided into 2 point prediction methods and 3 time series data prediction methods. The present invention selects representative methods in each category for comparison. The topic popularity prediction results of the proposed SELF-TF method and other four different time series prediction models on Weibo and Twitter datasets. The results show that the performance of SELF-TF is better than the other 7 models, such as Figure 3 As shown in the figure, because it integrates semantic and early life cycle feature enhancement modules, it improves the performance of the model by effectively learning time series data and topic static semantic information, thereby improving the utilization of topic data and making the prediction of new topics more accurate.
[0069] Topic popularity prediction step 102: The topic semantic information and early life cycle features of the topic are integrated and input into the semantic time fusion encoder. After learning the non-temporal semantic information and integrating the semantics and time-varying information of different time steps, the decoder is decoded and the predicted topic popularity value is output.
[0070] In this embodiment of the application, the topic popularity prediction step 102 includes:
[0071] Semantic encoding step: Receive the non-time series data and time series data of the list topic through the topic semantic encoder, form a semantic-based prediction, and output the learned semantic vector; among them, the non-time series data includes: the semantics, initiator, type and introduction of the list topic; the time series data includes: the discussion volume, like volume, popularity value and time of the list topic;
[0072] Among them, the semantic encoding consists of the topic name, topic type, and topic lead. The temporal features consist of topic features and features between the time list topics. The topic semantic encoder integrates the semantic information of the list topics into the network list topics, deeply understands information such as topic type and topic lead, forms a semantic-based prediction, solves the problem that the list topics cannot be predicted for newly listed topics due to the lack of historical information, and uses S to represent the learned semantic vector.
[0073] A pre-trained language model is introduced to carry out the semantic feature representation of the list topics. The topic semantic encoder integrates the semantic information of the list topics, enabling the model to more deeply understand the name, type, and details of the topics. The model organically combines the temporal data and static semantic data. This method solves the challenge of early prediction for new trend topics lacking historical data.
[0074] Technical effect: Ablation experiments show that the accuracy of the test data on the Weibo list has increased by 7.79% after using this structure. Control group experiments were carried out under different dataset sizes, and the effect of the semantic module is more obvious on larger datasets. On larger datasets, compared with the model without the semantic module, adding the semantic module reduces the loss by 3.8% (MAE) and 4.6% (RMSE). In contrast, on smaller datasets, the contribution of this module is relatively limited, and the loss reduction is only 0.3% (MAE) and 1.9% (RMSE). This shows that the semantic module is more effective when the dataset is larger because larger datasets have more topics and richer semantic information.
[0075] Steps of temporal feature encoding: Input the learned semantic vector, historical features, and early cycle features into the temporal feature encoder, integrate the time dynamic change features of the topic, and learn the topic information of the list topics changing dynamically over time by normalizing and encoding the topic temporal features, and output the learned list topic temporal vector. Among them, the topic information includes: the encoding information of the past known time series and the time series information of future list topics.
[0076] The list topic temporal feature encoder of the list topic integrates the time dynamic change features of the topic into the network list topic network. By normalizing and encoding the topic temporal features, the features are numerically encoded and spliced into a feature vector (each feature corresponds to one dimension of the vector). Subsequently, the feature vector is normalized, and the various features are uniformly normalized to between 0 and 10. To learn the topic information of the list topics changing dynamically over time, including the encoding of the past known time series and the known time series information of future list topics (that is, derivative time information, such as what time, day of the week, whether it is a holiday, weekend, etc. at the predicted future moment. Such time features can help understand the active rules of the topic at different time periods), and use Ht to represent the temporal vector of the list topic learned at time t.
[0077] An early life feature representation and integration module for the list topic is proposed. Ten features in three categories of time, achievement, and fluctuation are used to describe the historical global information of the list topic outside the current slice. The feature representation is as Figure 4 shown. By integrating the features into the prediction model, it better adapts to the variable length and periodic characteristics of the list topic, helps the model better capture the topic trend and locate the topic stage where the prediction interval is located, and also has better interpretability.
[0078] Technical effect: Ablation experiments show that adding statistical life cycle features to numerical features improves the model performance by 10.15%. The life cycle feature module combines the trend features of the current and historical stages, enabling relatively accurate prediction of trends in different stages of the life cycle. Control group experiments were conducted on datasets with different topic lengths. After introducing the statistical life cycle feature module, the prediction of long topics has been significantly improved, with the loss reduced by 8.5% (MAE) and 7.2% (RMSE). The improvement of the module for short topics is relatively small, with the loss reduced by only 3.7% (MAE) and 3.5% (RMSE). This shows that the statistical life cycle module not only compensates for the cumulative error problem of long topics but also takes advantage of the richer historical information of long topics. For long topics, the observed historical time span is longer, and the statistical life cycle module can capture more detailed information, which is very beneficial for the current prediction stage.
[0079] In the embodiment of this application, the above-mentioned topic popularity prediction step 102 further includes:
[0080] Time series fusion decoding step: Input the semantic vector learned by the topic semantic encoder and the list topic time series vector learned by the time series feature encoder into the time series fusion decoder for decoding and fusion, learn the semantic and time series relationships among the semantic input of the full life cycle of the list topic, the observable long-term input, and the short-term input that changes with the list topic at any time, and obtain the to-be-predicted topic popularity value according to the learning result.
[0081] The decoder for semantic-temporal fusion of the list topic adopts the decoder structure of TFT, denoted as the semantic-time fusion decoder. It jointly decodes the semantic information of the topic semantic encoder integrated with the list topic and the dynamic time features fused by the time feature encoder, and learns the semantic and temporal relationships among the semantic input of the full life cycle of the list topic, the observable long-term input, and the short-term input that changes with the list topic at any time. This decoder is based on the sequence-to-sequence framework, which integrates the temporal modeling ability of TFT (Temporal Fusion Transformer) and the masked interpretable feature selection mechanism of GRN (Gated Residual Network). Based on the model, at the prediction time step t, it can only be based on the input of time features from 0 to t-1, which conforms to the causality of the time series. Let ξ1 represent the input of the list topic of the decoder.
[0082] In the prediction stage of the list topic, the possible topic popularity values within each predicted list topic range are predicted by different heads in the attention, which increases the prediction robustness. Let y t represent the popularity value predicted for the list topic t.
[0083] In the embodiment of the present application, the above-mentioned topic semantic encoder uses BERT to learn non-temporal semantic information; the time series feature encoder uses LSTM for local time processing and multi-head attention to integrate semantic and time-varying information of different time steps; the time series fusion decoder adopts the decoder structure of TFT.
[0084] As described above, the method of the present invention can be better implemented.
[0085] The method proposed by the present invention overcomes three challenges of the current topic popularity prediction technology, namely, low accuracy, inability to conduct early prediction, and the model being easily outdated over time. By analyzing the data characteristics of the list scenario, a list topic popularity prediction model with the fusion of time series data and non-data data as the input is proposed. This model improves the effect of topic prediction in the list scenario.
[0086] Embodiment 2
[0087] As Figure 5 shown, the embodiment of the present application provides a social media hot list topic popularity prediction system, which adopts the social media hot list topic popularity prediction method as described above. The system includes:
[0088] Prediction model construction module 201: Construct the SELF-TF model for predicting the popularity of list topics. The SELF-TF model is used to fuse time-series data and non-time-series data, integrate topic semantic information, and add periodic features of the topic; the SELF-TF model includes: a semantic time fusion encoder and a time-series fusion decoder; the semantic time fusion encoder further includes: a topic semantic encoder and a time-series feature encoder;
[0089] Topic popularity prediction module 202: Input the integrated topic semantic information and the early life cycle characteristics of the topic into the semantic time fusion encoder. After learning the non-temporal semantic information and integrating the semantics and time-varying information of different time steps, it enters the decoder for decoding and completion, and outputs the predicted topic popularity value.
[0090] In the embodiment of the present application, the above semantic time fusion encoder further includes: a topic semantic encoder and a time-series feature encoder;
[0091] The topic semantic encoder receives the non-time-series data and time-series data of the list topic, forms a semantic-based prediction, and outputs the learned semantic vector; among them, the non-time-series data includes: the semantics, sponsors, types, and introductions of the list topic; the time-series data includes: the discussion volume, like volume, popularity value, and time of the list topic;
[0092] Input the learned semantic vector, historical features, and early cycle features into the time-series feature encoder to integrate the time dynamic change features of the topic. By normalizing and encoding the topic time-series features, learn the topic information of the list topic changing dynamically over time, and output the learned list topic time-series vector, where the topic information includes: the encoding information of the past known time series and the time series information of future list topics.
[0093] In the embodiment of the present application, input the semantic vector learned by the topic semantic encoder and the list topic time-series vector learned by the time-series feature encoder into the time-series fusion decoder for decoding and fusion, learn the semantic and time-series relationship between the semantic input of the entire life cycle of the list topic, the observable long-term input, and the short-term input changing with the list topic at any time, and obtain the predicted topic popularity value according to the learning result.
[0094] In the embodiment of the present application, the above topic semantic encoder uses BERT to learn non-temporal semantic information; the time-series feature encoder uses LSTM for local time processing and multi-head attention to integrate the semantics and time-varying information of different time steps; the time-series fusion decoder uses the decoder structure of TFT.
[0095] Embodiment III
[0096] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method for predicting the topic popularity of a social media hot list are implemented.
[0097] Embodiment 4
[0098] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for predicting the topic popularity of a social media hot list are implemented.
[0099] In addition, the method for predicting the topic popularity of a social media hot list described in the embodiments of the present application can be implemented by an electronic device, such as a computer device. Figure 1 FIG. is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present application. Figure 6 FIG. is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present application.
[0100] In some of these embodiments, the computer device may further include a communication interface 83 and a bus 80. Among them, as shown in FIG., the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 and complete communication with each other. Figure 6 As shown, the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 and complete communication with each other.
[0101] Specifically, the above-mentioned processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.
[0102] The memory 82 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81.
[0103] The processor 81 reads and executes the computer program instructions stored in the memory 82 to implement any one of the methods for predicting the topic popularity of a social media hot list in the above embodiments.
[0104] The technical features of the above-mentioned embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0105] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for predicting the popularity of topics on a social media hot list, characterized in that, The method includes: Prediction model construction step: Construct a list topic popularity prediction SELF-TF model, which is used to fuse time series data and non-time series data, integrate topic semantic information and add periodic features of the topic; the SELF-TF model includes: a semantic time fusion encoder and a time series fusion decoder; the semantic time fusion encoder further includes: a topic semantic encoder and a time series feature encoder; Topic popularity prediction step: Input the integrated topic semantic information and the early life cycle features of the topic into the semantic time fusion encoder. After learning the non-temporal semantic information and integrating the semantic and time-varying information of different time steps, enter the decoder for decoding to complete, and output the topic popularity value to be predicted.
2. The method for predicting the topic popularity of a social media hot list according to claim 1, wherein The topic popularity prediction step includes: Semantic encoding step: Receive the non-time series data and time series data of the list topic through the topic semantic encoder to form a semantic-based prediction and output the learned semantic vector; among them, the non-time series data includes: the semantics, initiator, type and introduction of the list topic; the time series data includes: the discussion volume, like volume, popularity value and time of the list topic; Time series feature encoding step: Input the learned semantic vector, historical features and early cycle features into the time series feature encoder to integrate the time dynamic change features of the topic. By normalizing and encoding the topic time series features, learn the topic information of the list topic changing dynamically over time, and output the learned list topic time series vector, where the topic information includes: the time series information encoding information of the past known time series and the future list topic.
3. The method for predicting the topic popularity of a social media hot list according to claim 1, wherein The topic popularity prediction step further includes: Time series fusion decoding step: Input the semantic vector learned by the topic semantic encoder and the list topic time series vector learned by the time series feature encoder into the time series fusion decoder for decoding and fusion, learn the semantic and time series relationship between the semantic input of the full life cycle of the list topic, the observable long-term input and the short-term input changing with the list topic at any time, and obtain the topic popularity value to be predicted according to the learning result.
4. The method for predicting the topic popularity of a social media hot list according to claim 1, wherein, The topic semantic encoder uses BERT to learn non-temporal semantic information; the time series feature encoder uses LSTM for local time processing and multi-head attention to integrate the semantic and time-varying information of different time steps; the time series fusion decoder uses the decoder structure of TFT.
5. A social media popular list topic popularity prediction system, adopting the social media popular list topic popularity prediction method described in any one of claims 1-4, characterized in that, The system includes: Prediction model construction module: Construct a list topic popularity prediction SELF-TF model, which is used to fuse time series data and non-time series data, integrate topic semantic information and add periodic features of the topic; the SELF-TF model includes: a semantic time fusion encoder and a time series fusion decoder; the semantic time fusion encoder further includes: a topic semantic encoder and a time series feature encoder; Topic popularity prediction module: Integrate topic semantic information and early life cycle characteristics of the topic, input them into the semantic temporal fusion encoder, learn non-temporal semantic information and integrate semantic and time-varying information of different time steps, then enter the decoder for decoding to complete, and output the predicted topic popularity value.
6. The social media popular list topic popularity prediction system according to claim 5, characterized in that, The semantic temporal fusion encoder further includes: a topic semantic encoder and a temporal feature encoder; The topic semantic encoder receives non-temporal data and temporal data of the list topic, forms a semantic-based prediction, and outputs the learned semantic vector; among them, the non-temporal data includes: the semantics, initiator, type, and profile of the list topic; the temporal data includes: the discussion volume, like volume, popularity value, and time of the list topic; Input the learned semantic vector, historical features, and early cycle features into the temporal feature encoder, integrate the time dynamic change features of the topic, and learn the topic information of the list topic changing dynamically over time by normalizing and encoding the topic temporal features, and output the learned list topic temporal vector, where the topic information includes: the encoded information of the past known time series and the time series information of the future list topic.
7. The method for predicting the topic popularity of the social media hot list according to claim 5, wherein, Input the semantic vector learned by the topic semantic encoder and the list topic temporal vector learned by the temporal feature encoder into the temporal fusion decoder for decoding and fusion, learn the semantic and temporal relationship between the semantic input of the full life cycle of the list topic, the observable long-term input, and the short-term input changing with the list topic at any time, and obtain the predicted topic popularity value according to the learning result.
8. The method for predicting the topic popularity of a social media hot list according to claim 5, wherein The topic semantic encoder uses BERT to learn non-temporal semantic information; the temporal feature encoder uses LSTM for local time processing and multi-head attention to integrate semantic and time-varying information of different time steps; the temporal fusion decoder uses the decoder structure of TFT.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method for predicting the popularity of social media hot list topics described in any one of claims 1-7.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for predicting the popularity of social media hot list topics described in any one of claims 1 to 7.
Citation Information
Cited By
Cross-topic rumor detection method based on time perception attention mechanism
CN120725025A