A method, apparatus, device, and readable storage medium for tracking follow-up news.

CN115858610BActive Publication Date: 2026-08-14AGRICULTURAL BANK OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-02
Publication Date
2026-08-14

AI Technical Summary

Benefits of technology

[0049]由上述技术方案可以看出,本申请实施例提供的后续新闻的追踪方法、装置、设备及可读存储介质,基于待追踪热点新闻的热度指标,构建多元逻辑回归模型,多元逻辑回归模型用于基于实时输入的热度指标,输出热度概率,获取后续新闻簇,计算后续新闻簇的热度指标。将后续新闻簇的热度指标输入至多元逻辑回归模型,获取多元逻辑回归模型输出的热度概率。进一步当后续新闻簇满足预设条件时,将后续新闻簇作为待追踪热点新闻的后续热点新闻簇,其中,热度指标包括传播度指标、发布者权威度指标、以及丰富度指标,也即热度指标能够指示新闻在多个维度上的热门程度。热度概率用于指示作为输入的热度指标对应的新闻处于爆发状态的概率。也即,本申请得到的热度概率结果能够指示后续新闻簇处于爆发状态的概率,又由于,预设条件包括热度概率大于预设的概率阈值,因此,本申请的后续新闻追踪结果包括处于爆发状态的概率大于预设概率阈值的后续新闻簇,由于,多元逻辑回归模型基于待追踪热点新闻构建,提高了后续新闻簇的热度概率的准确性,实现有针对性地追踪不同热点新闻。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115858610B_ABST
    Figure CN115858610B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and readable storage medium for tracking subsequent news. Based on the popularity index of the hot news to be tracked, a multivariate logistic regression model is constructed. This model is used to obtain subsequent news clusters by outputting popularity probabilities based on real-time input popularity indices and calculating the popularity index of each cluster. The popularity index of the subsequent news clusters is input into the multivariate logistic regression model to obtain the popularity probabilities output by the model. Furthermore, when a subsequent news cluster meets preset conditions, it is designated as a subsequent hot news cluster of the hot news to be tracked. The popularity index indicates the degree of popularity of news across multiple dimensions. The popularity probability indicates the probability that the news corresponding to the input popularity index is in a state of explosive growth. Therefore, this application improves the accuracy of the popularity probability of subsequent news clusters, enabling targeted tracking of different hot news.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data mining technology, and in particular to a method, apparatus, device, and readable storage medium for tracking subsequent news. Background Technology

[0002] In today's booming internet era, especially in the Web 2.0 era, various social media platforms provide every user with a channel to voice their opinions. Social media platforms generate massive amounts of news every day, and a large number of trending news stories break out in real time. Furthermore, as events unfold, trending news stories may continue to ferment and even reverse. Therefore, how to achieve targeted tracking of the follow-up news of trending news is an issue that urgently needs to be addressed. Summary of the Invention

[0003] This application provides a method, apparatus, device, and readable storage medium for tracking follow-up news, as follows:

[0004] A method for tracking follow-up news includes:

[0005] Obtain the popularity metrics of trending news to be tracked, including dissemination metrics, publisher authority metrics, and richness metrics;

[0006] Based on the popularity index of the hot news to be tracked, a multivariate logistic regression model is constructed. The multivariate logistic regression model is used to output the popularity probability based on the real-time input popularity index. The popularity probability is used to indicate the probability that the news corresponding to the input popularity index is in an explosive state.

[0007] Obtain subsequent news clusters, which are news sets obtained by clustering and include subsequent news of multiple hot news items to be tracked;

[0008] Calculate the popularity index of the subsequent news cluster;

[0009] The popularity index of the subsequent news cluster is input into the multivariate logistic regression model to obtain the popularity probability output by the multivariate logistic regression model.

[0010] Based on the popularity probability, it is determined whether the subsequent news cluster meets the preset conditions, the preset conditions including the popularity probability being greater than a preset probability threshold;

[0011] If so, the subsequent news clusters will be regarded as the subsequent hot news clusters of the hot news to be tracked.

[0012] Optional metrics for obtaining trending news topics to be tracked include:

[0013] Extract keywords from the trending news stories to be tracked;

[0014] Based on the keywords of the trending news to be tracked, obtain similar news articles.

[0015] Obtain historical hot news clusters, which include the hot news to be tracked and multiple similar news items to the hot news to be tracked;

[0016] The popularity characteristics of each news item in the historical hot news cluster are obtained, including the number of reads, the number of comments, the number of authoritative media, the number of top media, the number of mid-tier media, the number of ordinary media, and the richness of content;

[0017] Based on the popularity characteristics of each news item in the historical hot news cluster, the popularity index of the historical hot news cluster is calculated.

[0018] The popularity index of the historical hot news clusters is used as the popularity index for obtaining the hot news to be tracked.

[0019] Optionally, retrieve subsequent news clusters, including:

[0020] Monitor real-time news, obtain keywords from the real-time news, and obtain a set of keywords for the real-time news;

[0021] The similarity of the keyword set of the real-time news is compared with the keyword set of the historical hot news cluster to determine whether the keyword set of the historical hot news cluster is a subset of the keyword set of the real-time news.

[0022] Determine whether the text similarity between the real-time news and the trending news to be tracked is within a preset similarity range;

[0023] If the real-time news meets the preset subsequent determination conditions, then the real-time news is determined to be a subsequent news of the hot news to be tracked, and the real-time news is added to the subsequent news set. The subsequent determination conditions include that the keyword set of the historical hot news cluster is a subset of the keyword set of the real-time news, and the text similarity between the real-time news and the hot news to be tracked is within the similarity range.

[0024] Based on text similarity, the news in the subsequent news set is clustered to obtain at least one subsequent news cluster.

[0025] Optionally, the popularity index of the subsequent news cluster is calculated, including:

[0026] The popularity characteristics of each news item in the subsequent news cluster are acquired in real time, and the popularity index of the subsequent news cluster is calculated based on the popularity characteristics of each news item in the subsequent news cluster.

[0027] Optionally, based on the popularity characteristics of each news item in the target news cluster, a popularity index for the target news cluster is calculated. The target news cluster includes the subsequent news cluster and the historical hot news cluster, including:

[0028] The total number of reads for each news item in the target news cluster is summed to obtain the total number of reads in real time. The dissemination index of the target news cluster is obtained based on the growth rate of the total number of reads in real time. The dissemination index is positively correlated with the growth rate of the total number of reads in real time.

[0029] Based on the number of authoritative media, top media, mid-tier media, and ordinary media for each news item in the target news cluster, the total number of authoritative media, top media, mid-tier media, and ordinary media for the target news cluster is obtained. Based on the growth rate of the total number of authoritative media, the growth rate of the total number of top media, the growth rate of the total number of mid-tier media, and the growth rate of the total number of ordinary media, the publisher authority index of the target news cluster is calculated.

[0030] Based on the number of comments for each news item in the target news cluster, the total number of comments for the target news cluster over at least one time period is obtained. The average content richness of the target news cluster is obtained by averaging the content richness of each news item in the target news cluster. Based on the growth rate of the total number of comments and the average content richness over the at least one time period, the richness index of the target news cluster is calculated.

[0031] Optionally, the preset conditions also include:

[0032] The popularity index of the subsequent news cluster is greater than the popularity index of the hot news to be tracked.

[0033] And / or, the popularity index of the subsequent news cluster is greater than the preset popularity index threshold.

[0034] Optionally, this method also includes:

[0035] Obtain the keyword set of the subsequent hot news cluster as the subsequent hot word set;

[0036] The trending news to be tracked, the subsequent cluster of trending news, and the subsequent set of trending keywords are associated.

[0037] A follow-up news tracking device includes:

[0038] The first indicator calculation unit is used to obtain the popularity indicators of the hot news to be tracked. The popularity indicators include the dissemination index, the authoritative index of the publisher, and the richness index.

[0039] The model building unit is used to build a multivariate logistic regression model based on the popularity index of the hot news to be tracked. The multivariate logistic regression model is used to output the popularity probability based on the real-time input popularity index. The popularity probability is used to indicate the probability that the news corresponding to the input popularity index is in an explosive state.

[0040] The news monitoring unit is used to acquire subsequent news clusters, which are news sets obtained by clustering and include subsequent news of multiple hot news items to be tracked.

[0041] The second indicator calculation unit is used to calculate the popularity index of the subsequent news cluster.

[0042] The model application unit is used to input the popularity index of the subsequent news cluster into the multivariate logistic regression model and obtain the popularity probability output by the multivariate logistic regression model.

[0043] A condition judgment unit is used to determine whether the subsequent news cluster meets a preset condition based on the popularity probability, wherein the preset condition includes the popularity probability being greater than a preset probability threshold.

[0044] The subsequent hotspot determination unit is used to, if so, identify the subsequent news cluster as the subsequent hotspot news cluster of the hotspot news to be tracked.

[0045] A follow-up news tracking device, comprising: a memory and a processor;

[0046] The memory is used to store programs;

[0047] The processor is used to execute the program to implement the various steps of the subsequent news tracking method.

[0048] A readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the various steps of a method for tracking subsequent news.

[0049] As can be seen from the above technical solutions, the follow-up news tracking method, apparatus, device, and readable storage medium provided in this application embodiment construct a multivariate logistic regression model based on the popularity index of the hot news to be tracked. The multivariate logistic regression model is used to output popularity probability based on the real-time input popularity index, obtain follow-up news clusters, and calculate the popularity index of the follow-up news clusters. The popularity index of the follow-up news clusters is input into the multivariate logistic regression model to obtain the popularity probability output by the multivariate logistic regression model. Furthermore, when the follow-up news clusters meet preset conditions, the follow-up news clusters are used as follow-up hot news clusters of the hot news to be tracked. The popularity index includes the dissemination index, the publisher authority index, and the richness index, that is, the popularity index can indicate the popularity of news in multiple dimensions. The popularity probability is used to indicate the probability that the news corresponding to the input popularity index is in an explosive state. In other words, the heat probability results obtained by this application can indicate the probability that subsequent news clusters are in an outbreak state. Since the preset conditions include a heat probability greater than a preset probability threshold, the subsequent news tracking results of this application include subsequent news clusters with a probability greater than the preset probability threshold that are in an outbreak state. Since the multivariate logistic regression model is constructed based on the hot news to be tracked, it improves the accuracy of the heat probability of subsequent news clusters and enables targeted tracking of different hot news. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 A flowchart illustrating a specific implementation of a method for tracking subsequent news provided in this application embodiment;

[0052] Figure 2 A schematic diagram of the structure of a follow-up news tracking system provided in this application embodiment;

[0053] Figure 3 A flowchart illustrating a method for tracking subsequent news provided in an embodiment of this application;

[0054] Figure 4 A schematic diagram of the structure of a follow-up news tracking device provided in this application embodiment;

[0055] Figure 5 This is a schematic diagram of the structure of a follow-up news tracking device provided in an embodiment of this application. Detailed Implementation

[0056] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0057] The method provided in this application can be applied to scenarios where it is necessary to track follow-up news on historical hot news. Specifically, it can be applied to track various historical hot news on social media platforms, monitor follow-up news reports of historical hot news, and monitor whether follow-up news reports generate hot news. In this application, the situation in which follow-up news of historical hot news generates hot news is referred to as the "secondary outbreak" of historical hot news.

[0058] For example, news item N0, published on September 30, 2022, stating "City A will host a large-scale expo on October 30, 2022," became a high-profile news story, i.e., a historical hot news item. News item N1, published on October 31, 2022, stating "The first day of the large-scale expo in City A saw 20,000 visitors," and news item N2, stating "More than 200 valuable exhibits were displayed on the first day of the large-scale expo in City A," were also published. News items N1 and N2 are both follow-up reports to historical news item N0. By analyzing the popularity curves of historical hot news items, we can identify whether news items N1 and / or N2 have become hot topics. If so, we can link the hot topics of news items N1 and / or N2 with historical hot news items to form a hot news chain.

[0059] Furthermore, this solution can be applied to smart devices, such as computers, tablets, or smartphones. Additionally, this solution can also be applied to servers. Next, in conjunction with the appendix... Figure 1 This section describes the methods for tracking subsequent news related to this application, such as... Figure 1 This example demonstrates a flowchart of a method for tracking follow-up news, which includes the following details:

[0060] S101, Get the trending news to be tracked.

[0061] In this embodiment, trending news refers to news with a popularity value greater than a preset popularity threshold. The popularity value of the news is determined based on popularity features, which include, but are not limited to, the number of reads, the number of comments, and the number of reposts. For details, please refer to the prior art.

[0062] S102. Based on the hot news to be tracked, obtain historical hot news clusters.

[0063] In this embodiment, the historical hot news cluster includes hot news to be tracked and at least one similar news item to the hot news to be tracked.

[0064] In this embodiment, the method for obtaining similar news to the trending news to be tracked includes:

[0065] 1. Preprocess the text of the hot news to be tracked, and use the trained BERT model (a bidirectional structure model pre-built based on Transformer Encoder) to extract the keywords of the hot news to be tracked and the keywords of other news.

[0066] 2. Other news items whose keywords match the keywords of the trending news to be tracked with a keyword degree greater than a preset threshold are selected as candidate news items.

[0067] 3. Sort the candidate news according to the popularity value from high to low, and take the top n candidate news as similar news to the hot news to be tracked, and add them to the historical hot news cluster.

[0068] S103. Obtain the popularity characteristics and keyword set of historical hot news clusters.

[0069] In this embodiment, the method for obtaining the popularity characteristics of historical hot news clusters includes:

[0070] 1. Based on the news source (called news ID) of each news item in the historical hot news cluster, obtain the popularity characteristics of each news item.

[0071] In this embodiment, popularity characteristics include readership, number of comments, number of reposts, number of authoritative media outlets publishing the news, number of top-tier media outlets, number of mid-tier media outlets, number of ordinary media outlets, and the richness of the news content.

[0072] 2. Based on the popularity characteristics of each news item in the historical hot news cluster, obtain the popularity characteristics of the historical hot news cluster.

[0073] In this embodiment, the readership, comment volume, repost volume, number of authoritative media outlets, number of top-tier media outlets, number of mid-tier media outlets, and number of ordinary media outlets publishing the news in the historical hot news cluster are all obtained by summing or averaging the popularity characteristics of each news item. The richness of news content is obtained by averaging the popularity characteristics of each news item.

[0074] In this embodiment, the keyword set of the historical hot news cluster includes the keywords of the news with the highest popularity value in the historical hot news cluster.

[0075] It should be noted that the system obtains the popularity characteristics of each news item in the historical hot news cluster at various points in time within a preset historical time period through a real-time update service based on preset popularity features. The historical time period is from the time the news was published until the popularity of the hot news item to be tracked reaches its highest value. See existing technology for details.

[0076] S104. Calculate the popularity index of historical hot news clusters based on the popularity characteristics of historical hot news clusters.

[0077] In this embodiment, the popularity indicators include dissemination indicators, publisher authority indicators, and richness indicators.

[0078] Specifically, the dissemination index Used to indicate the breadth (i.e., extent) of dissemination of historical hot news clusters, optionally, it can be obtained by calculating the growth rate of readership.

[0079] Publisher Authority Index This term is used to indicate the news value of historical hot news clusters from the perspective of the authority level of the news disseminator. It's understandable that, based on summarizing historical news dissemination data, one can conclude that the higher the authority of the publisher, or the greater the number of publishers, the higher the news value of their published news. Optionally, This ranking is calculated based on the combined growth rate of reposts from the number of authoritative media outlets, top-tier media outlets, mid-tier media outlets, and ordinary media outlets that publish news. It should be noted that the authority level of the news disseminator is determined by analyzing its registration information; for example, official accounts have a higher authority level than private accounts.

[0080] Richness index This can be determined by the richness of news content and the number of comments. For example, the smaller the proportion of keywords in the news content, the higher the richness of the news content.

[0081] In this embodiment, the specific method for calculating the popularity index of historical hot news clusters is as follows:

[0082] r represents the number of reads. Used to represent growth rate calculations, for example, , and These represent the number of reads at different times, with % indicating a percentage conversion.

[0083] ,in, The number of authoritative media outlets, For the number of top media outlets, For the number of mid-tier media outlets, This refers to the number of regular users.

[0084] ,in, This represents the number of discussions in the last hour. This represents the number of discussions in the last 6 hours. This indicates the richness of news content, among which, This indicates the proportion of keywords in the news content, that is... This indicates the richness of news content. The function, specifically, and Inverse correlation. In this embodiment, the discussion volume can be the comment volume. For specific calculation methods, please refer to existing technologies.

[0085] S105. Construct a popularity trend curve and build a multivariate logistic regression model based on the popularity trend curve.

[0086] In this embodiment, the independent variables of the multivariate logistic regression model are the dissemination index, the publisher authority index, and the richness index, while the dependent variable is the popularity probability P. The popularity probability indicates the probability that the news (a single news item or a cluster of news items) corresponding to the popularity index as input is in a state of explosion.

[0087] In this embodiment, the model is as follows:

[0088]

[0089] Where y represents the set of popularity indicators, This indicates that at time j, the set of popularity indicators includes three popularity indicators, which are respectively equal to... The probability of popularity under certain circumstances For preset model parameters, These represent three popularity indicators; the calculation method is described in the steps above.

[0090] It should be noted that a multivariate logistic regression model was constructed based on the popularity trend curve, considering the news's dissemination, the publisher's authority, and the richness of information. Before constructing the model, data relevance verification was performed. The model passed the verification through goodness-of-fit test and parameter validation analysis. For specific model construction methods, please refer to existing technologies.

[0091] S106. Monitor real-time news and determine whether the real-time news belongs to the follow-up news of the hot news to be tracked. If so, add the real-time news to the follow-up news set.

[0092] In this embodiment, the method for determining whether real-time news belongs to the follow-up news of the hot news to be tracked includes:

[0093] 1. The trained BERT model extracts keywords from real-time news to obtain a keyword set for real-time news.

[0094] 2. Compare the similarity between the keyword set of real-time news and the keyword set of historical hot news clusters to determine whether the keyword set of historical hot news clusters is a subset of the keyword set of real-time news.

[0095] In this embodiment, the DFA algorithm is used to determine the similarity of keywords in different keyword sets. If the similarity is greater than the preset similarity threshold, the keywords are determined to be the same.

[0096] 3. If so, calculate the text similarity between real-time news and the trending news to be tracked, and check whether the text similarity is within the preset similarity range.

[0097] In this embodiment, doc2vec is used to obtain the vector representations of real-time news and trending news to be tracked, respectively. The cosine similarity distance between the vector representations of real-time news and trending news to be tracked is calculated, and the text similarity between real-time news and trending news to be tracked is obtained based on the cosine similarity distance. The preset similarity range is [0.65, 0.8].

[0098] 4. If so, then the real-time news is confirmed to be a follow-up to the hot news to be tracked.

[0099] It is evident that this application considers whether real-time news meets the characteristics of subsequent news from both keyword and full-text perspectives, resulting in a more accurate judgment of subsequent news.

[0100] S107. Cluster the news in the subsequent news set to obtain at least one subsequent news cluster.

[0101] In this embodiment, news in the subsequent news set is clustered based on the text similarity between subsequent news items. For specific methods, please refer to the prior art.

[0102] S108. Obtain the keyword set and popularity characteristics of subsequent news clusters.

[0103] In this embodiment, the method for obtaining the keyword set and popularity characteristics of subsequent news clusters is described in S103.

[0104] S109. Based on the popularity characteristics of subsequent news clusters, obtain the popularity index of subsequent news clusters.

[0105] For specific calculation methods of the popularity index, please refer to the steps above.

[0106] S110. Based on the popularity index of subsequent news clusters, the popularity probability output by the multivariate logistic regression model is obtained as the popularity probability of subsequent news clusters.

[0107] S111. If the probability of popularity of subsequent news clusters is greater than the preset probability threshold, the subsequent news clusters are determined to be hot news clusters.

[0108] S112. If the subsequent news cluster is determined to be a hot news cluster, obtain the subsequent hot words of the hot news to be tracked based on the keyword set of the subsequent news cluster.

[0109] In this embodiment, keywords that exist in the keyword set of subsequent news clusters but do not exist in the keyword set of historical hot news clusters are selected as subsequent hot keywords for the hot news to be tracked.

[0110] As can be seen, the subsequent hot keywords of the hot news to be tracked are used to represent the popular topics in the subsequent news of the hot news to be tracked. For example, the subsequent hot keywords of news No. 0, "City A will hold a large-scale expo on October 30, 2022," include visitor traffic, indicating that the popular topic in the subsequent news of news No. 0 is visitor traffic. The cluster of subsequent news with "visitor traffic" as the keyword has become a hot news cluster. In layman's terms, news No. 0 has experienced a "second boom" due to the topic of visitor traffic.

[0111] S113, related trending news, subsequent trending keywords, and subsequent news clusters.

[0112] In this embodiment, the relationships between the trending news to be tracked, subsequent trending keywords, and subsequent news clusters are stored and displayed in a preset format. For example, the display interface of the trending news to be tracked adds the display of subsequent trending keywords and jump links to subsequent news clusters.

[0113] As can be seen from the above technical solutions, the follow-up news tracking method provided in this application identifies the follow-up news clusters of the hot news to be tracked by monitoring and analyzing the keywords of real-time news. Furthermore, it determines whether the follow-up news cluster has become a hot topic by using the real-time popularity characteristics of the follow-up news cluster and a multivariate logistic regression model. The multivariate logistic regression model is constructed based on the popularity trend curve of the preceding news (i.e., the hot news to be tracked) of the follow-up news cluster. Therefore, the judgment result on whether the follow-up news cluster has become a hot topic is more accurate. It can be seen that this application not only realizes the real-time tracking of follow-up news, but also judges the hot news in the follow-up news based on the historical popularity trend curve of the hot news to be tracked. Combining the correlation between the follow-up news and the preceding news, it achieves targeted tracking of the follow-up news of different hot news.

[0114] Understandably, the solution provided in this application differs from existing methods for mining trending news. Existing methods typically determine whether a news story is in an explosive state based on its real-time popularity characteristics. However, it's clear that different categories of news exhibit different trends in their explosive states; for example, sports news and entertainment news show distinct differences in their explosive states. Furthermore, the explosive state of subsequent news stories is correlated with that of preceding news stories. Therefore, this application constructs a judgment model for each trending news story to determine the explosive state of subsequent news stories, offering highly targeted and accurate results, and making the subsequent news stories more valuable for tracking. Furthermore, a multivariate logistic regression model is constructed by combining three dimensions of popularity indicators: dissemination, publisher authority, and richness. These indicators are calculated using real-time monitored popularity characteristics, indicating the popularity of news across different dimensions, thus further improving the accuracy and objectivity of the judgment results.

[0115] Furthermore, by displaying related trending news, subsequent trending keywords, and subsequent news clusters, the tracking trajectory of trending news is visualized, making it more convenient for users to follow the development of news events and improving the user experience.

[0116] In summary, this invention effectively solves the problem of real-time tracking of trending topics by tracking the follow-up news reported by the media. It models the popularity of subsequent news by referencing the popularity trends of preceding news before their explosive emergence, thus improving the accuracy of identifying subsequent trending news.

[0117] It should be noted that, Figure 1 This application only provides one possible implementation of a follow-up news tracking method. Other possible implementation methods are also included, such as:

[0118] In an optional embodiment, when calculating the popularity index of historical hot news clusters based on the popularity characteristics of historical hot news clusters in S104, the popularity index may also include the commenter participation index.

[0119] Commentator engagement metrics are used to indicate the news value of historical hot news clusters from the perspective of news audiences. It can be understood that, based on the dissemination data of historical news, it can be concluded that the wider the range of news commentators (distinguishing account users (discussers) by occupation, years of account registration, age of account registration, etc.), or the richer the comment content, the higher the news value of the comment.

[0120] Furthermore, this embodiment does not limit the calculation method for each heat index. Figure 1 This only illustrates one specific method for calculating an optional popularity metric.

[0121] In an optional embodiment, the method further includes obtaining a timeout period based on the type of the trending news to be tracked, executing the method within the timeout period, and ignoring the trending news after the timeout period is reached.

[0122] In an optional embodiment, after determining that the probability of the popularity of the subsequent news cluster is greater than a preset probability threshold, S111 further includes determining whether the popularity index of the subsequent news cluster is greater than the popularity index of the historical hot news cluster or whether it is greater than a preset index value threshold. If so, the subsequent news cluster is determined to be a hot news cluster.

[0123] In an optional embodiment, the method further includes: unified management of trending news. For example:

[0124] 1. Store the keyword set, news popularity value and corresponding multivariate logistic regression model of historical hot news clusters into the previous news database for unified management and maintenance.

[0125] 2. Store subsequent news items in the subsequent news collection into the subsequent news monitoring queue for unified management and maintenance.

[0126] 3. Subsequent news clusters identified as hot news clusters will be stored in the news hot news cluster database for unified management and maintenance.

[0127] It should be noted that each news database is managed based on a corresponding preset timeout duration, and data exceeding the timeout duration is deleted from the news database in real time.

[0128] Figure 2 This paper illustrates the structural diagram of a follow-up news tracking system. Based on the principle that follow-up and preceding news stories share textual similarities and exhibit a certain correlation in terms of user attention, this application constructs a follow-up news tracking system. Figure 2 As shown, the tracking system 20 includes a hot news management module 201, a follow-up news identification module 202, a text aggregation processing module 203, and an outbreak determination module 204.

[0129] like Figure 2 As shown, the Hot News Management module, after acquiring the hot news to be tracked, calls the keyword extraction service to obtain a keyword set and the similar news query service to obtain similar news, thereby obtaining historical hot news clusters. It then calls the real-time popularity feature update service to obtain the popularity features of these historical hot news clusters. Based on these features, it calculates the popularity index of the historical hot news clusters, constructs a popularity trend curve, builds a multivariate logistic regression model based on the trend curve, and stores the multivariate logistic regression model, keyword set, and other information in the previous news database. The specific functions of the Hot News Management module are described in the above embodiment.

[0130] like Figure 2As shown, the follow-up news identification module obtains real-time news, calls the keyword comparison service to obtain the keyword set of current events news, and calls the similarity comparison service to perform keyword comparison and full-text comparison on the real-time news and historical hot news clusters in the previous news database to determine whether the real-time news is a follow-up news of the hot news to be tracked. The follow-up news and the keyword set of the follow-up news are then stored in the follow-up news monitoring queue. For the specific implementation process of the follow-up news identification module, please refer to the above embodiment.

[0131] like Figure 2 As shown, the text aggregation processing module calls the similarity comparison service to cluster subsequent news in the subsequent news monitoring queue to obtain subsequent news clusters, and stores the subsequent news clusters in the subsequent news cluster database. It should be noted that the text aggregation processing module can extract subsequent news from the subsequent news monitoring queue periodically, or it can extract subsequent news one by one. The clustering methods include full clustering, or determining whether newly added subsequent news should be added to the existing subsequent news clusters. The functional implementation process of the text aggregation processing module is described in the above embodiment.

[0132] like Figure 2 As shown, the outbreak determination module calls the real-time update service for popularity features to obtain the popularity features of subsequent news in the subsequent news cluster database. It then obtains the popularity index for each subsequent news cluster, and subsequently calls the multivariate logistic regression model built based on the popularity trend curve in the preceding news database to obtain the popularity probability of the multivariate logistic regression model. Based on the popularity probability, it determines whether the subsequent news cluster of the hot news to be tracked is in an outbreak state. Furthermore, it outputs the subsequent news clusters belonging to the hot news and the keyword set of the subsequent news clusters. The functional implementation process of the outbreak determination module is described in the above embodiment.

[0133] It should be noted that the aforementioned preceding news database, subsequent news monitoring queue, and news cluster database are all equipped with timed devices to delete timed-out data; see existing technologies for details.

[0134] In summary, the follow-up news tracking method provided in this application can be summarized as follows: Figure 2 The process shown is as follows: Figure 3 As shown, this method includes:

[0135] S301. Obtain the popularity index of trending news to be tracked.

[0136] In this embodiment, the popularity indicators include dissemination indicators, publisher authority indicators, and richness indicators.

[0137] Optionally, methods for obtaining the popularity metrics of trending news to be tracked include:

[0138] The popularity index is calculated directly based on the popularity characteristics of the trending news at various times in the historical period.

[0139] Alternatively, the popularity index of the historical hot news cluster can be calculated based on the popularity characteristics of each news item in the news cluster (i.e., the historical hot news cluster) to which the hot news to be tracked belongs, and then the popularity index of the historical hot news cluster can be used as the popularity index of the hot news to be tracked.

[0140] For a specific method of obtaining one of the optional popularity indicators, please refer to the above embodiment.

[0141] S302. Based on the popularity index of the hot news to be tracked, construct a multivariate logistic regression model.

[0142] In this embodiment, a multivariate logistic regression model is used to output a popularity probability based on a real-time input popularity index. The popularity probability indicates the probability that the news corresponding to the input popularity index is in a state of explosive growth. An optional method for constructing a multivariate logistic regression model can be found in the above embodiment.

[0143] S303. Obtain subsequent news clusters and calculate the popularity index of subsequent news clusters.

[0144] In this embodiment, the subsequent news cluster is a news set obtained through clustering, which includes subsequent news from multiple trending news stories to be tracked. A specific method for obtaining an optional subsequent news cluster can be found in the above embodiment.

[0145] S304. Input the popularity index of subsequent news clusters into the multivariate logistic regression model to obtain the popularity probability output by the multivariate logistic regression model.

[0146] In this embodiment, the popularity index of subsequent news clusters is input as the independent variable into the multivariate logistic regression model in real time. The multivariate logistic regression model obtains the probability that the subsequent news clusters corresponding to the input popularity index are in an explosive state at the current moment through fitting.

[0147] S305. Determine whether subsequent news clusters meet preset conditions based on popularity probability.

[0148] In this embodiment, the preset conditions include a heat probability greater than a preset probability threshold.

[0149] S306. If so, treat the subsequent news clusters as subsequent hot news clusters of the hot news to be tracked.

[0150] As can be seen from the above technical solution, the follow-up news tracking method provided in this application embodiment constructs a multivariate logistic regression model based on the popularity index of the hot news to be tracked. The multivariate logistic regression model is used to output popularity probability based on the real-time input popularity index, obtain follow-up news clusters, and calculate the popularity index of the follow-up news clusters. The popularity index of the follow-up news clusters is input into the multivariate logistic regression model to obtain the popularity probability output by the multivariate logistic regression model. Furthermore, when the follow-up news clusters meet preset conditions, the follow-up news clusters are used as follow-up hot news clusters of the hot news to be tracked. The popularity index includes the dissemination index, the publisher authority index, and the richness index, that is, the popularity index can indicate the popularity of the news in multiple dimensions. The popularity probability is used to indicate the probability that the news corresponding to the input popularity index is in an explosive state. In other words, the heat probability results obtained by this application can indicate the probability that subsequent news clusters are in an outbreak state. Since the preset conditions include a heat probability greater than a preset probability threshold, the subsequent news tracking results of this application include subsequent news clusters with a probability greater than the preset probability threshold that are in an outbreak state. Since the multivariate logistic regression model is constructed based on the hot news to be tracked, it improves the accuracy of the heat probability of subsequent news clusters and enables targeted tracking of different hot news.

[0151] Figure 4 This application provides a schematic diagram of the structure of a follow-up news tracking device according to an embodiment of the present application. Figure 4 As shown, the device may include:

[0152] The first indicator calculation unit 401 is used to obtain the popularity index of the hot news to be tracked. The popularity index includes the dissemination index, the authoritative index of the publisher, and the richness index.

[0153] The model building unit 402 is used to build a multivariate logistic regression model based on the popularity index of the hot news to be tracked. The multivariate logistic regression model is used to output the popularity probability based on the real-time input popularity index. The popularity probability is used to indicate the probability that the news corresponding to the input popularity index is in an outbreak state.

[0154] The news monitoring unit 403 is used to acquire subsequent news clusters, wherein the subsequent news clusters are news sets obtained by clustering and include multiple subsequent news of the hot news to be tracked;

[0155] The second indicator calculation unit 404 is used to calculate the popularity index of the subsequent news cluster.

[0156] The model application unit 405 is used to input the popularity index of the subsequent news cluster into the multivariate logistic regression model and obtain the popularity probability output by the multivariate logistic regression model.

[0157] The condition judgment unit 406 is used to judge whether the subsequent news cluster meets the preset conditions based on the popularity probability, the preset conditions including the popularity probability being greater than a preset probability threshold;

[0158] The subsequent hotspot determination unit 407 is used to, if so, identify the subsequent news cluster as the subsequent hotspot news cluster of the hotspot news to be tracked.

[0159] Optionally, the first indicator calculation unit is used to obtain the popularity index of the trending news to be tracked, including: The first indicator calculation unit is specifically used for:

[0160] Extract keywords from the trending news stories to be tracked;

[0161] Based on the keywords of the trending news to be tracked, obtain similar news articles.

[0162] Obtain historical hot news clusters, which include the hot news to be tracked and multiple similar news items to the hot news to be tracked;

[0163] The popularity characteristics of each news item in the historical hot news cluster are obtained, including the number of reads, the number of comments, the number of authoritative media, the number of top media, the number of mid-tier media, the number of ordinary media, and the richness of content;

[0164] Based on the popularity characteristics of each news item in the historical hot news cluster, the popularity index of the historical hot news cluster is calculated.

[0165] The popularity index of the historical hot news clusters is used as the popularity index for obtaining the hot news to be tracked.

[0166] Optionally, the news monitoring unit is used to acquire subsequent news clusters, including: The news monitoring unit is specifically used for:

[0167] Monitor real-time news, obtain keywords from the real-time news, and obtain a set of keywords for the real-time news;

[0168] The similarity of the keyword set of the real-time news is compared with the keyword set of the historical hot news cluster to determine whether the keyword set of the historical hot news cluster is a subset of the keyword set of the real-time news.

[0169] Determine whether the text similarity between the real-time news and the trending news to be tracked is within a preset similarity range;

[0170] If the real-time news meets the preset subsequent determination conditions, then the real-time news is determined to be a subsequent news of the hot news to be tracked, and the real-time news is added to the subsequent news set. The subsequent determination conditions include that the keyword set of the historical hot news cluster is a subset of the keyword set of the real-time news, and the text similarity between the real-time news and the hot news to be tracked is within the similarity range.

[0171] Based on text similarity, the news in the subsequent news set is clustered to obtain at least one subsequent news cluster.

[0172] Optionally, the second indicator calculation unit is used to calculate the popularity index of the subsequent news cluster, including: the second indicator calculation unit is specifically used for:

[0173] The popularity characteristics of each news item in the subsequent news cluster are acquired in real time, and the popularity index of the subsequent news cluster is calculated based on the popularity characteristics of each news item in the subsequent news cluster.

[0174] Optionally, the second indicator calculation unit is used to calculate the popularity index of the target news cluster based on the popularity characteristics of each news item in the target news cluster, wherein the target news cluster includes the subsequent news cluster and the historical hot news cluster, and includes: the second indicator calculation unit is specifically used for:

[0175] The total number of reads for each news item in the target news cluster is summed to obtain the total number of reads in real time. The dissemination index of the target news cluster is obtained based on the growth rate of the total number of reads in real time. The dissemination index is positively correlated with the growth rate of the total number of reads in real time.

[0176] Based on the number of authoritative media, top media, mid-tier media, and ordinary media for each news item in the target news cluster, the total number of authoritative media, top media, mid-tier media, and ordinary media for the target news cluster is obtained. Based on the growth rate of the total number of authoritative media, the growth rate of the total number of top media, the growth rate of the total number of mid-tier media, and the growth rate of the total number of ordinary media, the publisher authority index of the target news cluster is calculated.

[0177] Based on the number of comments for each news item in the target news cluster, the total number of comments for the target news cluster over at least one time period is obtained. The average content richness of the target news cluster is obtained by averaging the content richness of each news item in the target news cluster. Based on the growth rate of the total number of comments and the average content richness over the at least one time period, the richness index of the target news cluster is calculated.

[0178] Optionally, the preset conditions also include:

[0179] The popularity index of the subsequent news cluster is greater than the popularity index of the hot news to be tracked.

[0180] And / or, the popularity index of the subsequent news cluster is greater than the preset popularity index threshold.

[0181] Optionally, the device further includes: an association unit, used for:

[0182] Obtain the keyword set of the subsequent hot news cluster as the subsequent hot word set;

[0183] The trending news to be tracked, the subsequent cluster of trending news, and the subsequent set of trending keywords are associated.

[0184] Figure 5 The diagram shows the structure of a tracking device for this follow-up news, which may include: at least one processor 501, at least one communication interface 502, at least one memory 503, and at least one communication bus 504;

[0185] In this embodiment of the application, the number of processor 501, communication interface 502, memory 503 and communication bus 504 is at least one, and processor 501, communication interface 502 and memory 503 communicate with each other through communication bus 504.

[0186] The processor 501 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0187] The memory 503 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;

[0188] The memory stores a program, and the processor can execute the program stored in the memory to implement the various steps of the follow-up news tracking method provided in this application embodiment, as follows:

[0189] A method for tracking follow-up news includes:

[0190] Obtain the popularity metrics of trending news to be tracked, including dissemination metrics, publisher authority metrics, and richness metrics;

[0191] Based on the popularity index of the hot news to be tracked, a multivariate logistic regression model is constructed. The multivariate logistic regression model is used to output the popularity probability based on the real-time input popularity index. The popularity probability is used to indicate the probability that the news corresponding to the input popularity index is in an explosive state.

[0192] Obtain subsequent news clusters, which are news sets obtained by clustering and include subsequent news of multiple hot news items to be tracked;

[0193] Calculate the popularity index of the subsequent news cluster;

[0194] The popularity index of the subsequent news cluster is input into the multivariate logistic regression model to obtain the popularity probability output by the multivariate logistic regression model.

[0195] Based on the popularity probability, it is determined whether the subsequent news cluster meets the preset conditions, the preset conditions including the popularity probability being greater than a preset probability threshold;

[0196] If so, the subsequent news clusters will be regarded as the subsequent hot news clusters of the hot news to be tracked.

[0197] Optional metrics for obtaining trending news topics to be tracked include:

[0198] Extract keywords from the trending news stories to be tracked;

[0199] Based on the keywords of the trending news to be tracked, obtain similar news articles.

[0200] Obtain historical hot news clusters, which include the hot news to be tracked and multiple similar news items to the hot news to be tracked;

[0201] The popularity characteristics of each news item in the historical hot news cluster are obtained, including the number of reads, the number of comments, the number of authoritative media, the number of top media, the number of mid-tier media, the number of ordinary media, and the richness of content;

[0202] Based on the popularity characteristics of each news item in the historical hot news cluster, the popularity index of the historical hot news cluster is calculated.

[0203] The popularity index of the historical hot news clusters is used as the popularity index for obtaining the hot news to be tracked.

[0204] Optionally, retrieve subsequent news clusters, including:

[0205] Monitor real-time news, obtain keywords from the real-time news, and obtain a set of keywords for the real-time news;

[0206] The similarity of the keyword set of the real-time news is compared with the keyword set of the historical hot news cluster to determine whether the keyword set of the historical hot news cluster is a subset of the keyword set of the real-time news.

[0207] Determine whether the text similarity between the real-time news and the trending news to be tracked is within a preset similarity range;

[0208] If the real-time news meets the preset subsequent determination conditions, then the real-time news is determined to be a subsequent news of the hot news to be tracked, and the real-time news is added to the subsequent news set. The subsequent determination conditions include that the keyword set of the historical hot news cluster is a subset of the keyword set of the real-time news, and the text similarity between the real-time news and the hot news to be tracked is within the similarity range.

[0209] Based on text similarity, the news in the subsequent news set is clustered to obtain at least one subsequent news cluster.

[0210] Optionally, the popularity index of the subsequent news cluster is calculated, including:

[0211] The popularity characteristics of each news item in the subsequent news cluster are acquired in real time, and the popularity index of the subsequent news cluster is calculated based on the popularity characteristics of each news item in the subsequent news cluster.

[0212] Optionally, based on the popularity characteristics of each news item in the target news cluster, a popularity index for the target news cluster is calculated. The target news cluster includes the subsequent news cluster and the historical hot news cluster, including:

[0213] The total number of reads for each news item in the target news cluster is summed to obtain the total number of reads in real time. The dissemination index of the target news cluster is obtained based on the growth rate of the total number of reads in real time. The dissemination index is positively correlated with the growth rate of the total number of reads in real time.

[0214] Based on the number of authoritative media, top media, mid-tier media, and ordinary media for each news item in the target news cluster, the total number of authoritative media, top media, mid-tier media, and ordinary media for the target news cluster is obtained. Based on the growth rate of the total number of authoritative media, the growth rate of the total number of top media, the growth rate of the total number of mid-tier media, and the growth rate of the total number of ordinary media, the publisher authority index of the target news cluster is calculated.

[0215] Based on the number of comments for each news item in the target news cluster, the total number of comments for the target news cluster over at least one time period is obtained. The average content richness of the target news cluster is obtained by averaging the content richness of each news item in the target news cluster. Based on the growth rate of the total number of comments and the average content richness over the at least one time period, the richness index of the target news cluster is calculated.

[0216] Optionally, the preset conditions also include:

[0217] The popularity index of the subsequent news cluster is greater than the popularity index of the hot news to be tracked.

[0218] And / or, the popularity index of the subsequent news cluster is greater than the preset popularity index threshold.

[0219] Optionally, this method also includes:

[0220] Obtain the keyword set of the subsequent hot news cluster as the subsequent hot word set;

[0221] The trending news to be tracked, the subsequent cluster of trending news, and the subsequent set of trending keywords are associated.

[0222] This application embodiment also provides a readable storage medium that can store a computer program suitable for execution by a processor. When the computer program is executed by the processor, it implements the various steps of a follow-up news tracking method provided in this application embodiment, as follows:

[0223] A method for tracking follow-up news includes:

[0224] Obtain the popularity metrics of trending news to be tracked, including dissemination metrics, publisher authority metrics, and richness metrics;

[0225] Based on the popularity index of the hot news to be tracked, a multivariate logistic regression model is constructed. The multivariate logistic regression model is used to output the popularity probability based on the real-time input popularity index. The popularity probability is used to indicate the probability that the news corresponding to the input popularity index is in an explosive state.

[0226] Obtain subsequent news clusters, which are news sets obtained by clustering and include subsequent news of multiple hot news items to be tracked;

[0227] Calculate the popularity index of the subsequent news cluster;

[0228] The popularity index of the subsequent news cluster is input into the multivariate logistic regression model to obtain the popularity probability output by the multivariate logistic regression model.

[0229] Based on the popularity probability, it is determined whether the subsequent news cluster meets the preset conditions, the preset conditions including the popularity probability being greater than a preset probability threshold;

[0230] If so, the subsequent news clusters will be regarded as the subsequent hot news clusters of the hot news to be tracked.

[0231] Optional metrics for obtaining trending news topics to be tracked include:

[0232] Extract keywords from the trending news stories to be tracked;

[0233] Based on the keywords of the trending news to be tracked, obtain similar news articles.

[0234] Obtain historical hot news clusters, which include the hot news to be tracked and multiple similar news items to the hot news to be tracked;

[0235] The popularity characteristics of each news item in the historical hot news cluster are obtained, including the number of reads, the number of comments, the number of authoritative media, the number of top media, the number of mid-tier media, the number of ordinary media, and the richness of content;

[0236] Based on the popularity characteristics of each news item in the historical hot news cluster, the popularity index of the historical hot news cluster is calculated.

[0237] The popularity index of the historical hot news clusters is used as the popularity index for obtaining the hot news to be tracked.

[0238] Optionally, retrieve subsequent news clusters, including:

[0239] Monitor real-time news, obtain keywords from the real-time news, and obtain a set of keywords for the real-time news;

[0240] The similarity of the keyword set of the real-time news is compared with the keyword set of the historical hot news cluster to determine whether the keyword set of the historical hot news cluster is a subset of the keyword set of the real-time news.

[0241] Determine whether the text similarity between the real-time news and the trending news to be tracked is within a preset similarity range;

[0242] If the real-time news meets the preset subsequent determination conditions, then the real-time news is determined to be a subsequent news of the hot news to be tracked, and the real-time news is added to the subsequent news set. The subsequent determination conditions include that the keyword set of the historical hot news cluster is a subset of the keyword set of the real-time news, and the text similarity between the real-time news and the hot news to be tracked is within the similarity range.

[0243] Based on text similarity, the news in the subsequent news set is clustered to obtain at least one subsequent news cluster.

[0244] Optionally, the popularity index of the subsequent news cluster is calculated, including:

[0245] The popularity characteristics of each news item in the subsequent news cluster are acquired in real time, and the popularity index of the subsequent news cluster is calculated based on the popularity characteristics of each news item in the subsequent news cluster.

[0246] Optionally, based on the popularity characteristics of each news item in the target news cluster, a popularity index for the target news cluster is calculated. The target news cluster includes the subsequent news cluster and the historical hot news cluster, including:

[0247] The total number of reads for each news item in the target news cluster is summed to obtain the total number of reads in real time. The dissemination index of the target news cluster is obtained based on the growth rate of the total number of reads in real time. The dissemination index is positively correlated with the growth rate of the total number of reads in real time.

[0248] Based on the number of authoritative media, top media, mid-tier media, and ordinary media for each news item in the target news cluster, the total number of authoritative media, top media, mid-tier media, and ordinary media for the target news cluster is obtained. Based on the growth rate of the total number of authoritative media, the growth rate of the total number of top media, the growth rate of the total number of mid-tier media, and the growth rate of the total number of ordinary media, the publisher authority index of the target news cluster is calculated.

[0249] Based on the number of comments for each news item in the target news cluster, the total number of comments for the target news cluster over at least one time period is obtained. The average content richness of the target news cluster is obtained by averaging the content richness of each news item in the target news cluster. Based on the growth rate of the total number of comments and the average content richness over the at least one time period, the richness index of the target news cluster is calculated.

[0250] Optionally, the preset conditions also include:

[0251] The popularity index of the subsequent news cluster is greater than the popularity index of the hot news to be tracked.

[0252] And / or, the popularity index of the subsequent news cluster is greater than the preset popularity index threshold.

[0253] Optionally, this method also includes:

[0254] Obtain the keyword set of the subsequent hot news cluster as the subsequent hot word set;

[0255] The trending news to be tracked, the subsequent cluster of trending news, and the subsequent set of trending keywords are associated.

[0256] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0257] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0258] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for tracking follow-up news, characterized in that, include: Obtain the popularity metrics of trending news to be tracked, including dissemination metrics, publisher authority metrics, and richness metrics; Based on the popularity index of the hot news to be tracked, a multivariate logistic regression model is constructed. The multivariate logistic regression model is used to output the popularity probability based on the real-time input popularity index. The popularity probability is used to indicate the probability that the news corresponding to the input popularity index is in an explosive state. The process involves obtaining subsequent news clusters, where each cluster is a news set including multiple subsequent news items of the hot news to be tracked, obtained through clustering. This includes: monitoring real-time news, obtaining keywords from the real-time news to form a keyword set; comparing the keyword set of the real-time news with the keyword set of a historical hot news cluster to determine if the keyword set of the historical hot news cluster is a subset of the keyword set of the real-time news, where the historical hot news cluster includes the hot news to be tracked and at least one similar news item; determining if the text similarity between the real-time news and the hot news to be tracked is within a preset similarity range; if the real-time news meets preset subsequent determination conditions, then it is determined that the real-time news is a subsequent news item of the hot news to be tracked, and the real-time news is added to the subsequent news set, where the subsequent determination conditions include the keyword set of the historical hot news cluster being a subset of the keyword set of the real-time news, and the text similarity between the real-time news and the hot news to be tracked being within the similarity range; and clustering the news items in the subsequent news set based on text similarity to obtain at least one subsequent news cluster. Calculate the popularity index of the subsequent news cluster; The popularity index of the subsequent news cluster is input into the multivariate logistic regression model to obtain the popularity probability output by the multivariate logistic regression model. Based on the popularity probability, it is determined whether the subsequent news cluster meets the preset conditions, the preset conditions including the popularity probability being greater than a preset probability threshold; If so, the subsequent news clusters will be regarded as the subsequent hot news clusters of the hot news to be tracked.

2. The method according to claim 1, characterized in that, The heat index for acquiring trending news items to be tracked includes: Extract keywords from the trending news stories to be tracked; Based on the keywords of the trending news to be tracked, obtain similar news articles. Obtain historical hot news clusters, which include the hot news to be tracked and multiple similar news items to the hot news to be tracked; The popularity characteristics of each news item in the historical hot news cluster are obtained, including the number of reads, the number of comments, the number of authoritative media, the number of top media, the number of mid-tier media, the number of ordinary media, and the richness of content; Based on the popularity characteristics of each news item in the historical hot news cluster, the popularity index of the historical hot news cluster is calculated. The popularity index of the historical hot news clusters is used as the popularity index for obtaining the hot news to be tracked.

3. The method according to claim 1, characterized in that, The calculation of the popularity index of the subsequent news cluster includes: The popularity characteristics of each news item in the subsequent news cluster are acquired in real time, and the popularity index of the subsequent news cluster is calculated based on the popularity characteristics of each news item in the subsequent news cluster.

4. The method according to claim 3, characterized in that, Based on the popularity characteristics of each news item in the target news cluster, a popularity index for the target news cluster is calculated. The target news cluster includes the subsequent news cluster and the historical hot news cluster, including: The total number of reads for each news item in the target news cluster is summed to obtain the total number of reads in real time. The dissemination index of the target news cluster is obtained based on the growth rate of the total number of reads in real time. The dissemination index is positively correlated with the growth rate of the total number of reads in real time. Based on the number of authoritative media, top media, mid-tier media, and ordinary media for each news item in the target news cluster, the total number of authoritative media, top media, mid-tier media, and ordinary media for the target news cluster is obtained. Based on the growth rate of the total number of authoritative media, the growth rate of the total number of top media, the growth rate of the total number of mid-tier media, and the growth rate of the total number of ordinary media, the publisher authority index of the target news cluster is calculated. Based on the number of comments for each news item in the target news cluster, the total number of comments for the target news cluster over at least one time period is obtained. The average content richness of the target news cluster is obtained by averaging the content richness of each news item in the target news cluster. Based on the growth rate of the total number of comments and the average content richness over the at least one time period, the richness index of the target news cluster is calculated.

5. The method according to claim 1, characterized in that, The preset conditions also include: The popularity index of the subsequent news cluster is greater than the popularity index of the hot news to be tracked. And / or, the popularity index of the subsequent news cluster is greater than the preset popularity index threshold.

6. The method according to claim 1, characterized in that, The method further includes: Obtain the keyword set of the subsequent hot news cluster as the subsequent hot word set; The trending news to be tracked, the subsequent cluster of trending news, and the subsequent set of trending keywords are associated.

7. A follow-up news tracking device, characterized in that, include: The first indicator calculation unit is used to obtain the popularity indicators of the hot news to be tracked. The popularity indicators include the dissemination index, the authoritative index of the publisher, and the richness index. The model building unit is used to build a multivariate logistic regression model based on the popularity index of the hot news to be tracked. The multivariate logistic regression model is used to output the popularity probability based on the real-time input popularity index. The popularity probability is used to indicate the probability that the news corresponding to the input popularity index is in an explosive state. The news monitoring unit is used to acquire subsequent news clusters, which are news sets obtained by clustering and include subsequent news of multiple hot news items to be tracked. The second indicator calculation unit is used to calculate the popularity index of the subsequent news cluster. The model application unit is used to input the popularity index of the subsequent news cluster into the multivariate logistic regression model and obtain the popularity probability output by the multivariate logistic regression model. A condition judgment unit is used to determine whether the subsequent news cluster meets a preset condition based on the popularity probability, wherein the preset condition includes the popularity probability being greater than a preset probability threshold. The subsequent hotspot determination unit is used to, if so, identify the subsequent news cluster as the subsequent hotspot news cluster of the hotspot news to be tracked; The news monitoring unit is specifically used for: monitoring real-time news, obtaining keywords of the real-time news, and obtaining a set of keywords of the real-time news; comparing the similarity between the set of keywords of the real-time news and the set of keywords of historical hot news clusters, and determining whether the set of keywords of historical hot news clusters is a subset of the set of keywords of the real-time news, wherein the historical hot news clusters include hot news to be tracked and at least one similar news to the hot news to be tracked. Determine whether the text similarity between the real-time news and the trending news to be tracked is within a preset similarity range; if the real-time news meets preset subsequent determination conditions, then determine that the real-time news is a subsequent news of the trending news to be tracked, and add the real-time news to the subsequent news set. The subsequent determination conditions include that the keyword set of the historical trending news cluster is a subset of the keyword set of the real-time news, and the text similarity between the real-time news and the trending news to be tracked is within the similarity range; based on the text similarity, cluster the news in the subsequent news set to obtain at least one subsequent news cluster.

8. A follow-up news tracking device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement the steps of the follow-up news tracking method as described in any one of claims 1 to 6.

9. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the follow-up news tracking method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Hot news acquisition method, equipment and storage medium

    CN108897774A