A method and apparatus for analyzing the correlation heat of trending events

By analyzing the hot topic value set and relevance of online articles, a correlation heat model is constructed, which solves the problem of lagging trend prediction of hot events in existing technologies, realizes early and accurate prediction and public opinion control, and supports news production and social stability.

CN116860982BActive Publication Date: 2025-11-14BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310076435.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-17
Publication Date
2025-11-14
Estimated Expiration
2043-01-17

AI Technical Summary

Technical Problem

Existing methods for analyzing the popularity of trending events cannot effectively predict their development trends when a topic first emerges, resulting in an inability to timely and accurately control the direction of public opinion, and they also lack analysis of the connections between trending events.

Method used

By acquiring text data and basic information from online articles, a hot topic value set and popularity value are generated. Combining the relevance of time, location, people, and behavior, the correlation popularity is calculated, a topic relationship graph is constructed, and popularity prediction is performed.

Benefits of technology

It enables early and accurate prediction and trend analysis of trending events, helps users efficiently obtain valuable news information, supports the monitoring of trending events and news production, suppresses negative emotions, and maintains social stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116860982B_ABST
    Figure CN116860982B_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for analyzing the correlation heat of trending events. It generates a set of trending values ​​corresponding to online articles based on text data, and generates a heat value corresponding to the online articles based on the set of trending values ​​and a trending event thesaurus. It also generates a relevance degree corresponding to the online articles based on basic information, and generates a topic importance degree corresponding to the online articles based on the relevance degree. Finally, it generates a correlation heat corresponding to the online articles based on the heat value and the topic importance. This application provides an in-depth analysis of the heat calculation, correlation relevance analysis, and correlation heat calculation used in the whistleblowing system, listing relevant factors involved in the relevant formulas and models, such as different types of word features like time, location, people, and behavior, thereby enabling the calculation of the correlation between events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, and specifically to a method and apparatus for analyzing the correlation and popularity of trending events. Background Technology

[0002] Hot topics are an unavoidable phenomenon in today's internet age. Their occurrence often attracts intense public attention, with people continuously expressing their opinions, attitudes, and emotions. These online hot topics, from their inception to a certain period, often form a focal point, representing the interests and core emotions of netizens. In this era of data explosion, how to leverage massive amounts of historical data based on news information to provide editors, reporters, and other media professionals with fast, accurate, and personalized news lead recommendations and intelligent early warning support, enhance their awareness of hot topic situations and news insight, and effectively improve work efficiency and news creation capabilities, is a problem that needs to be solved. In recent years, research on the analysis of online hot topics has gradually deepened into experimental research projects by ordinary scholars. These projects generally focus on social networks or applications such as Weibo, WeChat, and forums, where a large number of active users exist. Once a hot topic emerges, its spread accelerates exponentially. Hot online hot topics primarily rely on the internet for dissemination; a hot topic attracts public attention, commentary, and dissemination, thereby generating wider social awareness. In terms of trend analysis, domestic researchers use the influence propagation model to describe trending events. This model counts the number of times keywords are spread; a higher number indicates higher influence, and vice versa. The influence propagation model can be used to assess the degree of interaction between different users on social networks. Furthermore, by analyzing related messages and the number of reposts, it determines whether a topic qualifies as trending. By utilizing user attention to construct the influence propagation model, the number of times keywords are spread reflects the magnitude of an event's influence.

[0003] A review of existing literature on trend analysis reveals two main perspectives for analyzing online trending topics. The first is from the user's perspective, analyzing user activity on platforms like forums and microblogs. Topics arise from user descriptions of events, and the key difference between trending and ordinary topics lies in the amount of information users use to describe them, the amount of online resources consumed, and the duration of discussion. The second perspective is from the media's perspective, analyzing the reposting and ranking of trending events on news websites like Sina and Sohu. The emergence and spread of a topic occurs after widespread public discussion and media reporting and reposting. Whether a topic becomes a trending topic is often measured by the quantity and frequency of reports. However, existing trend analysis methods fail to establish connections between trending events. Furthermore, most existing research uses trend calculations combined with historical data for verification, which often has a lag, failing to effectively predict a topic's development trend as soon as it emerges. This hinders timely and accurate regulation of public opinion by relevant departments and prevents continuous monitoring of topics according to established rules. Summary of the Invention

[0004] In view of the aforementioned problems, this application is proposed to provide a method and apparatus for analyzing the correlation heat of hot events to overcome or at least partially solve the aforementioned problems, comprising:

[0005] A method for analyzing the relevance of trending events, the method being used to analyze the relevance of online articles by using a terminology database of trending events corresponding to the current trending event, including:

[0006] Obtain the text data of online articles corresponding to the current trending event and the basic information corresponding to the online articles; wherein, the online articles include at least one article;

[0007] A hotspot value set corresponding to the online article is generated based on the text data, and a popularity value corresponding to the online article is generated based on the hotspot value set and the hotspot event lexicon; wherein, the hotspot value set includes the hotspot value corresponding to each word;

[0008] Based on the basic information, a relevance score is generated corresponding to the online article, and based on the relevance score, a topic importance score is generated corresponding to the online article.

[0009] Based on the popularity value and the importance of the topic, a related popularity value corresponding to the online article is generated.

[0010] Preferably, the step of generating a hotspot value set corresponding to the online article based on the text data includes:

[0011] Based on the text data, a binary distribution statistical result set corresponding to the online article is generated by performing binary distribution statistics; wherein, the binary distribution statistical result set includes the binary distribution statistical result corresponding to each word;

[0012] Based on the binary distribution statistical result set, determine the hotspot value set corresponding to the online article.

[0013] Preferably, the step of generating a popularity value corresponding to the online article based on the hotspot value set and a preset hotspot event thesaurus includes:

[0014] Based on the hotspot value set, generate a hotspot active term library and a hotspot inert term library corresponding to the online article;

[0015] A first co-occurrence threshold is generated based on the aforementioned hot topic active term library and the preset hot topic event term library;

[0016] A second co-occurrence threshold is generated based on the aforementioned hotspot inert lexicon and the preset hotspot event lexicon;

[0017] The popularity value corresponding to the online article is generated based on the first co-occurrence threshold and the second co-occurrence threshold.

[0018] Preferably, the relevance includes time relevance, location relevance, person relevance, and behavior relevance, and the step of generating a relevance score corresponding to the online article based on the basic information includes:

[0019] Based on the aforementioned basic information, the corresponding time information, corresponding location information, corresponding person information, and corresponding behavior information are determined;

[0020] The corresponding time correlation is generated based on the corresponding time information;

[0021] The corresponding location relevance is generated based on the corresponding location information;

[0022] The corresponding person relevance is generated based on the corresponding person information;

[0023] The corresponding behavioral relevance is generated based on the corresponding behavioral information.

[0024] Preferably, the step of generating a topic importance corresponding to the online article based on the relevance includes:

[0025] Generate a relationship matrix corresponding to the online article based on the relevance;

[0026] The topic importance corresponding to the online article is generated based on the relationship matrix.

[0027] Preferably, the step of generating a relationship matrix corresponding to the online article based on the relevance includes:

[0028] Construct a topic relationship graph corresponding to the online articles based on the relevance;

[0029] Based on the topic relationship diagram, determine the relationship matrix corresponding to the online article.

[0030] Preferably, the step of generating a hot topic active term library and a hot topic inert term library corresponding to the online article based on the hot topic value set includes:

[0031] The hot topic active word library corresponding to the online article is generated from words whose values ​​in the hot topic value set are greater than or equal to the upper threshold.

[0032] The hot topic inert word library corresponding to the online article is generated from words whose values ​​in the hot topic value set are less than or equal to the lower threshold.

[0033] A device for analyzing the relevance of trending events, the device being used to analyze the relevance of online articles using a terminology database corresponding to the current trending event, comprising:

[0034] The data acquisition module is used to acquire text data of online articles corresponding to the current hot topic and basic information corresponding to the online articles; wherein, the online articles include at least one article;

[0035] A popularity value generation module is used to generate a hotspot value set corresponding to the online article based on the text data, and to generate a popularity value corresponding to the online article based on the hotspot value set and the hotspot event lexicon; wherein, the hotspot value set includes a hotspot value corresponding to each word;

[0036] The topic importance generation module is used to generate a relevance score corresponding to the online article based on the basic information, and to generate a topic importance score corresponding to the online article based on the relevance score.

[0037] The related popularity generation module is used to generate related popularity corresponding to the online article based on the popularity value and the importance of the topic.

[0038] To achieve this, the application also includes an electronic device, comprising a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the steps of the hot topic association heat analysis method.

[0039] To realize the present application, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the steps of the hot topic event correlation heat analysis method.

[0040] This application has the following advantages:

[0041] In embodiments of this application, text data of online articles corresponding to current trending events and basic information corresponding to the online articles are obtained; wherein, the online articles include at least one article; a trending value set corresponding to the online articles is generated based on the text data, and a popularity value corresponding to the online articles is generated based on the trending value set and the trending event thesaurus; wherein, the trending value set includes a trending value corresponding to each word; a relevance degree corresponding to the online articles is generated based on the basic information, and a topic importance degree corresponding to the online articles is generated based on the relevance degree;

[0042] Based on the popularity value and the importance of the topic, a correlation popularity corresponding to the online article is generated. This application provides an in-depth analysis of the popularity calculation, correlation analysis, and correlation popularity calculation used in the whistleblowing system, listing relevant factors involved in the formulas and models, such as different types of word features like time, location, people, and behavior. This allows for the calculation of the correlation between events, enabling better prediction of whether they will develop into trending events when combined with other data. Through the above methods and practical applications, this application demonstrates good accuracy, allowing users to efficiently and intelligently obtain targeted news information that is of interest and value from massive amounts of news information, thus more effectively supporting business operations such as trending event monitoring, news tracking, and news production. This system can guide the direction of trending events and allow for rapid responses to major public opinion events. This can, to some extent, suppress negative emotions generated by public opinion events, which will help correctly guide the development trend of trending events and maintain social harmony and stability. Attached Figure Description

[0043] To more clearly illustrate the technical solution of this application, the drawings used in the description of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a flowchart illustrating the steps of a method for analyzing the correlation heat of trending events according to an embodiment of this application;

[0045] Figure 2This is a topic relationship diagram of a hot topic trend analysis method provided in one embodiment of this application;

[0046] Figure 3 This is a heat prediction framework diagram of a correlation heat analysis method for hot events provided in an embodiment of this application;

[0047] Figure 4 This is a structural block diagram of a hot topic correlation heat analysis device provided in an embodiment of this application;

[0048] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0049] To make the objectives, features, and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0050] Reference Figure 1 The diagram illustrates a flowchart of a method for analyzing the correlation heat of trending events according to an embodiment of this application, specifically including the following steps:

[0051] S110. Obtain the text data of online articles corresponding to the current hot topic and the basic information corresponding to the online articles; wherein, the online articles include at least one article;

[0052] S120. Generate a hotspot value set corresponding to the online article based on the text data, and generate a popularity value corresponding to the online article based on the hotspot value set and the hotspot event lexicon; wherein, the hotspot value set includes the hotspot value corresponding to each word;

[0053] S130. Generate a relevance score corresponding to the online article based on the basic information, and generate a topic importance score corresponding to the online article based on the relevance score;

[0054] S140. Generate a related popularity score corresponding to the online article based on the popularity score and the importance of the topic.

[0055] This application involves acquiring text data of online articles corresponding to current trending events and basic information about those articles; wherein each online article includes at least one article; generating a trending value set corresponding to the online article based on the text data, and generating a popularity value corresponding to the online article based on the trending value set and a trending event terminology database; wherein the trending value set includes a trending value corresponding to each word; generating a relevance degree corresponding to the online article based on the basic information, and generating a topic importance degree corresponding to the online article based on the relevance degree; and generating an association popularity degree corresponding to the online article based on the popularity value and the topic importance. This application provides an in-depth analysis of the popularity calculation, association relevance analysis, and association popularity calculation used in the whistleblowing system, listing relevant factors involved in the relevant formulas and models, such as different types of word features like time, location, people, and behavior, thereby enabling the calculation of the association degree between events and allowing for better prediction of whether an event will develop into a trending event when combined with other data. Through the aforementioned methods and practical applications, this application has demonstrated good accuracy, enabling users to efficiently and intelligently obtain relevant, interesting, and valuable target news information from massive amounts of news data. This provides stronger support for business operations such as monitoring hot topics, tracking news, and producing news. Furthermore, the system can guide the direction of hot topics and allow for rapid responses to major public opinion events. This can, to some extent, suppress negative public sentiment towards public opinion events, contributing to the correct guidance of the development trend of hot topics and maintaining social harmony and stability.

[0056] The following will further explain the correlation heat analysis method for hot events in this exemplary embodiment.

[0057] As described in step S110 above, obtain the text data of online articles corresponding to the current hot topic and the basic information corresponding to the online articles; wherein, the online articles include at least one article.

[0058] In one embodiment of the present invention, the specific process of step S110, "obtaining text data of online articles corresponding to the current hot topic and basic information corresponding to the online articles; wherein the online articles include at least one article," can be further described in conjunction with the following description.

[0059] The following steps describe the process of acquiring online articles about trending news events on the internet; automatically collecting online articles about trending news events from internet websites; where the websites can be specified and pre-configured, and there are no restrictions here. The sources of the trending news event data include mainstream news websites and new media websites. The acquired trending event data can include one or more of the following: news sources, forums, microblogs, and search engines.

[0060] It should be noted that the online articles in this application can be short articles, long articles, or a single topic, and are not limited to the number of words.

[0061] In one specific embodiment, the online articles on trending events obtained by this application may not initially have any popularity. Positive articles can be pushed, while negative articles are not pushed, thereby enabling better prediction and helping relevant departments to effectively and accurately regulate public opinion regarding trending events in a timely manner.

[0062] As described in step S120 above, a hotspot value set corresponding to the online article is generated based on the text data, and a popularity value corresponding to the online article is generated based on the hotspot value set and the hotspot event lexicon; wherein, the hotspot value set includes the hotspot value corresponding to each word.

[0063] In one embodiment of the present invention, the specific process of step S120, which involves "generating a hotspot value set corresponding to the online article based on the text data, and generating a popularity value corresponding to the online article based on the hotspot value set and the hotspot event lexicon; wherein the hotspot value set includes a hotspot value corresponding to each word", can be further explained in conjunction with the following description.

[0064] As described in the following steps, a binary distribution statistical result set corresponding to the online article is generated based on the text data; wherein, the binary distribution statistical result set includes the binary distribution statistical result corresponding to each word; and a hotspot value set corresponding to the online article is determined based on the binary distribution statistical result set.

[0065] As described in the following steps, a hot topic active term library and a hot topic inert term library corresponding to the online article are generated based on the hot topic value set; a first co-occurrence threshold is generated based on the hot topic active term library and the preset hot topic event term library; a second co-occurrence threshold is generated based on the hot topic inert term library and the preset hot topic event term library; and the popularity value corresponding to the online article is generated based on the first co-occurrence threshold and the second co-occurrence threshold.

[0066] In one embodiment of the present invention, the specific process of the step "generating a binary distribution statistical result set corresponding to the online article based on the text data by performing binary distribution statistics based on the text data; determining the hotspot value set corresponding to the online article based on the binary distribution statistical result set" can be further explained in conjunction with the following description.

[0067] In one specific embodiment, firstly, the text data of online articles is semantically decomposed, and then the semantically decomposed news hot topic data, i.e. word-based data, is subjected to binary distribution statistics to count the number of times each word appears, and the binary distribution statistics results are obtained.

[0068] Then, the binary distribution statistics are used to calculate the hotspot value of each word using the Z-Score algorithm. The formula is as follows:

[0069]

[0070] Among them, in the formula The number of times a term appears; This represents the average number of times each term appears. The standard deviation is 1; the result is 1. It is the deviation from the mean, expressed in standard deviation, used to represent the hotspot value of a word.

[0071] The hotspot value of each word in each article is calculated to obtain a hotspot value set corresponding to the online article. Each hotspot value set is a set of hotspot values ​​of each word in each online article, and the number of hotspot value sets is the same as the number of online articles.

[0072] In one embodiment of the present invention, the steps of "generating a hot topic active term library and a hot topic inert term library corresponding to the online article based on the hot topic value set; generating a first co-occurrence threshold based on the hot topic active term library and the preset hot topic event term library; generating a second co-occurrence threshold based on the hot topic inert term library and the preset hot topic event term library" can be further explained in conjunction with the following description.

[0073] The specific process of generating the popularity value corresponding to the online article based on the first co-occurrence threshold and the second co-occurrence threshold.

[0074] As described in the following steps, words with values ​​greater than or equal to the upper threshold in the hotspot value set are used to generate the hotspot active word library corresponding to the online article; words with values ​​less than or equal to the lower threshold in the hotspot value set are used to generate the hotspot inactive word library corresponding to the online article.

[0075] In one specific embodiment, hotspot values ​​greater than a preset upper threshold are stored in the active hotspot terminology within the hotspot terminology library, while hotspot values ​​less than a preset lower threshold are stored in the inactive hotspot terminology library. The hotspot terminology library is associated with a domain terminology library, which includes fields such as news, blogs, forums, and social networking sites. The domains from which hotspot words in each hotspot terminology library originate can be queried. Then, based on the word hotspot values ​​and the preset hotspot terminology library, the co-occurrence threshold of hotspot words in the word-based data is determined.

[0076] Based on the terms appearing in the news trending events data, the co-occurrence threshold P1 of trending active terms is calculated using the following formula:

[0077]

[0078] in A collection of news terms. This is a set of trending active words. Then, the co-occurrence threshold of trending inactive words is calculated using the following formula. :

[0079]

[0080] in A collection of news terms. This is a set of trending inert words. Then, based on the co-occurrence thresholds of trending active words and trending inert words... and A linear weighted calculation is performed to obtain the heat value. The formula for calculating the heat value is as follows:

[0081]

[0082] in For the first The hotspot value of each word The threshold for co-occurrence of trending active keywords. The threshold for co-occurrence of inert words in trending topics is set. Then, the trending news event data is judged based on the trending value, and the trending value is graded according to the preset trending level judgment criteria. News event data that meets the trending level judgment criteria are archived into trending documents, and news event data that does not meet the trending level judgment criteria are archived into non-trending documents.

[0083] In one specific embodiment, hot documents and non-hot documents are generated based on the popularity value.

[0084] In one specific embodiment, the method further includes: generating a news sensitivity corresponding to the online article based on the hotspot value set and a preset sensitive word library.

[0085] In one specific embodiment, in sensitivity analysis, the number of sensitive words contained in the hot topic active word library is obtained by comparing it with a preset sensitive word library, and then the sensitivity value is calculated as the news sensitivity S using the following formula:

[0086]

[0087] in, To the number of sensitive words contained, This refers to the number of trending and active words in the news within the domain thesaurus.

[0088] As described in step S130 above, a relevance score corresponding to the online article is generated based on the basic information, and a topic importance score corresponding to the online article is generated based on the relevance score.

[0089] In one embodiment of the present invention, the specific process of step S130, "generating a relevance degree corresponding to the online article based on the basic information, and generating a topic importance degree corresponding to the online article based on the relevance degree," can be further explained in conjunction with the following description.

[0090] As described in the following steps, based on the basic information, corresponding time information, corresponding location information, corresponding person information, and corresponding behavior information are determined; based on the corresponding time information, the corresponding time relevance is generated; based on the corresponding location information, the corresponding location relevance is generated; based on the corresponding person information, the corresponding person relevance is generated; and based on the corresponding behavior information, the corresponding behavior relevance is generated.

[0091] As described in the following steps, a relationship matrix corresponding to the online article is generated based on the relevance; and a topic importance corresponding to the online article is generated based on the relationship matrix.

[0092] In one embodiment of the present invention, the specific process of the steps "determining corresponding time information, corresponding location information, corresponding person information and corresponding behavior information based on the basic information; generating the corresponding time relevance based on the corresponding time information; generating the corresponding location relevance based on the corresponding location information; generating the corresponding person relevance based on the corresponding person information; and generating the corresponding behavior relevance based on the corresponding behavior information" can be further explained in conjunction with the following description.

[0093] In one specific embodiment, predicting trending events requires judging the future trend of a topic. Generally, the higher the popularity value of a related topic, the greater the probability that the topic will become a trending topic. In other words, the probability of a predicted topic becoming a trending topic is related to the popularity or number of related topics. The analysis of the relationship between topics mainly includes calculating the degree of correlation between different types of word features such as time, location, people, and behavior, and weighting them.

[0094] In one specific embodiment, time relevance is calculated as follows: The time relevance of topics primarily refers to whether the time difference between two topics occurring falls within a specified range. We need to calculate the time interval and use it to determine the relevance. If the interval is within the range, the two topics are considered to be related in time, and the shorter the time interval, the stronger the correlation. The formula is as follows: Representing the time of a certain topic, and This represents two topics whose relevance needs to be predicted. If it is necessary to analyze the order in which the topics appear, then... Simply arrange them in chronological order.

[0095]

[0096] In one specific embodiment, location relevance is calculated as follows: Location names and other information within the topic are the primary basis for calculating the relevance, and the distance between key locations is used to calculate the relevance value. Therefore, a set of location-related terms needs to be constructed, specifically at the district level in cities or the township level in rural areas, and a hierarchical tree needs to be established corresponding to higher administrative regions. If the predicted geographical locations of a topic are within a certain distance, they can be considered mutually related. The strength of the association can be calculated based on the distance between them; the closer the distance, the higher the degree of association. The formula is as follows, where... Indicates the main location where the topic takes place, and its relationship with The difference represents the path length between the two topics on the hierarchy tree.

[0097]

[0098] In one specific embodiment, the relevance of individuals is calculated as follows: Relevance primarily refers to whether the individuals or organizations involved in the predicted topic follow each other or have other relationships. If there are friends or other relationships, the two topics are considered related in terms of individuals. However, in practical applications, Weibo or WeChat friend relationships are often unavailable. Therefore, the relevance can be calculated using the names in the topic, for example, by calculating the number of repeated names. The formula is as follows, where... A collection of names of people, etc., related to a particular topic. and This represents two topics that need to be predicted.

[0099]

[0100] In one specific embodiment, behavioral relevance is calculated by collecting feature words related to topic behaviors. If the behaviors involved are the same or similar, they are considered related. The formula is as follows, where... and This represents a set of behavioral characteristic words in two topics. This refers to the semantic similarity of the words. It is obtained based on the statistical analysis of the amount of word information in the prediction database.

[0101]

[0102] In one embodiment of the present invention, the specific process of "generating a relationship matrix corresponding to the online article based on the relevance; generating a topic importance corresponding to the online article based on the relationship matrix" can be further explained in conjunction with the following description.

[0103] As described in the following steps, a topic relationship graph corresponding to the online article is constructed based on the relevance; and a relationship matrix corresponding to the online article is determined based on the topic relationship graph.

[0104] In one specific embodiment, while the calculation and prediction of trending event popularity has yielded some results in academia, most algorithms primarily analyze data without considering the inherent characteristics of the trending events themselves, particularly neglecting the interconnectedness of online information. Therefore, this study, based on popularity calculation, incorporates the concept of correlation analysis, comprehensively considering the relevance of time, location, people, and behavior. It mines the correlations between different attributes, constructs a trending event popularity prediction model with correlations, and establishes a corresponding regression model for popularity by analyzing the relationships between related events or information, making the popularity values ​​more closely reflect reality.

[0105] The main purpose of correlation heat calculation is to divide the topic heat into segments according to time, and then identify it according to named entities. For example, time correlation is calculated by time information, location correlation is calculated by location information, person correlation is calculated by person information, behavior correlation is calculated by behavior data, and finally a correlation relationship connection graph is established [9].

[0106] This whistleblowing system establishes a relationship graph between news topics, calculates a popularity value, and sets it as the initial weight value for calculating the related popularity within a certain time period. After the popularity calculation is completed, a relevance algorithm is used to predict and analyze the changing trends of topic popularity, thereby enabling the whistleblowing system to issue early warnings.

[0107] (1) Establishing relationships between topics

[0108] Set A=<V,E> for The relationship diagram is shown below, where... For a given topic, set For the retrieved and A collection of related topics It is a set of edges, where each edge represents the relevance between topics if and only if two vertices are related. When the correlation degree between edges is not less than the threshold, It exists. For example... Figure 2 As shown.

[0109] After establishing the relationship graph, the next step is to convert the graph into matrix form. In the matrix, rows and columns represent points in the relationship graph, and values ​​represent the degree between those points. As shown in the following matrix diagram, where... It represents the correlation between node i and node j. If the correlation is less than the threshold, i.e. there is no edge ij, the value is 0.

[0110]

[0111] In one specific embodiment, the importance of related topics is calculated as follows:

[0112] Define the transformation matrix M as follows:

[0113]

[0114] Where d is the damping coefficient, ranging from 0 to 1. This matrix primarily measures the influence of each point on the point to be predicted. Matrix M has a unique stable distribution. The matrix representation of this model is as follows:

[0115]

[0116] The obtained h value can then be used to represent the importance of a topic in the relationship graph, i.e., topic importance.

[0117] As described in step S140 above, a related popularity score corresponding to the online article is generated based on the popularity score and the importance of the topic.

[0118] In one embodiment of the present invention, the specific process of "generating a related popularity value corresponding to the online article based on the popularity value and the importance of the topic" in step S140 can be further explained in conjunction with the following description.

[0119] In one specific embodiment, to further determine the relevance popularity I, it is necessary to calculate the arithmetic sum of the product of popularity and topic importance, as shown in the formula:

[0120]

[0121] In one specific embodiment, the system also includes trend prediction. In a whistleblowing system, it is necessary to predict the short-term trend of trending events with limited current information to determine whether the topic will become a trending topic. This study uses a gray-scale prediction method for trend prediction. The GM(1,1) model is typically used to predict topic popularity, and the calculation process is as follows:

[0122] a. Input initial sequence ;

[0123] b. Generate the sequence by accumulating the initial sequence once.

[0124] ;

[0125] c. Generate the nearest neighbor mean sequence of X1

[0126]

[0127] d. That is, the grey differential equation model of GM(1,1) is

[0128]

[0129] In the formula, a is the development coefficient, and b is the grey effect quantity. Let... The vector of parameters to be estimated, i.e. Then the least squares estimate parameter sequence of the grey differential equation satisfies

[0130]

[0131] in, ,

[0132] e. Solving the differential equation yields the following solution:

[0133]

[0134] f. Restore to the original data, and obtain

[0135] Once the predicted range for the popularity trend is obtained, the process ends.

[0136] In one specific embodiment, the method further includes: generating the probability of becoming a trending topic based on the correlation popularity, specifically implemented through a model. In practical application, the main method used is to predict the trend of trending events based on event correlation and determine whether they will become trending topics. This model is primarily based on the assumption that "events are interconnected and mutually influential," meaning that there is a certain connection between events, and they may influence or constrain each other. Its algorithm framework is as follows: Figure 3 As shown.

[0137] The specific process can be seen to mainly include:

[0138] (1) Retrieve events related to the topic to be predicted in the recent period. When setting search terms, attention should be paid to the selection of feature words.

[0139] (2) Search the group's local database, compare it with the search results on the Internet, and analyze the relationships between topics to obtain textual information data related to hot events. However, after data collection, noise reduction and other processing are required to ensure a certain level of accuracy.

[0140] (3) Clustering algorithm is used to analyze the sorted text information and extract the number of topics it may contain.

[0141] (4) Sort the text data by time, set time periods according to actual needs, and calculate the relevance between topics based on the time, people, places, and behaviors of the events in each time period, so as to obtain the relationship between all topics, i.e., the relationship connection diagram.

[0142] (5) Analyze the importance of different topics and predict the related popularity, and finally calculate the probability that the topic or information will become a hot topic.

[0143] In one specific embodiment, the experimental results and analysis are then performed. After predicting the popularity of hot events, the whistleblowing system further uses the posterior difference test method to verify the experimental effect. The specific steps include:

[0144] (1) Calculate the average value of the original sequence;

[0145] (2) Calculate the mean square error S1 of the original sequence;

[0146] (3) Calculate the mean of the residuals;

[0147] (4) Calculate the root mean square error of the residuals, S2;

[0148] (5) Calculate the ratio C of S2 to S1.

[0149] (6) Calculate the probability of small residual P

[0150] This study uses P-value and C-value to measure the predictive effect of sudden hot events, and designs a corresponding posterior difference test discrimination reference table.

[0151] Table 1. Reference Table for Posterior Error Test Judgment

[0152]

[0153] The following are the experimental results obtained by predicting the popularity of data related to "Zhang San" in the database:

[0154] Table 2 Experimental Results

[0155]

[0156] The results above show that the correlation heat calculation method is very effective in predicting sudden hot events, verifying the feasibility and effectiveness of the heat analysis technology used in the whistleblowing system.

[0157] In one specific embodiment, this study provides an in-depth analysis of the heat calculation, correlation analysis, correlation heat calculation, and heat prediction used in the whistleblowing system. It lists the relevant factors involved in the formulas and models, such as different types of word features like time, location, people, and behavior, thereby calculating the correlation between events and predicting whether they will develop into hot topics. Through the above methods and practical applications, the newspaper group's whistleblowing system has been proven to have good accuracy, enabling users to efficiently and intelligently obtain targeted news information that is of interest and value from massive amounts of news information, thus more effectively supporting business operations such as hot topic monitoring, news tracking, and news production. The government can also use this system to guide the direction of hot topics and react quickly to major public opinion events. This can, to some extent, suppress negative emotions among the public regarding public opinion events, which will help the government correctly guide the development trend of hot topics and maintain social harmony and stability.

[0158] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0159] Reference Figure 4 This application illustrates a hot topic event correlation heat analysis device according to an embodiment of the present application, which specifically includes the following modules:

[0160] Data acquisition module 410: used to acquire text data of online articles corresponding to the current hot topic and basic information corresponding to the online articles; wherein, the online articles include at least one article;

[0161] Popularity value generation module 420: used to generate a hotspot value set corresponding to the online article based on the text data, and to generate a popularity value corresponding to the online article based on the hotspot value set and the hotspot event lexicon; wherein, the hotspot value set includes a hotspot value corresponding to each word;

[0162] Topic Importance Generation Module 430: Used to generate a relevance score corresponding to the online article based on the basic information, and to generate a topic importance score corresponding to the online article based on the relevance score;

[0163] Related popularity generation module 440: used to generate related popularity corresponding to the online article based on the popularity value and the importance of the topic.

[0164] In one embodiment of the present invention, the heat value generation module 420 includes:

[0165] Binary distribution statistical result set device: used to generate a binary distribution statistical result set corresponding to the online article based on the text data; wherein, the binary distribution statistical result set includes the binary distribution statistical result corresponding to each word;

[0166] Hotspot value set device: used to determine the hotspot value set corresponding to the online article based on the binary distribution statistical result set.

[0167] Thesaurus device: used to generate a hot topic active thesaurus and a hot topic inert thesaurus corresponding to the online articles based on the hot topic value set;

[0168] First co-occurrence threshold device: used to generate a first co-occurrence threshold based on the hot topic active word library and the preset hot topic event word library;

[0169] Second co-occurrence threshold device: used to generate a second co-occurrence threshold based on the hotspot lazy lexicon and the preset hotspot event lexicon;

[0170] Popularity value device: used to generate the popularity value corresponding to the online article based on the first co-occurrence threshold and the second co-occurrence threshold.

[0171] In one embodiment of the present invention, the lexicon device includes:

[0172] Hot Topic Active Keywords Submodule: Used to generate the hot topic active keywords corresponding to the online article by selecting words whose hot topic value is greater than or equal to the upper threshold value from the hot topic value set;

[0173] Hotspot Inertial Terminology Submodule: Used to generate the hotspot inertial terminology corresponding to the online article from words whose values ​​in the hotspot value set are less than or equal to the lower threshold.

[0174] In one embodiment of the present invention, the topic importance generation module 430 includes:

[0175] Basic information device: used to determine corresponding time information, corresponding location information, corresponding person information, and corresponding behavior information based on the basic information;

[0176] Time correlation device: used to generate the corresponding time correlation based on the corresponding time information;

[0177] Location relevance device: used to generate the corresponding location relevance based on the corresponding location information;

[0178] Person Relevance Device: Used to generate the corresponding person relevance based on the corresponding person information;

[0179] Behavioral relevance device: used to generate the corresponding behavioral relevance based on the corresponding behavioral information.

[0180] Relationship matrix device: used to generate a relationship matrix corresponding to the online article based on the relevance;

[0181] Topic Importance Device: Used to generate topic importance corresponding to the online article based on the relationship matrix.

[0182] In one embodiment of the present invention, the relation matrix device includes:

[0183] Topic Relationship Graph Submodule: Used to construct a topic relationship graph corresponding to the online article based on the relevance;

[0184] The relation matrix submodule is used to determine the relation matrix corresponding to the online article based on the topic relation graph.

[0185] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0186] This specific embodiment has some overlapping operation steps with the above specific embodiments. This specific embodiment is only briefly described. For other solutions, please refer to the description of the above specific embodiments.

[0187] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0188] Reference Figure 5 The computer device shown in this application, which provides a method for analyzing the correlation heat of trending events, may specifically include the following:

[0189] The aforementioned computer device 12 is manifested in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, memory 28, and a bus 18 connecting different system components (including memory 28 and processing unit 16).

[0190] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Audio / Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0191] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0192] Memory 28 may include computer system readable media in the form of volatile memory, such as random access memory 30 and / or cache memory 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (commonly referred to as a "hard disk drive"). Figure 5 As not shown, a disk drive for reading and writing to a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. The memory may include at least one program product having a set (e.g., at least one) of program modules 42 configured to perform the functions of the embodiments of this application.

[0193] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory. Such program modules 42 include—but are not limited to—an operating system, one or more application programs, other program modules 42, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of this application.

[0194] Computer device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, camera, etc.), and with one or more devices that enable an operator to interact with the computer device 12, and / or with any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed through I / O interface 22. Furthermore, computer device 12 can also communicate with one or more networks (e.g., local area network (LAN)), wide area network (WAN), and / or public networks (e.g., the Internet) via network adapter 20. Figure 5 As shown, network adapter 20 communicates with other modules of computer device 12 via bus 18. It should be understood that, although... Figure 5 Not shown, it can be combined with computer device 12 to use other hardware and / or software modules, including but not limited to: microcode, device drivers, redundant processing unit 16, external disk drive array, RAID system, tape drive and data backup storage system 34, etc.

[0195] The processing unit 16 executes various functional applications and data processing by running programs stored in memory 28, such as implementing a hot topic event correlation heat analysis method provided in the embodiments of this application.

[0196] That is, when the processing unit 16 executes the above procedure, it performs the following: obtaining text data of online articles corresponding to the current hot topic event and basic information corresponding to the online articles; wherein, the online articles include at least one article; generating a hot topic value set corresponding to the online articles based on the text data, and generating a popularity value corresponding to the online articles based on the hot topic value set and the hot topic event thesaurus; wherein, the hot topic value set includes a hot topic value corresponding to each word; generating a relevance degree corresponding to the online articles based on the basic information, and generating a topic importance degree corresponding to the online articles based on the relevance degree; generating an association popularity degree corresponding to the online articles based on the popularity value and the topic importance.

[0197] In this application embodiment, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a hot topic event correlation heat analysis method as provided in all embodiments of this application.

[0198] That is, when the program is executed by the processor, it performs the following: obtaining text data of online articles corresponding to the current hot topic event and basic information corresponding to the online articles; wherein, the online articles include at least one article; generating a hot topic value set corresponding to the online articles based on the text data, and generating a popularity value corresponding to the online articles based on the hot topic value set and the hot topic event thesaurus; wherein, the hot topic value set includes a hot topic value corresponding to each word; generating a relevance degree corresponding to the online articles based on the basic information, and generating a topic importance degree corresponding to the online articles based on the relevance degree; generating an association popularity degree corresponding to the online articles based on the popularity value and the topic importance.

[0199] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0200] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including—but not limited to—electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of transmitting, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0201] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the operator's computer, partially on the operator's computer, as a standalone software package, partially on the operator's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the operator's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider). The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably.

[0202] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0203] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0204] The above provides a detailed description of the method and apparatus for analyzing the correlation of trending events provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and its core ideas. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for analyzing the relevance of trending events, wherein the method is used to analyze the relevance of online articles by using a terminology database of trending events corresponding to the current trending event, characterized in that, include: Obtain the text data of online articles corresponding to the current trending event and the basic information corresponding to the online articles; wherein, the online articles include at least one article; A hotspot value set corresponding to the online article is generated based on the text data, and a popularity value corresponding to the online article is generated based on the hotspot value set and the hotspot event lexicon; wherein, the hotspot value set includes the hotspot value corresponding to each word; The process involves generating a relevance score corresponding to the online article based on the basic information, and then generating a topic importance score corresponding to the online article based on the relevance score. The relevance score includes time relevance, location relevance, person relevance, and behavior relevance. Specifically, the step of generating the relevance score based on the basic information includes: determining corresponding time information, corresponding location information, corresponding person information, and corresponding behavior information based on the basic information; generating the corresponding time relevance score based on the corresponding time information; generating the corresponding location relevance score based on the corresponding location information; generating the corresponding person relevance score based on the corresponding person information; and generating the corresponding behavior relevance score based on the corresponding behavior information. The step of generating the topic importance score corresponding to the online article based on the relevance score includes: generating a relationship matrix corresponding to the online article based on the relevance score; generating the topic importance score corresponding to the online article based on the relationship matrix. The step of generating the relationship matrix corresponding to the online article based on the relevance score includes: constructing a topic relationship graph corresponding to the online article based on the relevance score; and determining the relationship matrix corresponding to the online article based on the topic relationship graph. Based on the popularity value and the importance of the topic, a related popularity value corresponding to the online article is generated.

2. The method for analyzing the correlation heat of hot events according to claim 1, characterized in that, The step of generating a hotspot value set corresponding to the online article based on the text data includes: Based on the text data, a binary distribution statistical result set corresponding to the online article is generated by performing binary distribution statistics; wherein, the binary distribution statistical result set includes the binary distribution statistical result corresponding to each word; Based on the binary distribution statistical result set, determine the hotspot value set corresponding to the online article.

3. The method for analyzing the correlation heat of hot events according to claim 1, characterized in that, The step of generating a popularity value corresponding to the online article based on the hotspot value set and the preset hotspot event terminology database includes: Based on the hotspot value set, generate a hotspot active term library and a hotspot inert term library corresponding to the online article; A first co-occurrence threshold is generated based on the aforementioned hot topic active term library and the preset hot topic event term library; A second co-occurrence threshold is generated based on the aforementioned hotspot inert lexicon and the preset hotspot event lexicon; The popularity value corresponding to the online article is generated based on the first co-occurrence threshold and the second co-occurrence threshold.

4. The method for analyzing the correlation heat of hot events according to claim 3, characterized in that, The step of generating a hot topic active term library and a hot topic inert term library corresponding to the online article based on the hot topic value set includes: The hot topic active word library corresponding to the online article is generated from words whose values ​​in the hot topic value set are greater than or equal to the upper threshold. The hot topic inert word library corresponding to the online article is generated from words whose values ​​in the hot topic value set are less than or equal to the lower threshold.

5. A device for analyzing the relevance of trending events, the device being used to analyze the relevance of online articles using a terminology database corresponding to the current trending event, characterized in that, include: The data acquisition module is used to acquire text data of online articles corresponding to the current hot topic and basic information corresponding to the online articles; wherein, the online articles include at least one article; A popularity value generation module is used to generate a hotspot value set corresponding to the online article based on the text data, and to generate a popularity value corresponding to the online article based on the hotspot value set and the hotspot event lexicon; wherein, the hotspot value set includes a hotspot value corresponding to each word; The topic importance generation module is used to generate a relevance score corresponding to the online article based on the basic information, and to generate a topic importance score corresponding to the online article based on the relevance score. The relevance score includes time relevance, location relevance, person relevance, and behavior relevance. Specifically, the step of generating a relevance score corresponding to the online article based on the basic information includes: determining corresponding time information, corresponding location information, corresponding person information, and corresponding behavior information based on the basic information; generating a corresponding time relevance score based on the corresponding time information; generating a corresponding location relevance score based on the corresponding location information; generating a corresponding person relevance score based on the corresponding person information; and generating a corresponding behavior relevance score based on the corresponding behavior information. The step of generating a topic importance score corresponding to the online article based on the relevance score includes: generating a relationship matrix corresponding to the online article based on the relevance score; generating a topic importance score corresponding to the online article based on the relationship matrix. The step of generating a relationship matrix corresponding to the online article based on the relevance score includes: constructing a topic relationship graph corresponding to the online article based on the relevance score; and determining a relationship matrix corresponding to the online article based on the topic relationship graph. The related popularity generation module is used to generate related popularity corresponding to the online article based on the popularity value and the importance of the topic.

6. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, it implements the steps of the hot topic event correlation heat analysis method as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the hot topic association heat analysis method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Network information data popularity calculation method

    CN111324789A

  • Social analytics system and method for analyzing conversations in social media

    US20070214097A1