A news event discovery algorithm and device based on heterogeneous information network
By integrating heterogeneous information networks and emotional information processing into the news event discovery algorithm, the problem of failure to fully consider the emotional color of the article in the existing technology is solved, and a more accurate and efficient news recommendation effect is achieved.
Patent Information
- Application Number
- CN202110867857.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-07-28
AI Technical Summary
The prior art fails to fully consider the emotional color of the article in the discovery and recommendation of news events, resulting in inaccurate recommendation results, and the traditional text similarity algorithm has the disadvantage of the same frequency and different degrees of influence of keywords.
A news event discovery algorithm based on heterogeneous information network is proposed, which achieves more accurate news recommendations by preprocessing news on multiple topics, fusion of emotional information, metapathic path or metagraph construction, graph attention network feature extraction and recommendation cluster construction.
By integrating emotional information from the article, the accuracy of news topic recommendations is improved, and the time complexity of model training is reduced through heterogeneous information networks to achieve more efficient news recommendations.
Smart Images

Figure CN113742464B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information networks, and in particular, to a news event discovery algorithm based on a heterogeneous information network. Background Art
[0002] Quickly finding target information from massive text data and real-time tracking of the development trend of current hot topics have gradually become the real needs of users. Topic Detection and Tracking (TDT) technology with real-time detection and tracking as the goal has gradually entered an era of great splendor. Nowadays, in technology enterprises and government departments in order to real-time track the social public opinion orientation, the topic detection and tracking algorithm for news events has become the key research direction of computer researchers. However, the current topic detection or event discovery algorithms do not consider the emotional information of each keyword and cannot recommend articles with the same emotional color. Secondly, the current traditional text similarity algorithm calculates the term frequency-inverse document probability value of the text through TF-IDF, but there will be keywords with the same frequency, which will have different degrees of influence on the documents where they are located. And only through the similarity of keywords for event discovery will also lead to the loss of a large amount of hidden information of the user's emotions and articles. It is impossible to accurately complete the event discovery task.
[0003] Currently similar methods:
[0004] 1. News recommendation that only compares the relevance by calculating the text term frequency according to TF-IDF;
[0005] 2. After extracting keywords, directly recommend through a heterogeneous information network (HIN);
[0006] 3. Articles are recommended according to the graph attention network (GAT).
[0007] However, the current methods have the following disadvantages:
[0008] 1) Do not consider the emotional color of the article well
[0009] For example, for an article about Kobe Bryant's death, the key information is NBA, Kobe Bryant, and death. Its emotion is sad. What the user wants to focus on is why Kobe Bryant died and more reports on this information. Instead of what other NBA stars did during this period and who won the MVP.
[0010] 2) The complexity of the heterogeneous graph neural network
[0011] Although the heterogeneous graph neural network is a framework that can identify node features and semantic features based on multiple meta-paths for processing, this framework needs to specify the number of meta-paths in advance and will train the adjacency matrix of each meta-path together with the same feature matrix through a graph attention network once, which greatly increases the time complexity of model training.
[0012] The generation of heterogeneous graphs is mainly achieved by manually setting the styles of paths, such as N→K←N, where N represents news and K represents keywords. This meta-path indicates that there are the same keywords between news reports, and they are connected through the same keywords. Summary of the Invention
[0013] The present invention aims to solve at least one of the technical problems in the related art to some extent.
[0014] To this end, the first object of the present invention is to propose a news event discovery algorithm based on heterogeneous information networks to achieve more accurate recommendations for users.
[0015] The second object of the present invention is to propose a news event discovery device based on heterogeneous information networks.
[0016] To achieve the above object, the first aspect embodiment of the present invention proposes a news event discovery algorithm based on heterogeneous information networks, including the following steps:
[0017] Step S1: Extract news of multiple topics, preprocess the extracted news, select multiple keywords of the article according to the importance of each keyword, and generate a keyword set according to the multiple keywords;
[0018] Step S2: Integrate the sentiment information of the keyword set, and obtain an event group through prediction by a prediction model;
[0019] Step S3: Construct a meta-path or meta-graph for the event group to obtain a construction matrix, and generate a distance matrix according to the construction matrix;
[0020] Step S4: Extract features of the distance matrix and the event group through a graph attention network to obtain a feature matrix;
[0021] Step S5: Construct a recommendation cluster according to the feature matrix;
[0022] Step S6: Select news in the recommendation cluster that is greater than the preset threshold of the similarity of the original article for recommendation.
[0023] Optionally, in an embodiment of the present application, it is characterized in that
[0024] The preprocessing includes word segmentation of the article through Jieba, and the importance of each keyword is:
[0025] A-TFIDF = TFIDF + W
[0026] where the formula of W is defined as:
[0027] W = n * o
[0028] Wherein, n is the number of the same numerical values in the keywords with the same numerical values, and o is a unified minimum value of 1.0e-16.
[0029] Optionally, in an embodiment of the present application, it is characterized in that the S2 includes:
[0030] Training the prediction model;
[0031] Performing topic prediction on the trained prediction model to obtain a prediction result;
[0032] Corresponding the prediction result with the multiple topics to obtain a corresponding news event group set.
[0033] Optionally, in an embodiment of the present application, it is characterized in that the training of the prediction model includes:
[0034] Performing word embedding on the keyword set to obtain keyword set word vectors;
[0035] Performing word embedding on the sentiment information of the keywords to obtain keyword sentiment information word vectors;
[0036] Performing splicing processing on the keyword set word vectors and the keyword sentiment information word vectors, and performing dimensionality reduction through a fully connected layer;
[0037] Putting the dimensionality-reduced keyword set word vectors and the sentiment information word vectors of the keywords into the prediction model to perform topic prediction.
[0038] Optionally, in an embodiment of the present application, it is characterized in that the S3 includes:
[0039] Constructing a meta-path for the event group by selecting NKN path, NUN path, and NLN path to obtain a meta-path construction matrix;
[0040] Performing meta-graph construction on the event group by selecting NK(L\U)KN to obtain a meta-graph construction matrix;
[0041] Wherein, N represents a news instance, U represents a person name, K represents a keyword, and L represents a location.
[0042] Optionally, in an embodiment of the present application, it is characterized in that the S3 further includes:
[0043] Performing PathSim calculation on the meta-path construction matrix and the meta-graph construction matrix to obtain the distance matrix, and the calculation formula of the distance matrix is:
[0044]
[0045] Optionally, in an embodiment of the present application, it is characterized in that S4 includes:
[0046] When performing feature extraction through the graph attention network, ensure the association existing between the nodes of the graph attention network;
[0047] Use Softmax for normalization operation, and compare the attention coefficients that affect the nodes of the graph attention network, where the formula for the attention coefficient is:
[0048]
[0049] Optionally, in an embodiment of the present application, it is characterized in that S5 includes:
[0050] By adjusting the parameters of the clustering algorithm, so that the recommended clusters reach a preset threshold of accuracy.
[0051] The news event discovery method based on heterogeneous information network of the present invention extracts and preprocesses news of multiple topics at the same time, selects multiple keywords of an article according to the importance of each keyword, and generates a keyword set according to the multiple keywords; fuses the sentiment information of the keyword set, and obtains an event group through prediction by a prediction model; constructs a meta-path or a meta-graph for the event group to obtain a construction matrix, and generates a distance matrix according to the construction matrix; performs feature extraction on the distance matrix and the event group through a graph attention network to obtain a feature matrix; constructs a recommended cluster according to the feature matrix; selects news in the recommended cluster that is greater than a preset threshold of the similarity of the original article for recommendation. The method proposed in the present application can incorporate the sentiment information of the article, improve the accuracy of news topic recommendation to a certain extent; and can also construct a distance matrix through HIN, reducing the time complexity of model training.
[0052] To achieve the above object, an embodiment of the second aspect of the present application proposes an apparatus for discovering news events based on a heterogeneous information network of the present invention, including the following modules:
[0053] A data preprocessing module, configured to extract news of multiple topics, preprocess the extracted news, select multiple keywords of an article according to the importance of each keyword, and generate a keyword set according to the multiple keywords;
[0054] A prediction module, configured to fuse the sentiment information of the keyword set, and obtain an event group through prediction by a prediction model;
[0055] A construction module for constructing a meta-path or meta-graph for the event group to obtain a construction matrix, and calculating the distance matrix by calculating the construction matrix;
[0056] A feature extraction module for extracting features from the distance matrix and the event group through a graph attention network to obtain a feature matrix;
[0057] A feature clustering module for constructing a recommendation cluster according to the feature matrix;
[0058] A recommendation module for selecting news in the recommendation cluster that is greater than a preset threshold of the similarity of the original article for recommendation.
[0059] Optionally, in an embodiment of the present application, it is characterized in that the data preprocessing module includes:
[0060] Performing word segmentation on the article through Jieba word segmentation, and the importance degree of each keyword is:
[0061] A-TFIDF = TFIDF + W
[0062] Where the formula of W is defined as:
[0063] W = n * o
[0064] Where n is the number of the same numerical values in front of the keywords with the same numerical values, and o is a unified minimum value of 1.0e-16.
[0065] The news event discovery device based on heterogeneous information network of the present invention extracts news of multiple topics and performs preprocessing at the same time, selects multiple keywords of the article according to the importance degree of each keyword, generates a keyword set according to the multiple keywords; fuses the emotional information of the keyword set, and obtains an event group through prediction by a prediction model; constructs a meta-path or meta-graph for the event group to obtain a construction matrix, and generates a distance matrix according to the construction matrix; extracts features from the distance matrix and the event group through a graph attention network to obtain a feature matrix; constructs a recommendation cluster according to the feature matrix; selects news in the recommendation cluster that is greater than a preset threshold of the similarity of the original article for recommendation. The method proposed in the present application can integrate the emotional information of the article, improve the accuracy of news topic recommendation to a certain extent; and can also construct a distance matrix through HIN, reducing the time complexity of model training.
[0066] Technical effects of the present application: First, integrating the emotional information of the article can improve the accuracy of news topic recommendation to a certain extent, and thus better recommend news reports; Second, only constructing the distance matrix through HIN and adding emotional color during the formation of the feature matrix reduces the time complexity of model training. The above two points can more accurately recommend to users.
[0067] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the drawings, in which:
[0069] Figure 1 is a schematic flowchart of a news event discovery algorithm based on a heterogeneous information network according to an embodiment of the present invention;
[0070] Figure 2 is a schematic structural diagram of a news event discovery device based on a heterogeneous information network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0072] A news event discovery algorithm based on a heterogeneous information network according to an embodiment of the present invention will be described below with reference to the drawings.
[0073] As Figure 1 shown, to achieve the above object, a news event discovery algorithm based on a heterogeneous information network according to a first aspect embodiment of the present invention includes the following steps:
[0074] Step S1, extracting news of multiple topics, preprocessing the extracted news, selecting multiple keywords of the article according to the importance of each keyword, and generating a keyword set according to the multiple keywords;
[0075] Step S2, fusing the emotional information of the keyword set, and obtaining an event group through prediction by a prediction model;
[0076] Step S3, constructing a meta-path or meta-graph for the event group to obtain a construction matrix, and generating a distance matrix according to the construction matrix;
[0077] Step S4: Extract features from the distance matrix and the event group through a graph attention network to obtain a feature matrix;
[0078] Step S5: Construct a recommendation cluster based on the feature matrix;
[0079] Step S6: Select news in the recommendation cluster that is greater than the preset threshold of the similarity of the original article for recommendation.
[0080] In an embodiment of the present application, the preprocessing includes word segmentation of the article through Jieba word segmentation, and the importance degree of each keyword is:
[0081] A - TFIDF = TFIDF + W
[0082] Where the formula of W is defined as:
[0083] W = n * o
[0084] Where n is the number of the same numerical values in front of the keywords with the same numerical values, and o is a unified minimum value of 1.0e - 16.
[0085] In an embodiment of the present application, further, S2 includes:
[0086] Train the improved model;
[0087] S211: Perform word embedding on the keyword set;
[0088] S212: Perform word embedding on the sentiment information of the keywords;
[0089] S213: Concatenate the above two word vectors and perform dimensionality reduction through a fully connected layer;
[0090] S214: Put the vector after dimensionality reduction into the model for topic prediction;
[0091] S215: Repeat the process of S211 - S214 until the accuracy of topic prediction no longer improves;
[0092] S22: Perform topic prediction on the trained model;
[0093] S23: Correlate the prediction result with the topics in the database to obtain a corresponding news event group set.
[0094] In an embodiment of the present application, further, S3 includes:
[0095] S31: Select the NKN path, NUN path, and NLN path to construct meta - paths for the event group. Where N represents news instances, U represents person names, K represents keywords, and L represents locations.
[0096] S32. Select NK(L\U)KN as the meta-graph for construction, which represents that a news report can be related to another in multiple ways through location and users, indicating stronger relevance of the documents.
[0097] S33. Perform PathSim calculation on the constructed matrix to generate a distance matrix. The calculation formula is:
[0098]
[0099] In an embodiment of the present application, further, S4 includes:
[0100] S41. Perform stronger feature extraction on the distance matrix through a graph attention network to ensure a certain correlation between nodes;
[0101] S42. To compare the attention coefficients that affect the nodes, we use Softmax for normalization operation. The formula is:
[0102]
[0103] In an embodiment of the present application, further, S5 includes:
[0104] By continuously adjusting the parameters eps and min_samples in the DBSCAN algorithm, the accuracy of the recommended clusters is ensured, and articles with high similarity are prevented from becoming noise points.
[0105] Based on the news event discovery algorithm based on heterogeneous information network according to the embodiments of the present application, by extracting news on multiple topics and performing preprocessing at the same time, multiple keywords of the article are selected according to the importance of each keyword, and a keyword set is generated according to the multiple keywords; the sentiment information of the keyword set is fused, and an event group is obtained through prediction by a prediction model; the event group is constructed into a meta-path or meta-graph to obtain a constructed matrix, and a distance matrix is generated according to the constructed matrix; the distance matrix and the event group are subjected to feature extraction through a graph attention network to obtain a feature matrix; a recommended cluster is constructed according to the feature matrix; news in the recommended cluster that is greater than the preset similarity threshold of the original article is recommended. The method proposed in the present application can incorporate the sentiment information of the article, improve the accuracy of news topic recommendation to a certain extent; and can also construct a distance matrix through HIN, reducing the time complexity of model training.
[0106] As Figure 2 shown, to achieve the above object, an embodiment of the second aspect of the present application proposes a news event discovery device 10 based on a heterogeneous information network of the present invention, including the following modules:
[0107] The data preprocessing module 100 is used to extract news on multiple topics, preprocess the extracted news, select multiple keywords of the article according to the importance of each keyword, and generate a keyword set based on the multiple keywords;
[0108] The prediction module 200 is used to fuse the sentiment information of the keyword set and obtain an event cluster through prediction by a prediction model;
[0109] The construction module 300 is used to construct a meta-path or a meta-graph for the event cluster to obtain a construction matrix, and calculate a distance matrix by calculating the construction matrix;
[0110] The feature extraction module 400 is used to extract features from the distance matrix and the event cluster through a graph attention network to obtain a feature matrix;
[0111] The feature clustering module 500 is used to construct a recommendation cluster according to the feature matrix;
[0112] The recommendation module 600 is used to select news in the recommendation cluster that is greater than a preset threshold of the similarity of the original article for recommendation.
[0113] Optionally, in an embodiment of the present application, the above data preprocessing module includes:
[0114] Performing word segmentation on the article through Jieba word segmentation, and the importance of each keyword is:
[0115] A-TFIDF = TFIDF + W
[0116] Where the formula of W is defined as:
[0117] W = n * o
[0118] Where n is the number of the same numerical values in front of the keywords with the same numerical values, and o is a unified minimum value of 1.0e-16.
[0119] The news event discovery device based on the heterogeneous information network according to the embodiments of the present application extracts news on multiple topics and performs preprocessing at the same time, selects multiple keywords of an article according to the importance of each keyword, and generates a keyword set according to the multiple keywords; fuses the sentiment information of the keyword set, and obtains an event group through prediction by a prediction model; constructs a meta-path or a meta-graph for the event group to obtain a construction matrix, and generates a distance matrix according to the construction matrix; extracts features from the distance matrix and the event group through a graph attention network to obtain a feature matrix; constructs a recommendation cluster according to the feature matrix; and selects news in the recommendation cluster that is greater than a preset threshold of the similarity of the original article for recommendation. The method proposed in the present application can incorporate the sentiment information of the article, improving the accuracy of news topic recommendation to a certain extent; and can also construct a distance matrix through HIN, reducing the time complexity of model training.
[0120] In the description of the present specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In the present specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art may combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples.
[0121] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
Claims
1. A news event discovery algorithm based on heterogeneous information network, characterized in that, it includes the following steps: S1. Extract news of multiple topics, preprocess the extracted news, select multiple keywords of the article according to the importance of each keyword, and generate a keyword set according to the multiple keywords; S2. Integrate the sentiment information of the keyword set, and obtain an event group through prediction by a prediction model; S3. Construct a meta-path or meta-graph for the event group to obtain a construction matrix, and generate a distance matrix according to the construction matrix; S4. Extract features from the distance matrix and the event group through a graph attention network to obtain a feature matrix; S5. Construct a recommendation cluster according to the feature matrix; S6. Select news in the recommendation cluster that is greater than a preset threshold of the similarity of the original article for recommendation; The S2 includes: Training the prediction model; Performing topic prediction on the trained prediction model to obtain a prediction result; Corresponding the prediction result with the multiple topics to obtain a corresponding news event group set; The training of the prediction model includes: Performing word embedding on the keyword set to obtain keyword set word vectors; Performing word embedding on the sentiment information of the keyword to obtain keyword sentiment information word vectors; Performing splicing processing on the keyword set word vectors and the keyword sentiment information word vectors, and performing dimensionality reduction through a fully connected layer; Putting the dimensionality-reduced keyword set word vectors and the sentiment information word vectors of the keyword into the prediction model for topic prediction.
2. A news event discovery algorithm based on heterogeneous information network according to claim 1, characterized in that, the preprocessing includes performing word segmentation on the article through Jieba word segmentation, and the importance of each keyword is: A-TFIDF = TFIDF + W where the formula of W is defined as: W = n * o where n is the number of the same previous values among the keywords with the same value, and o is a unified minimum value of 1.0e-16.
3. A news event discovery algorithm based on heterogeneous information network according to claim 1, characterized in that, the S3 includes: Constructing a meta-path for the event group by selecting NKN path, NUN path, NLN path to obtain a meta-path construction matrix; Constructing a meta-graph for the event group by selecting NK(L\U)KN to obtain a meta-graph construction matrix; where N represents a news instance, U represents a person's name, K represents a keyword, and L represents a location.
4. A news event discovery algorithm based on heterogeneous information network according to claim 3, characterized in that, the S3 further includes: Performing PathSim calculation on the meta-path construction matrix and the meta-graph construction matrix to obtain the distance matrix, and the calculation formula of the distance matrix is:
5. A news event discovery algorithm based on heterogeneous information network according to claim 1, characterized in that S4 includes: When performing feature extraction through the graph attention network, ensuring the relevance between the nodes of the graph attention network; Perform a normalization operation using Softmax and compare the attention coefficients that affect the nodes of the graph attention network. Among them, the formula for the attention coefficient is:
6. A news event discovery algorithm based on a heterogeneous information network according to claim 1, characterized in that, S5 includes: Adjust the parameters of the clustering algorithm to make the recommended clusters reach a preset threshold of accuracy.
7. A news event discovery device based on a heterogeneous information network, characterized in that, includes: A data preprocessing module for extracting news on multiple topics, preprocessing the extracted news, selecting multiple keywords of the article according to the importance of each keyword, and generating a keyword set according to the multiple keywords; A prediction module for fusing the sentiment information of the keyword set and predicting an event group through a prediction model; A construction module for constructing a meta-path or a meta-graph for the event group to obtain a construction matrix, and calculating the distance matrix by calculating the construction matrix; A feature extraction module for extracting features of the distance matrix and the event group through a graph attention network to obtain a feature matrix; A feature clustering module for constructing a recommended cluster according to the feature matrix; A recommendation module for selecting news in the recommended cluster that is greater than a preset threshold of the similarity of the original article for recommendation; The prediction module is further used for: Training the prediction model; Performing topic prediction on the trained prediction model to obtain a prediction result; Corresponding the prediction result to the multiple topics to obtain a corresponding news event group set; The training of the prediction model includes: Performing word embedding on the keyword set to obtain a keyword set word vector; Performing word embedding on the sentiment information of the keyword to obtain a keyword sentiment information word vector; Performing splicing processing on the keyword set word vector and the keyword sentiment information word vector, and performing dimensionality reduction through a fully connected layer; Putting the dimensionality-reduced keyword set word vector and the keyword sentiment information word vector into the prediction model for topic prediction.
8. A news event discovery device based on a heterogeneous information network according to claim 7, characterized in that, The data preprocessing module includes: Performing word segmentation on the article through Jieba word segmentation, and the importance of each keyword is: A-TFIDF = TFIDF + W where the formula for W is defined as: W = n * o where n is the number of the same values in front of the keywords with the same value, and o is a unified minimum value of 1.0e-16.
Citation Information
Patent Citations
Hot event detection method and system
CN110232149A
Recommendation method for aggregating knowledge graph neural network and adaptive attention
CN112989064A