Hot Event Discovery Method, Device, Equipment and Storage Medium

Through the collection of articles by network crawlers, combined with deep learning models and event graph segmentation technology, the problem of poor clustering of hot-spot events in the existing technology is solved, and more efficient hot-spot event discovery is achieved.

CN111291182BActive Publication Date: 2025-06-24PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010033828.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-01-13
Publication Date
2025-06-24
Estimated Expiration
2040-01-13

AI Technical Summary

Technical Problem

The existing unsupervised text clustering algorithm is not effective in hot event discovery. Similar articles that aggregate may not necessarily talk about the same event, and articles of the same event cannot be effectively aggregated due to different expressions.

Method used

Articles are collected using network crawler technology, rough clustering is performed through preset clustering algorithms, text pairs are constructed and feature extraction and similarity calculation are calculated using deep learning models, event graphs are constructed and segmented to accurately cluster hotspot events.

Benefits of technology

It improves the accuracy of clustering of hot events, ensures that articles of the same event can be effectively aggregated, and improves the practicality of hot events discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111291182B_ABST
    Figure CN111291182B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for discovering hot events, including: collecting articles published on a specified website; using a preset clustering algorithm to cluster all articles to obtain a rough clustering result; successively taking articles within the same category from the rough clustering result and pairing them two by two to construct text pairs; preprocessing each text pair and then successively inputting it into a preset text pair model for processing, outputting the similarity between any two articles and whether they belong to the same event; constructing an event graph with each article as a vertex of the graph, connecting the articles of the same event two by two as the edges of the graph, and using the similarity between the articles as the weight of the corresponding edge; segmenting the event graph to obtain multiple sub-event graphs, and the articles corresponding to all vertices in the same sub-event graph are the articles corresponding to the same hot event. The present invention also discloses a device, equipment and computer-readable storage medium for discovering hot events. The present invention effectively improves the clustering accuracy of hot events on the Internet.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment and storage medium for discovering hot events. Background Art

[0002] Hot events refer to events that are widely discussed or spread in society (the Internet). Discovering hot events is an important field of natural language processing. Discovering hot events in a large number of Internet articles requires a series of text processing and the use of multiple algorithms and models for mining. Discovering hot events plays an important role in aspects such as public opinion monitoring, discovery of customer marketing opportunities, intelligent recommendation, and public opinion guidance.

[0003] Traditional hot event discovery algorithms mainly use unsupervised text clustering algorithms. For example, after using the TF-IDF algorithm to extract document word frequency features, clustering algorithms such as K-means or LDA are used to aggregate similar files together to form hot event articles. However, there is a common problem with unsupervised text clustering algorithms, that is, the aggregated similar articles may not necessarily be about the same event, or due to differences in text expressions within the same event, they are not aggregated together. Therefore, the clustering effect is not good, that is, the practicality of extracting hot events through existing unsupervised text clustering algorithms is not high. Summary of the Invention

[0004] The main purpose of the present invention is to provide a method, device, equipment and storage medium for discovering hot events, aiming to solve the technical problem of how to improve the clustering accuracy of hot events.

[0005] To achieve the above object, the present invention provides a method for discovering hot events, and the method for discovering hot events includes the following steps:

[0006] Collect articles published on a specified website using web crawler technology;

[0007] Based on the number of crawled articles, determine the number of categories for rough clustering, and use a preset clustering algorithm to cluster all articles to obtain a rough clustering result;

[0008] Take articles within the same category from the rough clustering result in pairs in sequence to construct text pairs;

[0009] Preprocess each text pair and then input it into a preset text pair model for processing, and output the similarity between any two articles and whether they belong to the same event;

[0010] Construct an event graph with each article as a vertex of the graph, connect articles of the same event in pairs as edges of the graph, and use the similarity between articles as the weight of the corresponding edges;

[0011] Segment the event graph to obtain multiple sub-event graphs, where all vertices in the same sub-event graph correspond to articles of the same hot event.

[0012] Optionally, before the step of collecting articles published on a specified website using web crawler technology, it further includes:

[0013] Obtain training samples for training a text pair model, where the training samples are text pairs and include positive samples and negative samples. A text pair contains two articles. Positive samples are obtained by pairing different articles within the same event pairwise, and negative samples are obtained by pairing articles between different events pairwise with sampling and by pairing articles between similar events pairwise;

[0014] Perform word segmentation on the two articles in each text pair to obtain multiple independent words or characters;

[0015] Use a preset dictionary or lexicon to encode each word or character to obtain character encoding vectors corresponding to the two articles in each text pair;

[0016] Input the character encoding vectors corresponding to the two articles in each text pair into an embedding layer to convert them into matrix vectors, and input the matrix vectors corresponding to the two articles in the same text pair into two independent and same-layer convolutional layers for feature extraction;

[0017] Input the features extracted by the two convolutional layers into a pooling layer respectively to calculate the similarity between the features corresponding to the two articles in the same text pair and to reduce the dimension of the extracted features;

[0018] Input the similarity between the features corresponding to the two articles in each text pair and the reduced-dimensional features into a fully connected layer for classification and normalization processing to obtain the text pair model.

[0019] Optionally, the segmenting the event graph to obtain multiple sub-event graphs includes:

[0020] Initialize the event graph to divide each vertex into different partitions, where the initial number of partitions is the same as the number of vertices;

[0021] Perform a partition trial division on each vertex one by one to divide each vertex into the partition where its adjacent neighbor vertices are located, calculate the modularity change value of the event graph corresponding to each vertex before and after the division, and record the neighbor vertex corresponding to the maximum modularity change value;

[0022] If the maximum modularity change value is greater than 0, divide the corresponding vertex into the partition where the neighbor vertex corresponding to the maximum modularity change value is located, otherwise abandon the current vertex trial division;

[0023] Repeat the processing flow of partition trial division until the partitions corresponding to all vertices no longer change;

[0024] Compress all vertices within the same partition into a new vertex to construct a new event graph, and set the weight of the edges between vertices within the same partition to the weight of the loop of the new vertex and set the weight of the edges between different partitions to the weight of the edges between the new vertices;

[0025] Repeat the processing flow of constructing a new event graph until the modularity of the entire event graph no longer changes, where one partition corresponds to a sub-event graph.

[0026] Optionally, determining the number of categories for rough clustering based on the number of crawled articles, and using a preset clustering algorithm to cluster all articles to obtain a rough clustering result, including:

[0027] Determine the number of categories for rough clustering based on the number of crawled articles;

[0028] Perform word segmentation on the crawled articles to obtain multiple independent words or characters;

[0029] Convert the words or characters after word segmentation into word vectors, and perform rough clustering on the word vectors corresponding to each crawled article using a preset clustering algorithm according to the number of categories to obtain a rough clustering result.

[0030] Optionally, after the step of successively taking articles within the same category from the rough clustering result and pairing them in pairs to construct text pairs, further include:

[0031] Obtain the titles of the articles in each constructed text pair in the same category;

[0032] Judge whether the titles of the articles in the same text pair are the same;

[0033] If they are the same, retain the corresponding text pair, otherwise eliminate the corresponding text pair.

[0034] Furthermore, to achieve the above object, the present invention also provides a hot event discovery device, and the hot event discovery device includes:

[0035] A collection module, configured to collect articles published on a specified website using web crawler technology;

[0036] A clustering module, configured to determine the number of categories for rough clustering based on the number of crawled articles, and use a preset clustering algorithm to cluster all articles to obtain a rough clustering result;

[0037] A pairing module, configured to successively take articles within the same category from the rough clustering result and pair them in pairs to construct text pairs;

[0038] A model processing module, configured to preprocess each text pair and then sequentially input it into a preset text pair model for processing, and output the similarity between any two articles and whether they belong to the same event;

[0039] A graph construction module, configured to use each article as a vertex of the graph, connect the articles of the same event in pairs as the edges of the graph, and use the similarity between the articles as the weight of the corresponding edges to construct an event graph;

[0040] A segmentation module, configured to segment the event graph to obtain a plurality of sub-event graphs, wherein all the vertices in the same sub-event graph correspond to the articles of the same hot event.

[0041] Optionally, the hot event discovery device further includes:

[0042] An acquisition module, configured to acquire training samples for training the text pair model, wherein the training samples are text pairs and include positive samples and negative samples. A text pair contains two articles. The positive samples are obtained by pairing different articles within the same event in pairs, and the negative samples are obtained by pairing articles between different events in pairs and taking samples, as well as pairing articles between similar events in pairs;

[0043] A word segmentation module, configured to perform word segmentation on the two articles in each text pair to obtain a plurality of independent words or characters;

[0044] An encoding module, configured to encode each word or character by using a preset dictionary or lexicon to obtain the character encoding vectors corresponding to the two articles in each text pair;

[0045] A model construction module, configured to input the character encoding vectors corresponding to the two articles in each text pair into an embedding layer to be converted into matrix vectors, and input the matrix vectors corresponding to the two articles in the same text pair into two independent and same-layer convolutional layers for feature extraction respectively; input the features extracted by the two convolutional layers into a pooling layer respectively to calculate the similarity between the features corresponding to the two articles in the same text pair and reduce the dimension of the extracted features; input the similarity between the features corresponding to the two articles in each text pair and the reduced features into a fully connected layer for classification and normalization processing to obtain the text pair model.

[0046] Optionally, the segmentation module includes:

[0047] An initialization unit, configured to initialize the event graph to divide each vertex into different partitions, wherein the initial number of partitions is the same as the number of vertices;

[0048] A partition division unit is used to perform partition trial division on each vertex one by one, so as to divide each vertex into the partition where its adjacent neighbor vertices are located, calculate the modularity change value of the event graph corresponding to each vertex before and after the division, and record the neighbor vertex corresponding to the maximum modularity change value; if the maximum modularity change value is greater than 0, then divide the corresponding vertex into the partition where the neighbor vertex corresponding to the maximum modularity change value is located, otherwise abandon the current vertex trial division; repeat the processing flow of partition trial division until the partitions corresponding to all vertices no longer change;

[0049] A new graph construction unit is used to compress all vertices in the same partition into a new vertex to construct a new event graph, and set the weight of the edges between the vertices in the same partition as the weight of the loop of the new vertex and set the weight of the edges between different partitions as the weight of the edges between the new vertices; repeat the processing flow of constructing a new event graph until the modularity of the entire event graph no longer changes, where a partition corresponds to a sub-event graph.

[0050] Optionally, the clustering module is specifically used for:

[0051] Determine the number of categories for rough clustering based on the number of crawled articles;

[0052] Perform word segmentation on the crawled articles to obtain multiple independent words or characters;

[0053] Convert the words or characters after word segmentation into word vectors, and perform rough clustering on the word vectors corresponding to each crawled article by using a preset clustering algorithm according to the number of categories to obtain a rough clustering result.

[0054] Optionally, the hot event discovery device further includes:

[0055] A text pair elimination module is used to obtain the titles of the articles in each constructed text pair in the same category; judge whether the titles of the articles in the same text pair are the same; if they are the same, then retain the corresponding text pair, otherwise eliminate the corresponding text pair.

[0056] Furthermore, to achieve the above object, the present invention also provides a hot event discovery device, which includes a memory, a processor, and a hot event discovery program stored on the memory and executable on the processor. When the hot event discovery program is executed by the processor, the steps of the hot event discovery method described in any one of the above are implemented.

[0057] Furthermore, to achieve the above object, the present invention also provides a computer-readable storage medium, on which a hot event discovery program is stored. When the hot event discovery program is executed by a processor, the steps of the hot event discovery method described in any one of the above are implemented.

[0058] Based on text preprocessing and traditional clustering methods, the present invention adds a classification model based on deep learning for accurate classification, which improves the accuracy of event discovery and can be trained using corpora according to requirements. The text pair model of the present invention can learn for text pair matching tasks to learn the important features for distinguishing two texts in the text pair and calculate the similarity of the features, so as to determine whether the two texts are about the same event, and at the same time, the similarity of the text pair can be output for further use in graph-based aggregation algorithms. The present invention effectively improves the effect of discovering hot events on the Internet, and the discovered event results can be applied to multiple fields such as public opinion monitoring, customer marketing opportunity discovery, and intelligent recommendation. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a schematic structural diagram of the operating environment of the hot event discovery device related to the embodiment solution of the present application;

[0060] Figure 2 It is a schematic flowchart of the first embodiment of the hot event discovery method of the present invention;

[0061] Figure 3 It is a schematic flowchart of the second embodiment of the hot event discovery method of the present invention;

[0062] Figure 4 For Figure 2 It is a detailed flowchart of an embodiment of step S160 in

[0063] Figure 5 For Figure 2 It is a detailed flowchart of an embodiment of step S120 in

[0064] Figure 6 It is a schematic flowchart of the third embodiment of the hot event discovery method of the present invention;

[0065] Figure 7 It is a schematic diagram of the functional modules of the first embodiment of the hot event discovery device of the present invention;

[0066] Figure 8 It is a schematic diagram of the functional modules of the second embodiment of the hot event discovery device of the present invention.

[0067] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0069] The present invention provides a hot event discovery device.

[0070] Refer to Figure 1 , Figure 1 which is a schematic structural diagram of the operating environment of the hot event discovery device involved in the solution of the embodiment of the present application.

[0071] As Figure 1 shown, the hot event discovery device includes: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may optionally also be a storage device independent of the aforementioned processor 1001.

[0072] Those skilled in the art can understand that Figure 1 the hardware structure of the hot event discovery device shown in does not constitute a limitation on the hot event discovery device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0073] As Figure 1 shown, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a computer program. Among them, the operating system is a program for managing and controlling the hot event discovery device and software resources, and supports the operation of the hot event discovery program and other software and / or programs.

[0074] In Figure 1 the hardware structure of the hot event discovery device shown, the network interface 1004 is mainly used to access the network; the user interface 1003 is mainly used to detect and confirm instructions and edit instructions, etc. The processor 1001 may be used to call the hot event discovery program stored in the memory 1005 and execute the operations of the following embodiments of the hot event discovery method.

[0075] Based on the above hardware structure of the hot event discovery device, various embodiments of the hot event discovery method of the present application are proposed.

[0076] Refer to Figure 2 , Figure 2 which is a schematic flowchart of the first embodiment of the hot event discovery method of the present invention. In this embodiment, the hot event discovery method includes the following steps:

[0077] Step S110, use web crawler technology to collect articles published on a specified website;

[0078] In this embodiment, it is preferably to directly collect the articles published on the specified website for hot event recognition, and there is no limit to the crawling method of the text published on the website. Preferably, Docker containers are used as the medium to deploy the specified crawler program on multiple machines to achieve crawling of specified content on multiple machines.

[0079] Step S120, determine the number of categories for rough clustering based on the number of crawled articles, and use a preset clustering algorithm to cluster all articles to obtain a rough clustering result;

[0080] In this embodiment, to obtain a better clustering effect, it is preferably to preset the correspondence between the number of articles and the number of clustering categories in advance. For example, 10,000 articles can be clustered into more than 10 categories, and more than 100,000 articles can be clustered into 100 categories, so as to ensure that the number of articles in each category after clustering is less than 1,000.

[0081] In this embodiment, through the preset clustering algorithm, a large number of articles can be roughly divided into several categories. Although the rough clustering division can form hot event articles, there may be cases where the aggregated similar articles do not describe the same event, or the same event is not aggregated in the same category due to differences in text expressions. Therefore, the effect of the rough clustering result needs to be further optimized.

[0082] Step S130, take out the articles within the same category in the rough clustering result in pairs in turn to construct text pairs;

[0083] In this embodiment, a text pair specifically refers to a small group composed of two articles. For example, if there are three articles A, B, and C in a certain category, the constructed text pairs include: (A, B), (A, C), (B, C).

[0084] In this embodiment, only the articles within the same category are paired in pairs to construct text pairs. Articles in different categories usually do not belong to the same or similar events. Therefore, there is no need to construct text pairs.

[0085] Step S140, preprocess each text pair and then input it into a preset text pair model for processing, and output the similarity between any two articles and whether they belong to the same event;

[0086] In this embodiment, after constructing the text pairs, each text pair is preprocessed first, such as performing word segmentation and converting it into word vectors, and then the data of each preprocessed text pair is input into the text pair model in sequence for related processing, such as feature extraction and feature similarity comparison, and then the similarity between any two articles is output and it is determined whether they belong to the same event.

[0087] In this embodiment, the text pair model is pre-trained by machine learning. Through this model, feature extraction and feature similarity comparison can be performed on the content of two articles, so as to identify whether the two articles belong to the same event. The text pair model can not only identify whether two articles in the same clustering category belong to the same event, but also identify whether two articles between different clustering categories belong to the same event.

[0088] Step S150: Construct an event graph with each article as the vertex of the graph, the articles of the same event connected pairwise as the edges of the graph, and the similarity between the articles as the weight of the corresponding edges.

[0089] In this embodiment, the text pair model can identify the similarity between any two articles in the same category and different categories and whether they belong to the same event, but it cannot identify all the articles belonging to the same event as a whole. Therefore, in this embodiment, the method of constructing an event graph is adopted to realize hot event clustering again for the articles corresponding to multiple same events, so as to determine the articles corresponding to the same hot event.

[0090] In this embodiment, an event graph is constructed with all articles as the vertices of the graph, the articles of the same event connected pairwise as the edges of the graph, and the similarity between the articles as the weight of the corresponding edges.

[0091] Step S160: Segment the event graph to obtain multiple sub-event graphs, where all the vertices in the same sub-event graph correspond to the articles of the same hot event.

[0092] In this embodiment, after constructing an event graph containing all articles and the articles corresponding to the same event, through a preset graph segmentation algorithm, based on the degree of association between the vertices, the event graph is segmented, so as to obtain multiple sub-event graphs, and all the vertices in the same sub-event graph correspond to the articles of the same hot event, so as to further realize hot event clustering on the basis of multiple same events.

[0093] Based on text preprocessing and traditional clustering methods, this embodiment adds a classification model based on deep learning for accurate classification, which improves the accuracy of event discovery and can be trained using a corpus according to requirements. The text pair model in this embodiment can learn for the text pair matching task to learn the important features for distinguishing two texts in the text pair and calculate the similarity of the features, so as to determine whether the two texts are about the same event, and at the same time, the similarity of the text pair can be output for further use in graph-based aggregation algorithms. This embodiment can effectively improve the effect of discovering hot events on the Internet and apply the discovered event results to multiple fields such as public opinion monitoring, customer marketing opportunity discovery, and intelligent recommendation.

[0094] Referring to Figure 3 , Figure 3 is a schematic flowchart of the second embodiment of the hot event discovery method of the present invention. In this embodiment, before the above step S110, it further includes:

[0095] Step S210, obtaining training samples for training the text pair model, where the training samples are text pairs and include positive samples and negative samples. A text pair contains two articles. The positive samples are obtained by pairing different articles within the same event pairwise, and the negative samples are obtained by pairing articles between different events pairwise and taking samples, as well as pairing articles between similar events pairwise;

[0096] In this embodiment, each training sample is a text pair containing two articles. Using text pairs as training samples can facilitate feature comparison to determine whether they belong to the same event.

[0097] In this embodiment, there are two types of positive and negative samples:

[0098] (1) Positive samples

[0099] The positive samples are obtained by pairing different articles within the same event pairwise. For example, if there are three articles A, B, and C all reporting the same event, then these three articles can be used as positive samples, and specifically, three samples (A, B), (A, C), and (B, C) can be formed.

[0100] (2) Negative samples

[0101] The negative samples are obtained by pairing articles between different events pairwise and taking samples, and all pairs of articles between similar events are selected as negative samples.

[0102] Step S220, performing word segmentation on the two articles in each text pair to obtain multiple independent words or characters;

[0103] Step S230, encoding each word or character using a preset dictionary or dictionary to obtain the character encoding vectors corresponding to the two articles in each text pair;

[0104] In this embodiment, for the convenience of feature extraction, it is necessary to first perform word segmentation on the articles in each text pair, so as to decompose the articles into multiple independent words or characters. Then, a pre-set dictionary or lexicon is used to encode the words or characters obtained by word segmentation, so as to form a character encoding vector, which is convenient for feature extraction of the text.

[0105] Step S240: Input the character encoding vectors corresponding to the two articles in each text pair into the embedding layer to convert them into matrix vectors, and input the matrix vectors corresponding to the two articles in the same text pair into two independent and same-layer convolutional layers respectively for feature extraction.

[0106] Step S250: Input the features extracted by the two convolutional layers into the pooling layer respectively to calculate the similarity between the features corresponding to the two articles in the same text pair and to reduce the dimension of the extracted features.

[0107] Step S260: Input the similarity between the features corresponding to the two articles in each text pair and the dimension-reduced features into the fully connected layer for classification and normalization processing to obtain the text pair model.

[0108] In this embodiment, it is preferably to use the PairCNN model to perform deep learning on the text pair. First, the character encoding vector is converted into a matrix vector through an embedding layer, and then two CNN networks with the same parameters are used to perform feature extraction on the two texts respectively. Then, the similarity of the output of the pooling layer is calculated, and the similarity and the features output by the pooling layer are used as the input of the classifier. The classifier uses a fully connected layer, and the output of the fully connected layer calculates the Softmax and then calculates the Loss with the label of the sample. The model parameters are iteratively updated using the stochastic gradient descent algorithm, so as to finally obtain the text pair model.

[0109] In this embodiment, text pairs are introduced as training samples, and two independent and same-layer convolutional layers are used for feature extraction during the model training process, so as to realize feature extraction and similarity comparison between different articles, providing a technical basis for the clustering of hot events. The text pair model of this embodiment not only performs feature recognition between any two articles in the same category, but also performs feature recognition between any two articles in different categories, so as to improve the accuracy of hot event discovery.

[0110] Refer to Figure 4 , Figure 4 For Figure 2 a detailed process schematic diagram of an embodiment of step S160 in

[0111] Step S1601, initialize the event graph to partition each vertex into different partitions, where the initial number of partitions is the same as the number of vertices;

[0112] In this embodiment, before partitioning the event graph, the event graph is first initialized. Specifically, each vertex in the graph is regarded as an independent partition, and the number of initial partitions is the same as the number of vertices. For example, if there are 100 vertices in the event graph, the event graph is divided into 100 partitions, and one vertex corresponds to one partition.

[0113] Step S1602, perform a partition trial division on each vertex one by one to partition each vertex into the partition where its adjacent neighbor vertices are located, calculate the modularity change value of the event graph corresponding to each vertex before and after the partition, and record the neighbor vertex corresponding to the maximum modularity change value;

[0114] In this embodiment, a metric for evaluating the quality of the partition network is introduced: modularity. Through the change value of modularity, the quality of the vertex partition method can be reflected. The calculation formula of modularity is as follows:

[0115]

[0116] Among them, W represents modularity, m represents an intermediate parameter, c represents a partition, and i, j represent vertices i, j in partition c; A ij represents the weight between vertex i and vertex j, ∑in represents the sum of the weights of the edges in partition c, and ∑tot represents the sum of the weights of all the edges connected to the vertices in partition c.

[0117] In this embodiment, for each vertex, try to assign the vertex to the partition where its neighbor vertices are located, calculate the modularity change value of the event graph before and after the assignment, and record the neighbor vertex with the largest modularity change value.

[0118] Step S1603, if the maximum modularity change value is greater than 0, partition the corresponding vertex into the partition where the neighbor vertex corresponding to the maximum modularity change value is located, otherwise abandon the current vertex trial partition;

[0119] Step S1604, repeat the processing flow of the partition trial division until the partitions corresponding to all vertices no longer change;

[0120] In this embodiment, if the maximum modularity change value is greater than 0, that is, the corresponding partition method is better, so the current trial partition is feasible. Therefore, the corresponding vertex is partitioned into the partition where the neighbor vertex corresponding to the maximum modularity change value is located, otherwise the current vertex trial partition is abandoned.

[0121] In this embodiment, after one round of tentative partitioning for all vertices is completed, the next round of tentative partitioning for all vertices is continued, that is, S1602 and S1603 are repeatedly executed until the tentative partitioning stops when the partitions corresponding to all vertices no longer change.

[0122] Step S1605: Compress all vertices within the same partition into a new vertex to construct a new event graph, set the weight of the edges between vertices within the same partition as the weight of the loop of the new vertex, and set the weight of the edges between different partitions as the weight of the edges between the new vertices.

[0123] Step S1606: Repeatedly execute the processing flow of constructing a new event graph until the modularity of the entire event graph no longer changes, where one partition corresponds to a sub-event graph.

[0124] In this embodiment, after the partition tentative partitioning ends, that is, the current event graph can no longer be divided. At this time, the original event graph needs to be compressed, and then a new event graph is constructed. Then, the partition tentative partitioning of the new event graph is repeated, that is, S1601 - S1605 are repeatedly executed until the modularity of the entire event graph no longer changes. That is, if the modularity corresponding to the new constructed event graph no longer changes during the partition tentative partitioning, the event graph segmentation stops. Each partition in the last event graph corresponds to a sub-event graph, and all vertices corresponding to the articles in the same sub-event graph are the articles corresponding to the same hot event. Thus, based on multiple identical events, hot event clustering is further realized.

[0125] Refer to Figure 5 , Figure 5 For Figure 2 a refinement process schematic diagram of step S120 in an embodiment in

[0126] Step S1201: Determine the number of categories for rough clustering based on the number of crawled articles.

[0127] Step S1202: Perform word segmentation on the crawled articles to obtain multiple independent words or characters.

[0128] Step S1203: Convert the words or characters after word segmentation into word vectors, and use a preset clustering algorithm to perform rough clustering on the word vectors corresponding to each crawled article according to the number of categories to obtain a rough clustering result.

[0129] To obtain a better clustering effect, it is therefore preferable to preset the correspondence between the number of articles and the number of clustering categories in advance. For example, 10,000 articles can be clustered into more than 10 categories, and more than 100,000 articles can be clustered into 100 categories, so as to ensure that the number of articles in each category after clustering is less than 1000.

[0130] In this embodiment, through techniques such as Jieba word segmentation, the sentences in the article are decomposed into multiple independent words or characters, so as to facilitate the subsequent conversion of the words or characters after word segmentation into word vectors. There is no limit to the conversion method of word vectors in this embodiment. For example, each word or character is encoded based on a preset dictionary or lexicon, and then the word or character encoding is converted into a word vector by using methods such as word2vec.

[0131] In this embodiment, the preset clustering method can use traditional K-means or LDA topic models.

[0132] In this embodiment, through the preset clustering algorithm, a large number of articles can be roughly divided into several categories. Although the rough clustering division can form articles on hot events, there may be a situation where the aggregated similar articles do not describe the same event, or the same event is not aggregated in the same category due to differences in text expressions. Therefore, it is necessary to further optimize the effect of the rough clustering results.

[0133] Refer to Figure 6 , Figure 6 is a schematic flowchart of the third embodiment of the hot event discovery method of the present invention. In this embodiment, after the above step S130, the following steps are further included:

[0134] Step S310, obtaining the titles of the articles in each constructed text pair in the same category;

[0135] Step S320, determining whether the titles of the articles in the same text pair are the same;

[0136] Step S330, if they are the same, retaining the corresponding text pair, otherwise removing the corresponding text pair.

[0137] In this embodiment, to reduce the model input and improve the model processing efficiency, before inputting the constructed text pairs into the model for processing, each text pair is filtered, specifically by comparing the article titles to remove completely dissimilar text pairs.

[0138] There is no limit to the method for determining whether the article titles are the same in this embodiment. For example, by comparing keywords, if the keywords in the article titles are exactly the same, it is determined that the article titles are the same. Or by using a semantic recognition method to determine whether the titles of the two articles in the text pair are the same. If they are the same, they are retained as the objects to be classified, otherwise the text pair corresponding to the title is removed.

[0139] The present invention also provides a hot event discovery device.

[0140] Refer to Figure 7 , Figure 7Schematic diagram of the functional modules of the first embodiment of the hot event discovery device of the present invention. In this embodiment, the hot event discovery device includes:

[0141] A collection module 10, configured to collect articles published on a specified website by using web crawler technology;

[0142] A clustering module 20, configured to determine the number of categories for rough clustering based on the number of crawled articles, and use a preset clustering algorithm to cluster all articles to obtain a rough clustering result;

[0143] A pairing module 30, configured to sequentially pair up articles within the same category from the rough clustering result to construct text pairs;

[0144] A model processing module 40, configured to preprocess each text pair and then sequentially input it into a preset text pair model for processing, and output the similarity between any two articles and whether they belong to the same event;

[0145] A graph construction module 50, configured to construct an event graph with each article as a vertex of the graph, connect pairwise articles of the same event as edges of the graph, and use the similarity between articles as the weight of the corresponding edges;

[0146] A segmentation module 60, configured to segment the event graph to obtain a plurality of sub-event graphs, where all vertices in the same sub-event graph correspond to articles of the same hot event.

[0147] Based on the same embodiment description content as the above-mentioned hot event discovery method of the present invention, therefore, the embodiment content of the hot event discovery device in this embodiment will not be elaborated too much.

[0148] In this embodiment, on the basis of text preprocessing and traditional clustering methods, a classification model based on deep learning is added for accurate classification, so as to improve the accuracy of event discovery, and the corpus can be used for training according to requirements. The text pair model in this embodiment can learn for the text pair matching task to learn the important features used to distinguish two texts in the text pair and calculate the similarity of the features, so as to judge whether the two texts are about the same event, and at the same time, the similarity of the text pair can be output for further use in the graph-based aggregation algorithm. This embodiment can effectively improve the effect of discovering hot events on the Internet and apply the discovered event results to multiple fields such as public opinion monitoring, customer marketing opportunity discovery, and intelligent recommendation.

[0149] Refer to Figure 8 , Figure 8 Schematic diagram of the functional modules of the second embodiment of the hot event discovery device of the present invention. Based on the first embodiment of the above device, in this embodiment, the hot event discovery device further includes:

[0150] An acquisition module 70 is configured to acquire training samples for training a text pair model. The training samples are text pairs and include positive samples and negative samples. A text pair contains two articles. The positive samples are obtained by pairing different articles within the same event pairwise, and the negative samples are obtained by pairing articles between different events pairwise with sampling and pairing articles between similar events pairwise.

[0151] A word segmentation module 80 is configured to perform word segmentation on the two articles in each text pair to obtain a plurality of independent words or characters.

[0152] An encoding module 90 is configured to encode each word or character using a preset dictionary or lexicon to obtain character encoding vectors corresponding to the two articles in each text pair.

[0153] A model construction module 100 is configured to input the character encoding vectors corresponding to the two articles in each text pair into an embedding layer to be converted into matrix vectors, and input the matrix vectors corresponding to the two articles in the same text pair into two independent and same-layer convolutional layers respectively for feature extraction; input the features extracted by the two convolutional layers into a pooling layer respectively to calculate the similarity between the features corresponding to the two articles in the same text pair and reduce the dimension of the extracted features; input the similarity between the features corresponding to the two articles in each text pair and the reduced features into a fully connected layer for classification and normalization processing to obtain the text pair model.

[0154] Based on the same embodiment description content as the above-mentioned hot event discovery method of the present invention, the embodiment content of the hot event discovery device is not described in detail in this embodiment.

[0155] Optionally, in a specific embodiment, the segmentation module includes:

[0156] An initialization unit is configured to initialize the event graph to divide each vertex into different partitions, where the initial number of partitions is the same as the number of vertices.

[0157] A partition division unit is configured to perform partition trial division on each vertex one by one to divide each vertex into the partition where the neighbor vertex adjacent to the vertex is located, calculate the modularity change value of the event graph corresponding to each vertex before and after the division, and record the neighbor vertex corresponding to the maximum modularity change value; if the maximum modularity change value is greater than 0, divide the corresponding vertex into the partition where the neighbor vertex corresponding to the maximum modularity change value is located, otherwise abandon the current vertex trial division; repeat the processing flow of the partition trial division until the partitions corresponding to all vertices no longer change.

[0158] A new graph construction unit is used to compress all vertices within the same partition into a new vertex to construct a new event graph, set the weight of the edges between vertices within the same partition as the weight of the loop of the new vertex, and set the weight of the edges between different partitions as the weight of the edges between the new vertices; repeat the process of constructing the new event graph until the modularity of the entire event graph no longer changes, where a partition corresponds to a sub-event graph.

[0159] Optionally, in a specific embodiment, the clustering module is specifically configured to:

[0160] Determine the number of categories for rough clustering based on the number of crawled articles;

[0161] Perform word segmentation on the crawled articles to obtain multiple independent words or characters;

[0162] Convert the words or characters after word segmentation into word vectors, and based on the number of categories, use a preset clustering algorithm to perform rough clustering on the word vectors corresponding to each crawled article to obtain a rough clustering result.

[0163] Optionally, in a specific embodiment, the hot event discovery device further includes:

[0164] A text pair elimination module is used to obtain the titles of the articles in each constructed text pair in the same category; determine whether the titles of the articles in the same text pair are the same; if they are the same, retain the corresponding text pair, otherwise eliminate the corresponding text pair.

[0165] The present invention also provides a non-volatile computer-readable storage medium.

[0166] In this embodiment, a hot event discovery program is stored on the computer-readable storage medium. When the hot event discovery program is executed by a processor, the steps of the hot event discovery method described in any of the above embodiments are implemented. Among them, the method implemented when the hot event discovery program is executed by the processor can refer to the various embodiments of the hot event discovery method of the present invention, so it will not be elaborated here.

[0167] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM) and includes several instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0168] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, all fall within the protection scope of the present invention.

Claims

1. A method for discovering hot events, characterized in that, The hot event discovery method includes the following steps: Collect articles published on a specified website using web crawler technology; Determine the number of categories for rough clustering based on the number of crawled articles, and use a preset clustering algorithm to cluster all articles to obtain a rough clustering result; Take articles within the same category from the rough clustering result in sequence and pair them up two by two to construct text pairs; Preprocess each text pair and then input it into a preset text pair model for processing, and output the similarity between any two articles and whether they belong to the same event; Construct an event graph with each article as a vertex of the graph, the articles of the same event connected by a line as an edge of the graph, and the similarity between articles as the weight of the corresponding edge; Segment the event graph to obtain multiple sub-event graphs, where all vertices corresponding to articles in the same sub-event graph are articles corresponding to the same hot event; The segmenting the event graph to obtain multiple sub-event graphs includes: Initialize the event graph to divide each vertex into different partitions, where the initial number of partitions is the same as the number of vertices; Perform partition trial division on each vertex one by one to divide each vertex into the partition where its adjacent neighbor vertices are located, calculate the modularity change value of the event graph corresponding to each vertex before and after the division, and record the neighbor vertex corresponding to the maximum modularity change value; If the maximum modularity change value is greater than 0, divide the corresponding vertex into the partition where the neighbor vertex corresponding to the maximum modularity change value is located, otherwise abandon the current vertex trial division; Repeat the processing flow of partition trial division until the partitions corresponding to all vertices no longer change; Compress all vertices in the same partition into a new vertex to construct a new event graph, and set the weight of the edge between vertices in the same partition as the weight of the loop of the new vertex and set the weight of the edge between different partitions as the weight of the edge between new vertices; Repeat the processing flow of constructing a new event graph until the modularity of the entire event graph no longer changes, where one partition corresponds to one sub-event graph.

2. The hot event discovery method according to claim 1, wherein Before the step of collecting articles published on a specified website using web crawler technology, it further includes: Obtain training samples for training the text pair model, where the training samples are text pairs and include positive samples and negative samples. A text pair contains two articles. Positive samples are obtained by pairing different articles within the same event two by two, and negative samples are obtained by pairing articles between different events two by two with sampling and pairing articles between similar events two by two; Perform word segmentation on the two articles in each text pair to obtain multiple independent words or characters; Use a preset dictionary or lexicon to encode each word or character to obtain the character encoding vectors corresponding to the two articles in each text pair; Input the character encoding vectors corresponding to the two articles in each text pair into an embedding layer to convert them into matrix vectors, and input the matrix vectors corresponding to the two articles in the same text pair into two independent and same-layer convolutional layers for feature extraction; The features extracted by the two convolutional layers are input into the pooling layer to calculate the similarity between the corresponding features of the two articles in the same text pair and to reduce the dimensionality of the extracted features; The similarity between the corresponding features of the two articles in each text pair and the features after dimensionality reduction are input into the fully connected layer for classification and normalization to obtain the text pair model.

3. The hot event discovery method according to claim 1, characterized in that, The number of categories for rough clustering is determined based on the number of crawled articles, and all articles are clustered using a preset clustering algorithm. The rough clustering results include: Based on the number of crawled articles, determine the number of categories for rough clustering; Segment the crawled articles to obtain multiple independent words or characters; The words or characters after word segmentation processing are converted into word vectors, and according to the number of categories, a preset clustering algorithm is used to roughly cluster the word vectors corresponding to each crawled article to obtain a rough clustering result.

4. The hot event discovery method according to any one of claims 1-3, characterized in that After the step of sequentially selecting articles in the same category from the rough clustering results and pairing them up in pairs to construct text pairs, the method further includes: Get the titles of the articles in each constructed text pair in the same category; Determine whether the titles of articles in the same text pair are the same; If they are the same, the corresponding text pair is retained, otherwise the corresponding text pair is discarded.

5. A hot event discovery device, characterized in that, The hot event discovery device comprises: The collection module is used to collect articles published on designated websites using web crawler technology; The clustering module is used to determine the number of categories for rough clustering based on the number of crawled articles, and cluster all articles using a preset clustering algorithm to obtain a rough clustering result; A pairing module, used for sequentially pairing articles in the same category from the rough clustering results to construct text pairs; The model processing module is used to pre-process each text pair and then sequentially input it into the preset text pair model for processing, and output the similarity between any two articles and whether they belong to the same event; A graph construction module is used to construct an event graph by using each article as a vertex of the graph, connecting two articles of the same event as edges of the graph, and using the similarity between articles as the weight of the corresponding edge; A segmentation module is used to segment the event graph to obtain multiple sub-event graphs, wherein all vertex corresponding articles in the same sub-event graph are articles corresponding to the same hot event; The segmentation module comprises: An initialization unit, used for initializing the event graph to divide each vertex into different partitions, wherein the number of initial partitions is the same as the number of vertices; A partition division unit is used to perform a partition trial division on each vertex one by one, so as to divide each vertex into a partition where a neighbor vertex adjacent to the vertex is located, and calculate the modularity change value of the event graph corresponding to each vertex before and after the division, and record the neighbor vertex corresponding to the maximum modularity change value; if the maximum modularity change value is greater than 0, the corresponding vertex is divided into the partition where the neighbor vertex corresponding to the maximum modularity change value is located, otherwise the vertex trial division is abandoned; the processing flow of the partition trial division is repeatedly executed until the partitions corresponding to all vertices no longer change; A new graph construction unit is used to compress all vertices within the same partition into a new vertex to construct a new event graph, set the weights of the edges between vertices within the same partition as the weights of the loops of the new vertex, and set the weights of the edges between different partitions as the weights of the edges between the new vertices; repeat the process of constructing the new event graph until the modularity of the entire event graph no longer changes, where a partition corresponds to a sub-event graph.

6. The hot event discovery device according to claim 5, wherein The hot event discovery device further includes: An acquisition module for acquiring training samples for training the text pair model, where the training samples are text pairs and include positive samples and negative samples. A text pair contains two articles. The positive samples are obtained by pairing different articles within the same event pairwise, and the negative samples are obtained by pairing articles between different events pairwise with sampling and pairing articles between similar events pairwise. A word segmentation module for performing word segmentation on the two articles in each text pair to obtain a plurality of independent words or characters. An encoding module for encoding each word or character using a preset dictionary or lexicon to obtain the character encoding vectors corresponding to the two articles in each text pair. A model construction module for inputting the character encoding vectors corresponding to the two articles in each text pair into an embedding layer to be converted into matrix vectors, and inputting the matrix vectors corresponding to the two articles in the same text pair into two independent and same-layer convolutional layers for feature extraction respectively; inputting the features extracted by the two convolutional layers into a pooling layer respectively to calculate the similarity between the features corresponding to the two articles in the same text pair and perform dimensionality reduction on the extracted features; inputting the similarity between the features corresponding to the two articles in each text pair and the dimensionality-reduced features into a fully connected layer for classification and normalization processing to obtain the text pair model.

7. A hot event discovery device, characterized in that The hot event discovery device includes a memory, a processor, and a hot event discovery program stored on the memory and executable on the processor. When the hot event discovery program is executed by the processor, the steps of the hot event discovery method according to any one of claims 1-4 are implemented.

8. A computer-readable storage medium, characterized in that, A hot event discovery program is stored on the computer-readable storage medium. When the hot event discovery program is executed by the processor, the steps of the hot event discovery method according to any one of claims 1-4 are implemented.

Citation Information

Patent Citations

  • Method, system and device for discovering and tracking hot topics based on network media data stream

    CN108804432A