Hotspot event mining method and device, equipment and medium

By acquiring document keywords for frequent itemset mining and event graph construction, the problem of finding hot events in massive amounts of information is solved, achieving efficient and accurate identification and display of hot events.

CN114547340BActive Publication Date: 2025-12-30BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210179886.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-12-30
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively uncover trending events from massive amounts of information, leading to difficulties for users in obtaining information, a decline in the reputation of internet products, and user churn.

Method used

By obtaining keywords from the original documents, frequent itemset mining is performed to construct an event graph and obtain event clusters, thus determining a list of hot events.

Benefits of technology

It improved the accuracy and efficiency of identifying trending events, saved computing resources, and enhanced the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114547340B_ABST
    Figure CN114547340B_ABST
Patent Text Reader

Abstract

The present disclosure provides a hotspot event mining method and device, equipment and medium, relates to the technical field of data processing, and particularly relates to the technical field of big data and artificial intelligence. The implementation scheme is: obtaining a plurality of original documents; for each original document, obtaining at least one keyword included in the original document; based on a plurality of keywords included in each of the plurality of original documents, obtaining at least one keyword frequent item set; based on the at least one keyword frequent item set, determining a plurality of preliminary screening documents from the plurality of original documents; at least based on a plurality of keywords included in each of the plurality of preliminary screening documents, constructing an event graph; based on the event graph, obtaining at least one event cluster; and based on the at least one event cluster, determining a hotspot event list.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to the fields of big data and artificial intelligence technology, specifically to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for mining hot events. Background Technology

[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0003] With the development of information technology, people have access to more and more channels for obtaining information, and the amount of information they acquire is also increasing. People's attention to trending events is growing daily, making the ability to quickly obtain social information and news a pressing need for users. However, the vast amount of unorganized information poses a significant obstacle to users' information acquisition, making it increasingly difficult for them to effectively obtain events of interest, thus leading to a decline in the reputation of internet products and user churn. Therefore, extracting trending events from massive amounts of information has become a crucial aspect of attracting and retaining users in the era of big data.

[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0005] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for hotspot event mining.

[0006] According to one aspect of this disclosure, a method for mining hot events is provided, comprising: acquiring multiple original documents; for each original document, acquiring at least one keyword included in the original document; acquiring at least one frequent itemset of keywords included in each of the multiple original documents; determining multiple preliminary screening documents from the multiple original documents based on the at least one frequent itemset of keywords; constructing an event graph based on at least one of the keywords included in each of the multiple preliminary screening documents; acquiring at least one event cluster based on the event graph; and determining a list of hot events based on the at least one event cluster.

[0007] According to another aspect of this disclosure, a hotspot event mining apparatus is provided, comprising: a first acquisition unit configured to acquire a plurality of original documents; a second acquisition unit configured to acquire at least one keyword included in each original document; a third acquisition unit configured to acquire at least one frequent itemset of keywords based on the plurality of keywords included in each of the plurality of original documents; a first determination unit configured to determine a plurality of preliminary screening documents from the plurality of original documents based on the at least one frequent itemset of keywords; a construction unit configured to construct an event graph based on the plurality of keywords included in each of the plurality of preliminary screening documents; a fourth acquisition unit configured to acquire at least one event cluster based on the event graph; and a second determination unit configured to determine a list of hotspot events based on the at least one event cluster.

[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described hotspot event mining method.

[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the above-described hotspot event mining method.

[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program, when executed by a processor, is capable of implementing the aforementioned hotspot event mining method.

[0011] According to one or more embodiments of this disclosure, the accuracy and efficiency of hot topic discovery can be improved.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0014] Figure 1A schematic diagram of an exemplary system in which various methods described herein may be implemented, according to exemplary embodiments of the present disclosure;

[0015] Figure 2 A flowchart of a hotspot event mining method according to an exemplary embodiment of the present disclosure is shown;

[0016] Figure 3 A schematic diagram of a tree data structure in the FP-Growth frequent itemset mining algorithm according to an exemplary embodiment of the present disclosure is shown;

[0017] Figure 4 A schematic diagram illustrating the process of a hotspot event mining method according to an exemplary embodiment of the present disclosure is shown;

[0018] Figure 5 A flowchart of a hotspot event mining method according to an exemplary embodiment of the present disclosure is shown;

[0019] Figure 6 A structural block diagram of a hotspot event mining apparatus according to an exemplary embodiment of the present disclosure is shown;

[0020] Figure 7 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0021] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0022] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to define the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0023] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0024] In related technologies, one approach to hot topic mining is based on feature extraction. This involves designing semi-structured feature extraction algorithms for hot topic data, segmenting the semi-structured data, and then identifying hot topics. However, this method requires extensive analysis of hot topics and manual feature design, resulting in high labor costs and hindering large-scale application. Another approach is based on text vector clustering. This involves obtaining article vector representations and using clustering algorithms to group these vectors into event clusters, thus identifying hot topics. However, article vectors may contain biases, failing to effectively represent the semantics of the articles, leading to unsatisfactory hot topic mining results. Furthermore, vector clustering is computationally intensive, resulting in low efficiency in hot topic mining.

[0025] Based on this, this disclosure provides a method for hotspot event mining. By obtaining keywords from the original documents, frequent itemset mining is performed, and the original documents are initially screened based on this. Furthermore, an event graph is constructed based on the initially screened documents to obtain hotspot event clusters. This method can utilize keywords in documents to characterize document content features, thereby improving the accuracy of hotspot event mining. Moreover, the initial document screening based on frequent itemsets of keywords saves computational resources required for hotspot event mining and improves efficiency.

[0026] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0027] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0028] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of hotspot event mining methods.

[0029] In some embodiments, server 120 may also provide other services or software applications that may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) model.

[0030] exist Figure 1In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.

[0031] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to send raw documents for hotspot event mining. The client devices can provide an interface that allows users to interact with them. The client devices can also output information to users through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.

[0032] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0033] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0034] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0035] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0036] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.

[0037] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0038] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.

[0039] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.

[0040] Figure 1 The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.

[0041] Figure 2 A flowchart of a hotspot event mining method according to an exemplary embodiment of this disclosure is shown. Figure 2 As shown, the method includes: step S201, obtaining multiple original documents; step S202, for each original document, obtaining at least one keyword included in the original document; step S203, based on the multiple keywords included in each of the multiple original documents, obtaining at least one frequent itemset of the keyword; step S204, based on the at least one frequent itemset of the keyword, determining multiple preliminary screening documents from the multiple original documents; step S205, constructing an event graph based on at least the multiple keywords included in each of the multiple preliminary screening documents; step S206, based on the event graph, obtaining at least one event cluster; and step S207, based on the at least one event cluster, determining a list of hot events. Therefore, keywords in documents can be used to characterize document content features, thereby improving the accuracy of hot event mining. Furthermore, using frequent itemsets of the keyword for preliminary document screening saves computational resources required for hot event mining and improves efficiency.

[0042] For example, the original document may be a media article, news content, webpage content from a specific website, etc. It should be understood that the content of each original document corresponds to at least one event, and the keywords included in the original document can characterize the main content of the corresponding event.

[0043] For example, obtaining the original document set in step S201 can be done automatically using a specific computer program. For instance, it can automatically crawl the content of all web pages published by a specific website according to a preset time period. Step S201 can also be done manually, and there is no limitation on this.

[0044] For example, in step S202, obtaining at least one keyword included in each original document can be achieved using the textRank keyword extraction algorithm, or it can be achieved using other types of keyword acquisition strategies, such as the LAD algorithm, Jieba word segmentation, etc., without limitation.

[0045] According to some embodiments, the original document includes at least one content module, and step S202, for each original document, obtaining at least one keyword included in the original document, includes: for each content module in the original document, determining a target acquisition strategy based on the position of the content module; and using the target acquisition strategy to obtain at least one keyword included in the content module. Therefore, different keyword acquisition strategies can be used for content at different positions in the document, fully considering the density differences of effective information contained in the document content at different positions, and improving the accuracy of document content features represented by keywords.

[0046] For example, the content modules included in the original document may include the document's first paragraph, middle paragraph, and last paragraph. Generally, the first and last paragraphs are more likely to be summary sections of the entire document, containing a higher density of effective information. Therefore, based on the location of the content modules included in the document, different keyword retrieval strategies can be used for the first, last, and middle paragraphs, thereby enabling more rational allocation of computing resources, improving efficiency, and enhancing the accuracy of document content features represented by keywords.

[0047] According to some embodiments, the at least one content module includes a document title and a document body, and step S202, for each original document, obtaining at least one keyword included in the original document, includes: for the document title included in the original document, determining a first acquisition strategy as the target acquisition strategy, and using the first acquisition strategy to obtain at least one title keyword included therein; for the document body included in the original document, determining a second acquisition strategy different from the first acquisition strategy as the target acquisition strategy, and using the second acquisition strategy to obtain at least one body keyword included therein; and based on the at least one title keyword and the at least one body keyword, determining at least one keyword included in the original document. Therefore, different keyword acquisition strategies can be used to obtain the keywords included in the document title and document content based on the density differences of the effective information contained therein, thereby improving the accuracy of document content features represented by keywords.

[0048] For example, the first acquisition strategy can be the Jieba word segmentation method, which can fully utilize the high-level semantic representation of the document content by the document title to extract the entity words included therein, so as to represent the document content features more efficiently and accurately. For example, the second acquisition strategy can be the TextRank keyword extraction algorithm. Generally speaking, document content contains a lot of descriptive and introductory information, and its effective information density is low. Using the TextRank keyword extraction algorithm can fully filter and select the effective information contained therein, discard invalid information, thereby saving the computing resources occupied by hot event mining and improving efficiency.

[0049] For example, in step S203, the FP-Growth algorithm can be used to obtain at least one frequent itemset of keywords based on the multiple keywords included in each of the multiple original documents, but it is not limited to this. For example, the Apriori algorithm or the Eclat algorithm can also be used to obtain the frequent itemset of keywords.

[0050] According to some embodiments, the method further includes: for each keyword included in each original document, determining a corresponding weight coefficient for the keyword based on the keyword's position in the original document and / or the keyword's part-of-speech tag; and in step S203, obtaining at least one keyword frequent itemset based on the multiple keywords included in each of the multiple original documents and the corresponding weight coefficient of each keyword. Thus, it is possible to assign weight coefficients to keywords based on their position and / or part-of-speech tag, enabling the multiple keywords corresponding to each document to more accurately represent the document content features.

[0051] For example, when the original document includes a document title and document content, different weight coefficients can be assigned to keywords in the document title and keywords in the document content based on their positions in the original document. It should be understood that the effective information density contained in the document title is usually higher than that in the document content. Therefore, by assigning higher weight coefficients to keywords in the document title and performing frequent itemset mining based on this, more accurate frequent itemsets can be obtained.

[0052] For example, different weight coefficients can be assigned to each keyword included in the original document based on its part of speech. For instance, different weight coefficients can be assigned to entity noun keywords and verb keywords, thereby fully utilizing the differences in effective information density represented by keywords with different parts of speech, and thus improving the accuracy of keyword frequent itemset mining.

[0053] The exemplary embodiments of this disclosure will be further described below with reference to examples.

[0054] In one example, the plurality of original documents comprises ten original documents, and the multiple keywords included in each original document and the corresponding weight coefficients for each keyword are shown in Table 1. The data in Table 1 can be obtained by iterating through the original documents and their multiple keywords and the corresponding weight coefficients for each keyword.

[0055] Table 1:

[0056]

[0057]

[0058] When m = 1.8 and n = 1.5, frequent itemset mining can be further performed based on the keywords and their corresponding weight coefficients in Table 1, building upon the FP-Growth algorithm. For example, the support of each keyword can be calculated based on its corresponding weight coefficient, and the itemset head table shown in Table 2 can be constructed by sorting the keywords in descending order of their support.

[0059] Table 2:

[0060] Keywords Support B 11.2 A 11 C 9 D 4.8 E 4.5

[0061] In some examples, based on the item header table shown in Table 2, keyword filtering can be achieved. For example, the keyword E with the lowest support can be discarded to improve the efficiency and accuracy of keyword frequent itemset mining.

[0062] For example, the data in Table 1 can be traversed again, and the keywords included in each original document in Table 1 can be sorted according to the keyword order in the header table shown in Table 2. Based on the sorting result, the keywords included in each original document can be added to a tree data structure with an empty node as the root node, to obtain the following... Figure 3 The tree-like data structure shown is an FP-tree. Further, based on the FP-tree, the conditional pattern base for each keyword in the item header table can be mined. This involves using the node corresponding to the keyword as the FP subtree corresponding to the leaf node, thus recursively mining the frequent itemsets of the keyword. Correspondingly, the support of each node in the conditional pattern base of the keyword can be calculated based on the weight coefficient of the corresponding keyword, and filtering can be performed based on the node support as needed.

[0063] Therefore, calculating the support of a keyword based on its corresponding weight coefficient can more accurately indicate the accuracy and importance of the document content features represented by that keyword, thereby enabling more precise and efficient acquisition of frequent keyword itemsets and improving the accuracy of hot topic mining.

[0064] For example, step S204, which involves determining multiple preliminary screening documents from the plurality of original documents based on at least one keyword frequent itemset, may include: for each of the plurality of original documents, in response to the original document including at least one keyword frequent itemset, determining the original document as a preliminary screening document. This allows for preliminary document screening based on keyword frequent itemsets, constructing an event graph and mining event clusters only for the preliminary screening documents, thereby saving computational resources required for hotspot event mining and effectively improving efficiency.

[0065] According to some embodiments, the construction of the event graph in step S205 includes: for any two of the plurality of pre-screened documents, in response to the fact that the multiple keywords included in each of the two pre-screened documents satisfy a first preset condition, establishing edges of the event graph with the two pre-screened documents as vertices. Thus, the correlation between the content of two documents can be represented using the multiple keywords corresponding to each of the two documents, and an event graph can be constructed based on this, improving accuracy.

[0066] For example, in response to the fact that the multiple keywords included in each of the two initial screening documents meet the first preset condition, the establishment of the edge of the event graph with the two initial screening documents as vertices may include: inputting the multiple keywords included in each of the two initial screening documents into the relevance prediction model, and obtaining the relevance prediction result output by the relevance prediction model; in response to the relevance prediction result being higher than a preset threshold, establishing the edge of the event graph with the two initial screening documents as vertices.

[0067] For example, the relevance prediction model can be trained through the following steps: obtaining multiple keywords included in each of two sample documents; calculating the true relevance of the two sample documents based on the multiple keywords included in each of the two sample documents; inputting the multiple keywords included in each of the two sample documents into the relevance prediction model and obtaining the relevance prediction result output by the relevance prediction model; calculating a loss value based on the true relevance result and the relevance prediction result; and adjusting the parameters of the relevance prediction model based on the loss value. Therefore, it is possible to obtain the relevance of corresponding documents based on the keywords included in the documents using the relevance prediction model, and construct an event graph based on this, which is more convenient and accurate.

[0068] According to some embodiments, the step of constructing the edges of an event graph with the two initial screening documents as vertices in response to the multiple keywords included in each of the two initial screening documents satisfying a first preset condition includes: constructing the edges of an event graph with the two initial screening documents as vertices in response to the intersection of the multiple keywords included in each of the two initial screening documents satisfying a second preset condition. Thus, the intersection of the multiple keywords included in each of the two initial screening documents can be used to characterize the correlation between the contents of two documents, and an event graph can be constructed based on this, improving accuracy.

[0069] For example, the step of establishing the edges of the event graph with the two initial screening documents as vertices in response to the intersection of the multiple keywords included in each of the two initial screening documents satisfying a second preset condition may include: establishing the edges of the event graph with the two initial screening documents as vertices in response to the ratio of the number of keywords included in the intersection of the multiple keywords included in each of the two initial screening documents to the number of keywords included in each of the two initial screening documents being greater than a preset threshold.

[0070] In one example, the two initial screening documents contain 8 and 10 keywords respectively. When the preset threshold is 0.5, that is, when the ratio of the number of keywords in the intersection of the keywords in each of the two initial screening documents to 8 and 10 is greater than 0.5, then an event graph edge can be constructed using any two initial screening documents as vertices. This allows for a simpler and more efficient determination of the conditions for constructing event graph edges, improving the efficiency and accuracy of hotspot event mining.

[0071] According to some embodiments, when each keyword included in each initial screening document has a corresponding weight coefficient, the step of establishing the edges of an event graph with the two initial screening documents as vertices in response to the intersection of multiple keywords included in any two initial screening documents satisfying a second preset condition includes: calculating the sum of the corresponding weight coefficients of each keyword included in the intersection of multiple keywords included in any two initial screening documents; and establishing the edges of an event graph with the two initial screening documents as vertices in response to the sum of the corresponding weight coefficients of each keyword included in the intersection of multiple keywords included in any two initial screening documents being greater than a preset threshold. Therefore, the keyword weight coefficients can be fully considered to represent the accuracy and importance of the document content features that the keyword can characterize, thereby more accurately representing the correlation between the contents of any two initial screening documents.

[0072] In one example, the two initial screening documents include document A and document B, and both document A and document B include two content modules: document title and document content. The weight coefficient of each keyword in the original document is determined based on the keyword's position and part of speech within the original document. Different weight coefficients can be assigned to keywords in the document based on their position and part of speech. For example, the weight coefficient of entity noun keywords in the document title is x, the weight coefficient of verb keywords in the document title is y, and the weight coefficient of keywords in the document content is z. Furthermore, the weight coefficient of keywords in the document title can be set to m, and the weight coefficient of keywords in the document content can be set to n, with the sum of m and n limited to 1, to balance the importance of keywords in the document title and document content.

[0073] The keywords included in each original document in Document A and Document B are shown in Table 3.

[0074] Table 3:

[0075]

[0076] It can be seen that the intersection of keywords in the titles of documents A and B is {b, c, d, e}, and the sum of the weight coefficients for each keyword is (2x + 2y). The intersection of keywords in the content of documents A and B is {h, i}, and the sum of the weight coefficients for each keyword is (2z). Further, we can set the weight coefficient of keywords in the document titles to m and the weight coefficient of keywords in the document content to n, and limit the sum of m and n to 1 to balance the importance of keywords in the document titles and content. The intersection of keywords in the title of document A and the content of document B is {a, c}, and the sum of the weight coefficients for each keyword is (2mx + 2nz). The intersection of keywords in the content of document A and the title of document B is {b, f}, and the sum of the weight coefficients for each keyword is (2mx + 2nz). Furthermore, by calculating the sum of the weight coefficients of each keyword included in the intersection of the above types, it can be determined whether to establish the edges of the event graph with the two initially screened documents as vertices.

[0077] According to some embodiments, constructing the event graph in step S205 includes: for any two of the plurality of initial screening documents, in response to the fact that the two initial screening documents contain the same frequent itemset of keywords, establishing edges of the event graph with the two initial screening documents as vertices. Thus, the frequent itemsets of keywords contained in each of the two initial screening documents can be used to characterize the correlation between the contents of the two documents, and an event graph can be constructed based on this, improving accuracy.

[0078] For example, the determination of whether to establish an edge graph with any two initial screening documents as vertices can also be based on a combination of conditions from various embodiments described above. For instance, in response to the intersection of multiple keywords included in each of the two initial screening documents satisfying a second preset condition, and the two initial screening documents containing the same frequent itemset of a keyword, when both conditions are met simultaneously, an edge graph with the two initial screening documents as vertices can be established, thereby enabling a more accurate representation of the correlation between the contents of any two initial screening documents.

[0079] According to some embodiments, the initial screening documents include publication time information, and the construction of the event graph in step S205 further includes: in response to the publication time of the initial screening documents included in the event graph satisfying a third preset condition, deleting the corresponding vertices and edges of the initial screening documents from the event graph. This enables dynamic adjustment of the event graph and more rational allocation of computing resources.

[0080] For example, deleting the corresponding vertex and edge of the preliminary screening document from the event graph in response to the publication time of the document included in the event graph meeting a third preset condition may include: deleting the corresponding vertex and edge of the preliminary screening document from the event graph in response to the difference between the publication time of the preliminary screening document included in the event graph and the current time being greater than a preset threshold. This enables the extraction of expired vertices and edges from the event graph, achieving automatic updates and preventing the event graph from continuously growing in size, thus allowing for the rational allocation of computing resources.

[0081] For example, in response to the publication time of the preliminary screening document included in the event graph falling within a preset time period, the corresponding vertex and edge of the preliminary screening document can be deleted from the event graph, thereby enabling manual intervention in the corresponding time period of the content in the event graph and dynamic adjustment of the event graph according to actual needs.

[0082] According to some embodiments, in step S206, obtaining at least one event cluster based on the event graph includes: obtaining at least one event cluster included in the event graph based on a community detection algorithm. This allows for efficient and accurate acquisition of event clusters from the event graph.

[0083] For example, the event clusters can be obtained using the Louvain community discovery algorithm based on modularity, but it is not limited to this. For example, it can also be obtained using the LPA algorithm, HANP algorithm, SLPA algorithm, etc., without limitation.

[0084] For example, event clusters in an event graph can also be obtained based on other types of algorithms, such as using a connected subgraph algorithm.

[0085] According to some embodiments, step S207, determining the hot event list based on the at least one event cluster, includes: for each event cluster, obtaining the initial screening documents corresponding to at least one vertex included in that event cluster; and determining the hot event list based on the at least one initial screening document corresponding to each event cluster. Therefore, it is possible to determine hot events based on the initial screening documents indexed by the event clusters and the information in the documents, thereby improving the accuracy of hot event mining.

[0086] For example, the initial screening can be based on the event cluster index of the corresponding documents, and the hot events can be further determined based on the original source information of the documents. For instance, when the initial screening document is web page content, the hot events can be further determined based on the credibility of the original source website of the document, thereby improving the accuracy of hot event mining.

[0087] Furthermore, according to some embodiments, the initial screening of documents includes document popularity information, and wherein a list of hot events is determined based on the document popularity information of at least one initially screened document corresponding to each of the at least one event cluster. This allows for the identification of hot events based on document popularity information, improving the accuracy of hot event discovery.

[0088] For example, when the at least one event cluster includes multiple event clusters, the initial screening documents corresponding to each vertex included in each event cluster can be indexed and their document popularity information can be obtained. The document popularity information may include, for example, the number of reads and distributions of the document. Based on the document popularity information of the initial screening documents corresponding to each vertex included in each event cluster, the events corresponding to the initial screening documents with a document popularity higher than a certain threshold can be identified as hot events, thereby obtaining a more accurate list of hot events.

[0089] According to some embodiments, when the initial screening documents include document popularity information, the method further includes: sorting the at least one initial screening documents based on the document popularity information of at least one initial screening document corresponding to each event cluster; and displaying the hot event list based on the sorting result of the at least one initial screening document. Therefore, hot events can be displayed based on the sorting result of the document popularity information of the initial screening documents corresponding to each event cluster, thereby providing users with an accurate list of hot events and fully meeting user needs.

[0090] According to some embodiments, when the initial screening documents include publication time information, the method further includes: sorting the at least one initial screening documents based on the publication time information of at least one initial screening document corresponding to each event cluster; and displaying the list of hot events based on the sorting result of the at least one initial screening document. Therefore, hot events can be displayed based on the sorting result of the publication time information of the initial screening documents corresponding to each event cluster, thereby providing users with an accurate timeline of hot events, meeting user needs, and improving user experience.

[0091] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0092] Figure 4 A schematic diagram of a hotspot event mining method according to an exemplary embodiment of the present disclosure is shown. Figure 5 A flowchart of a hotspot event mining method according to an exemplary embodiment of the present disclosure is shown.

[0093] like Figure 5As shown, the hot topic event mining method includes:

[0094] Step S501: Obtain multiple original documents, each of the multiple original documents including at least one content module;

[0095] Step S502: For each content module in each original document, determine the target acquisition strategy based on the location of the content module;

[0096] Step S503: Using the target acquisition strategy, acquire at least one keyword included in the content module;

[0097] Step S504: For each keyword included in each original document, determine the corresponding weight coefficient of the keyword based on the position of the keyword in the original document and / or the part of speech of the keyword;

[0098] Step S505: Based on the multiple keywords included in each of the multiple original documents and the corresponding weight coefficient of each keyword, obtain at least one keyword frequent itemset;

[0099] Step S506: Based on at least one keyword frequent itemset, determine multiple initial screening documents from the multiple original documents;

[0100] Step S507: Construct an event graph based on at least the multiple keywords included in each of the multiple initially screened documents;

[0101] Step S508: Based on the community detection algorithm, obtain at least one event cluster included in the event graph;

[0102] Step S509: For each event cluster in the at least one event cluster, obtain the initial screening document corresponding to at least one vertex included in the event cluster;

[0103] Step S510: Based on the document popularity information of at least one pre-screened document corresponding to each event cluster in the at least one event cluster, determine the hot event list;

[0104] Step S511: Sort the at least one preliminary screening document based on the document popularity information of at least one preliminary screening document corresponding to each event cluster in the at least one event cluster;

[0105] Step S512: Based on the sorting results of the at least one initial screening document, display the list of hot events.

[0106] In this example, the original document includes a document title and document content. This allows for full utilization of the density differences in the effective information contained in the document title and document content, enabling the use of different keyword acquisition strategies to obtain the keywords contained in each, rationally allocating computing resources, improving the efficiency of hotspot mining, and also improving the accuracy of using keywords to represent document content features.

[0107] According to another aspect of this disclosure, a hotspot event mining apparatus is provided. Figure 6 A structural block diagram of a hotspot event mining apparatus 600 according to an exemplary embodiment of the present disclosure is shown. Figure 6 As shown, the hotspot event mining device 600 includes: a first acquisition unit 601 configured to acquire multiple original documents; a second acquisition unit 602 configured to acquire at least one keyword included in each original document; a third acquisition unit 603 configured to acquire at least one frequent itemset of keywords based on the multiple keywords included in each of the multiple original documents; a first determination unit 604 configured to determine multiple preliminary screening documents from the multiple original documents based on the at least one frequent itemset of keywords; a construction unit 605 configured to construct an event graph based on the multiple keywords included in each of the multiple preliminary screening documents; a fourth acquisition unit 606 configured to acquire at least one event cluster based on the event graph; and a second determination unit 607 configured to determine a list of hotspot events based on the at least one event cluster. The operations of units 601-607 of the hotspot event mining device 600 are similar to the operations of steps S201-S207 described above, and will not be repeated here.

[0108] According to another aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the hotspot event mining method described above.

[0109] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to cause the computer to perform the above-described hotspot event mining method.

[0110] According to another aspect of this disclosure, a computer program product is also provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the above-described hotspot event mining method.

[0111] refer to Figure 7The present invention describes a structural block diagram of an electronic device 700 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0112] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded from storage unit 708 into random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0113] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, output unit 707, storage unit 708, and communication unit 709. Input unit 706 can be any type of device capable of inputting information to device 700. Input unit 706 can receive input numerical or character information and generate key signal inputs related to user settings and / or function control of the electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 707 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 708 may include, but is not limited to, a hard disk and an optical disk. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0114] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the hotspot event mining method. For example, in some embodiments, the hotspot event mining method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the hotspot event mining method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the hotspot event mining method by any other suitable means (e.g., by means of firmware).

[0115] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0116] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0117] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0119] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0120] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0121] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0122] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A hotspot event mining method, comprising: obtaining a plurality of original documents; for each original document, obtaining at least one keyword included in the original document, comprising: obtaining at least one title keyword included in a document title of the original document by using a first obtaining strategy, the first obtaining strategy comprising a word segmentation strategy; obtaining at least one body keyword included in a document body of the original document by using a second obtaining strategy different from the first obtaining strategy, the second obtaining strategy comprising a keyword extraction strategy; and determining at least one keyword included in the original document based on the at least one title keyword and the at least one body keyword; for each keyword included in each original document, determining a corresponding weight coefficient of the keyword based on a position of the keyword in the original document and a part of speech of the keyword; obtaining at least one keyword frequent item set based on a plurality of keywords included in each of the plurality of original documents and the corresponding weight coefficient of each keyword; determining a plurality of pre-screening documents from the plurality of original documents based on the at least one keyword frequent item set; constructing an event graph based on at least a plurality of keywords included in each of the plurality of pre-screening documents, comprising: for any two pre-screening documents in the plurality of pre-screening documents, establishing an edge of the event graph with the two pre-screening documents as vertices in response to the plurality of keywords included in the two pre-screening documents satisfying the following conditions: a sum of the corresponding weight coefficients of each keyword included in an intersection of the plurality of keywords included in the two pre-screening documents is greater than a preset threshold or the two pre-screening documents contain a same keyword frequent item set; and a ratio of a number of keywords included in the intersection of the plurality of keywords included in the two pre-screening documents to a number of keywords included in each of the two pre-screening documents is greater than a preset threshold; obtaining at least one event cluster based on the event graph; and determining a hotspot event list based on the at least one event cluster.

2. The method of claim 1, wherein, the pre-screening document includes publication time information, and wherein the constructing the event graph further comprises: in response to a publication time of a pre-screening document included in the event graph satisfying a third preset condition, deleting a corresponding vertex and edge of the pre-screening document from the event graph.

3. The method of claim 1 or 2, wherein, the obtaining at least one event cluster based on the event graph comprises: obtaining at least one event cluster included in the event graph based on a community discovery algorithm.

4. The method of claim 1 or 2, wherein, the determining a hotspot event list based on the at least one event cluster comprises: for each event cluster in the at least one event cluster, obtaining at least one pre-screening document corresponding to at least one vertex included in the event cluster; and determining a hotspot event list based on the at least one pre-screening document corresponding to each event cluster in the at least one event cluster.

5. The method of claim 4, wherein, the pre-screening document includes document popularity information, and wherein the determining a hotspot event list based on the at least one event cluster comprises:

6. The method of claim 4, when the pre-screening document includes document popularity information, the method further comprising: sort the at least one pre-screening document based on the document hotness information of the at least one pre-screening document corresponding to each of the at least one event cluster; and display the hot event list based on the sorting result of the at least one pre-screening document.

7. The method of claim 4, when the pre-screening document comprises publication time information, the method further comprises: sort the at least one pre-screening document based on the publication time information of the at least one pre-screening document corresponding to each of the at least one event cluster; and display the hot event list based on the sorting result of the at least one pre-screening document.

8. A hot event mining apparatus, comprising: a first obtaining unit configured to obtain a plurality of original documents; a second obtaining unit configured to, for each original document, obtain at least one keyword included in the original document, the second obtaining unit is configured to: obtain at least one title keyword in a document title included in the original document by using a first obtaining strategy, the first obtaining strategy comprising a word segmentation strategy; obtain at least one body keyword in a document body included in the original document by using a second obtaining strategy different from the first obtaining strategy, the second obtaining strategy comprising a keyword extraction strategy; and determine at least one keyword included in the original document based on the at least one title keyword and the at least one body keyword; the hot event mining apparatus is further configured to, for each keyword included in each original document, determine a weight coefficient of the keyword based on a position of the keyword in the original document and a part of speech of the keyword; a third obtaining unit configured to obtain at least one keyword frequent item set based on a plurality of keywords included in each of the plurality of original documents and the weight coefficient of each keyword; a first determining unit configured to determine a plurality of pre-screening documents from the plurality of original documents based on at least one keyword frequent item set; a constructing unit configured to construct an event graph based on a plurality of keywords included in each of the plurality of pre-screening documents, comprising: for any two pre-screening documents in the plurality of pre-screening documents, in response to the plurality of keywords included in the two pre-screening documents satisfying the following conditions, establishing an edge of the event graph with the two pre-screening documents as vertices: a sum of the weight coefficients of each keyword included in the intersection of the plurality of keywords included in the two pre-screening documents is greater than a preset threshold or the two pre-screening documents contain the same keyword frequent item set; and a ratio of the number of keywords included in the intersection of the plurality of keywords included in the two pre-screening documents to the number of keywords included in each of the two pre-screening documents is greater than a preset threshold; a fourth obtaining unit configured to obtain at least one event cluster based on the event graph; and a second determining unit configured to determine a hot event list based on the at least one event cluster.

9. An electronic device, comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein ​ The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.

10. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are for causing a computer to perform the method of any one of claims 1-7.

11. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Hotspot mining method, server and computer readable storage medium

    CN110232126A

  • Event development venation diagram generation method

    CN111382276A