A method, device, and storage medium for automatically constructing a main event library
Automatically construct the subject event database through DBSCAN algorithm and LAC technology, solving the problem of time-consuming and laborious search of subject events and data omissions manually, and achieving efficient event information acquisition and easy backtracking functions.
Patent Information
- Application Number
- CN202211662243.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-12-23
AI Technical Summary
In the prior art, when researchers look up subject events in public opinion analysis, it is time-consuming and laborious to focus on subject keywords through manual search and data omissions, especially in the case of large time spans.
The DBSCAN algorithm is used to cluster the data titles, and entities are extracted through LAC and the DBSCAN algorithm is used to merge the subjects, build the main event library, obtain data in segments and sort it according to the time span.
It realizes automatic construction of the main event database under the large time span, reduces data omissions, improves data processing efficiency, facilitates and quickly obtains event-related information, and has easy traceback function.
Smart Images

Figure CN115858718B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly to an automatic method, device and storage medium for a subject event library. Background Art
[0002] In public opinion analysis, event analysis is a very important research point. Usually, researchers will search for keywords of the concerned subject to find relevant information. However, in the process of manually sorting data, it is inevitable to have the problem of data omission, and this approach is time-consuming and laborious, especially in the case of a large time span, the efficiency is relatively low. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide an automatic method, device and storage medium for a subject event library, so as to solve the problem in the prior art that when researchers search for relevant information by manually searching for keywords of the concerned subject, it is time-consuming and laborious, and it is easy to have the problem of data omission, especially for the case of a large time span.
[0004] According to the first aspect of the embodiments of the present invention, an automatic construction method for a subject event library is provided, including:
[0005] Set a time interval according to the time span, and obtain data in segments from the message queue at the preset time interval. The data includes a data title, a data resource locator, and a release time;
[0006] According to the data titles obtained within each time interval, perform clustering through the DBSCAN algorithm, select the top X clustering results in the clustering results, and regard each selected clustering result as a clustering result cluster;
[0007] Extract the entities of each title in each clustering result cluster through LAC. If the number of entities in the title is greater than 1, then select the first entity that appears as the entity of the title. After obtaining the entities of all titles in the clustering result cluster, regard the entity with the largest number of occurrences as the subject of the clustering result cluster;
[0008] Merge the subjects within each time interval through the DBSCAN algorithm respectively, merge the titles within the same subject, and then merge all the subjects within the time span; sort the data resource locators according to the release time for all the titles within the same subject to construct a subject event library.
[0009] Preferably,
[0010] The clustering of different titles through the DBSCAN algorithm respectively includes:
[0011] Split all the obtained data titles into several segments, cluster each segment of data titles using the DBSCAN algorithm, and select the top X clustering results from the clustering results of this segment of data titles;
[0012] Combine the X clustering results of each segment of data titles, and then perform secondary clustering using the DBSCAN algorithm, and select the top X clustering results to form a clustering result cluster.
[0013] Preferably,
[0014] The clustering using the DBSCAN algorithm includes:
[0015] Select a group of titles from the data titles, vectorize this group of titles, and calculate the similarity matrix of two vectors;
[0016] Preset the neighborhood radius E for defining density and the threshold M for defining core points in the DBSCAN algorithm;
[0017] Substitute the similarity matrix into the DBSCAN algorithm, and calculate the respective labels of the two titles according to the preset neighborhood radius E for defining density and the threshold M for defining core points;
[0018] Repeat the above steps until all titles are labeled with corresponding labels;
[0019] Combine the title sets with the same label into one class. After all titles are classified, rank them according to the number of titles in each class, and select the top X classes. Each class is a clustering result cluster.
[0020] Preferably,
[0021] The merging of all entities using the DBSCAN algorithm includes:
[0022] Select a group of entities, and obtain the similarity between two entities in this group;
[0023] Preset the neighborhood radius E for defining density and the threshold M for defining core points in the DBSCAN algorithm;
[0024] Substitute the similarity between the two entities into the DBSCAN algorithm, and calculate the respective labels of the two entities according to the preset neighborhood radius E for defining density and the threshold M for defining core points;
[0025] Repeat the above steps until all entities are labeled with corresponding labels;
[0026] Merge the entities with the same label.
[0027] Preferably,
[0028] Selecting a group of entities and obtaining the similarity between two entities in the group includes:
[0029] Searching for relevant words of one entity through a trained word vector model, and taking the top Y relevant words with the highest occurrence frequency;
[0030] If the other entity to be compared appears in the top Y relevant words with the highest occurrence frequency, the similarity value of the two groups of entities is 1. If the other entity to be compared does not appear in the top Y relevant words with the highest occurrence frequency;
[0031] Then calculate the similarity between the two entities through the Jaccard similarity calculation formula.
[0032] Preferably,
[0033] The Jaccard similarity calculation formula is shown as follows:
[0034]
[0035] In the formula, A and B respectively represent two entities.
[0036] According to the second aspect of the embodiments of the present invention, a device for automatically constructing an entity event library is provided. The device includes:
[0037] A data acquisition module: configured to set a time interval according to a time span, and segmentally acquire data from a message queue at the preset time interval. The data includes a data title, a data resource locator, and a release time;
[0038] A title clustering module: configured to cluster according to the data titles acquired within each time interval through the DBSCAN algorithm, select the top X clustering results in the clustering results, and use each selected clustering result as a clustering result cluster;
[0039] An entity selection module: configured to extract entities of each title in each clustering result cluster through LAC. If the number of entities of the title is greater than 1, select the first entity that appears as the entity of the title. After obtaining the entities of all titles in the clustering result cluster, use the entity with the most occurrences as the entity of the clustering result cluster;
[0040] An entity merging module: configured to merge the entities within each time interval through the DBSCAN algorithm, merge the titles within the same entity, and then merge all the entities within the time span; sort the data resource locators according to the release time for all titles within the same entity to construct an entity event library.
[0041] According to a third aspect of an embodiment of the present invention, a storage medium is provided. The storage medium stores a computer program, and when the computer program is executed by a main controller, each step in the above-mentioned method is implemented.
[0042] The technical solution provided by the embodiment of the present invention may include the following beneficial effects:
[0043] In this application, time intervals are set according to the time span, data is obtained segment by segment from the message queue at the preset time intervals, clustering results clusters are obtained through the DBSCAN algorithm according to the data titles, the main body of each clustering results cluster is obtained by means of LAC main body extraction, and finally the main body is clustered and merged again through the DBSCAN algorithm, all the titles of the same main body are merged, and then the data resource locators of the titles of the same main body are sorted according to the release time to construct a main body event library. Through the above method, automatic data search and classification of events are realized. Compared with the existing manual search method, the possibility of data omission is reduced, and data is obtained segment by segment according to the time span set by the time intervals. In the case of a relatively large time span, event backtracking can also be realized.
[0044] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. Brief Description of the Drawings
[0045] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0046] Figure 1 is a flowchart showing a method for automatically constructing a main body event library according to an exemplary embodiment;
[0047] Figure 2 is a system diagram showing an apparatus for automatically constructing a main body event library according to another exemplary embodiment;
[0048] In the drawings: 1 - data acquisition module, 2 - title clustering module, 3 - main body selection module, 4 - main body merging module. Detailed Embodiments
[0049] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are only examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0050] Embodiment 1
[0051] Figure 1 is a schematic flowchart of a method for automatically constructing a main event library shown according to an exemplary embodiment. As Figure 1 shown, the method includes:
[0052] S1. Set a time interval according to the time span, and segmentally obtain data from the message queue at the preset time interval. The data includes a data title, a data resource locator, and a release time;
[0053] S2. Cluster according to the data titles obtained within each time interval by using the DBSCAN algorithm, select the top X clustering results in the clustering results, and regard each selected clustering result as a clustering result cluster;
[0054] S3. Extract the entities of each title in each clustering result cluster through LAC. If the number of entities of the title is greater than 1, select the first entity that appears as the entity of the title. After obtaining the entities of all titles in the clustering result cluster, regard the entity with the largest number of occurrences as the main body of the clustering result cluster;
[0055] S4. Merge the main bodies within each time interval respectively through the DBSCAN algorithm, merge the titles within the same main body, and then merge all the main bodies within the time span; sort the data resource locators according to the release time for all the titles within the same main body to construct a main event library;
[0056] It is understandable that the present application sets time intervals according to the time span. Generally, the time interval is 1 day. If the time span is 3 days, then the data for the first day, the second day, and the third day need to be obtained respectively. The obtained data includes the data title, the data resource locator, and the release time. Among them, the data resource locator, also known as the Uniform Resource Locator (URL), is a resource naming or location format used to specify or address resources. Then, the DBSCAN algorithm is used to cluster the data titles of each day, and the top X clustering results in the clustering results are selected. Each selected clustering result is used as a clustering result cluster. Then, the LAC is used to extract the entities of each title in each clustering result cluster. The full name of LAC is Lexical Analysis of Chinese, which is a lexical analysis tool developed by Baidu's NLP (Natural Language Processing Department) and can realize functions such as Chinese word segmentation, part-of-speech tagging, and proper name recognition. Here, the proper name recognition function is applied to specifically recognize the organizational structure name, which is both the required entity and the final subject. If the number of entities in the title is greater than 1, the first entity that appears is selected as the entity of the title. After obtaining the entities of all titles in the clustering result cluster, the entity that appears the most times is used as the subject of the clustering result cluster. Finally, the DBSCAN algorithm is used to merge the subjects of each day respectively, merge the titles within the same subject of each day, and then merge all the subjects within three days; sort the data resource locators of all titles within the same subject according to the release time to construct the subject event library. In the case of a large time span, the present application adopts a method of segmented processing and then fusion, which changes the disadvantage of time-consuming processing on the overall data in the past and greatly improves the data processing efficiency. Through the subject event library automatically constructed by the present application, it is convenient for personnel to quickly obtain information related to the subject event. When people are concerned about the development context of an event, they can conveniently extract relevant subject information from the library without being affected by the time span and has the function of easy backtracking; compared with the existing method of manually searching for subject keywords, it reduces the possibility of data omission.
[0057] Preferably,
[0058] The clustering of different titles through the DBSCAN algorithm respectively includes:
[0059] Divide all the obtained data titles into several segments, cluster each segment of data titles through the DBSCAN algorithm, and select the top X clustering results in the clustering results of this segment of data titles;
[0060] Collect the X clustering results of each segment of data titles, and then perform secondary clustering through the DBSCAN algorithm, and select the top X clustering results to form a clustering result cluster;
[0061] It is understandable that considering the problem of the long clustering time of DBSCAN on large-scale data, the data of each day is segmented into several segments, each segment is clustered, the results of the first X before clustering of each segment are selected, and then the results of the first X of each segment are gathered together, and then secondary clustering is performed, and then the first X results before clustering are selected to form a clustering result cluster.
[0062] Preferably,
[0063] The clustering by the DBSCAN algorithm includes:
[0064] Select a group of titles in the data titles, vectorize the group of titles, and calculate the similarity matrix of the two vectors;
[0065] Preset the neighborhood radius E when defining density and the threshold M when defining core points in the DBSCAN algorithm;
[0066] Substitute the similarity matrix into the DBSCAN algorithm, and calculate the respective labels of the two titles according to the preset neighborhood radius E when defining density and the threshold M when defining core points;
[0067] Repeat the above steps until all titles are labeled with corresponding labels;
[0068] Aggregate the title sets with the same label into one class. After all titles are classified, rank them according to the number of titles in each class, and select the top X classes. Each class is a clustering result cluster;
[0069] It can be understood that the DBSCAN (Density—Based Spatial Clustering of Application with Noise) algorithm is a typical density-based clustering method. It defines a cluster as the largest set of density-connected points, can divide regions with sufficient density into clusters, and can discover clusters of arbitrary shapes in spatial datasets with noise. There are two important parameters in the DBSCAN algorithm: Eps and MinPtS. Eps is the neighborhood radius when defining density, and MinPts is the threshold when defining core points. A core point is a point that contains more than MinPts number of points within the radius Eps. The selection of Eps is related to the text distance calculation method selected. In this application, MinPts and Eps are preset in advance, and then a group of titles is selected from the daily data titles. The group of titles is vectorized, and the similarity matrix of the two vectors is calculated. The process of title vectorization includes: in the process of vectorizing the text, TF-IDF is selected to calculate the word weights, TF-IDF = TF * IDF; the term frequency (TF) refers to the number of times a given word appears in the document. This number is usually normalized (generally, the term frequency is divided by the total number of words in the article) to prevent it from biasing towards long documents (the same word may have a higher term frequency in a long document than in a short document, regardless of the importance of the word); The inverse document frequency (IDF). The main idea of IDF is that if the number of documents containing the term t is less and the IDF is larger, it indicates that the term has good category discrimination ability. The IDF of a specific word can be obtained by dividing the total number of documents by the number of documents containing the word, and then taking the logarithm of the obtained quotient, The reason for adding 1 to the denominator is to avoid the denominator being 0. After the title is vectorized, the similarity between the two titles can be calculated. The similarity algorithm of vectors is a quite mature existing technology and will not be elaborated here. Substituting the similarity into the DBSCAN algorithm of preset M and E, the labels of the two titles can be obtained. Repeat the above steps until all titles in the data of this day are labeled with corresponding labels. Group the titles with the same label into a class. After all titles are classified, rank them according to the number of titles in each class, and select the top X classes. Each class is a clustering result cluster.
[0070] Preferably,
[0071] The merging of all entities through the DBSCAN algorithm includes:
[0072] Select a group of entities and obtain the similarity between two entities in the group;
[0073] Preset the neighborhood radius E for defining density and the threshold M for defining core points in the DBSCAN algorithm;
[0074] Substitute the similarity between the two entities into the DBSCAN algorithm, and calculate the respective labels of the two entities according to the preset neighborhood radius E for defining density and the threshold M for defining core points;
[0075] Repeat the above steps until all entities are labeled accordingly;
[0076] Merge the entities with the same label;
[0077] It can be understood that the DBSCAN algorithm has been described above and will not be elaborated here. After all entities are labeled accordingly, merge the entities with the same label, that is, collect all the titles within the entities with the same label.
[0078] Preferably,
[0079] The selection of a group of entities and obtaining the similarity between two entities in the group includes:
[0080] Search for the related words of one of the entities through the trained word vector model, and take the top Y related words with the highest occurrence frequency;
[0081] If the other entity to be compared appears in the top Y related words with the highest occurrence frequency, the similarity value between the two groups of entities is 1. If the other entity to be compared does not appear in the top Y related words with the highest occurrence frequency;
[0082] Then calculate the similarity between the two entities through the Jaccard similarity calculation formula;
[0083] It can be understood that for the calculation of the similarity between entities, in this application, the related words of one of the entities are searched through the trained word vector model, and the top Y related words with the highest occurrence frequency are taken. If the other entity to be compared appears in the top Y related words with the highest occurrence frequency, the similarity value between the two groups of entities is 1. For example, when searching for the related words of one entity and taking the top 10 words, if the other entity to be compared is among them, it is considered that the two are the same entity. For example, when comparing "China Minsheng Bank" and "Minsheng Bank", when searching for the related words of "Minsheng Bank", "China Minsheng Bank" appears among the top 10 words, then it is considered that the two are the same entity and the similarity value is 1. If it does not appear, the Jaccard similarity calculation is used.
[0084] Preferably,
[0085] The Jaccard similarity calculation formula is shown as follows:
[0086]
[0087] In the formula, A and B respectively represent two entities;
[0088] It can be understood that the Jaccard similarity calculation regards the entities as sets of words. The closer the entities are, the more common elements there are. In this application, the element n of entities A and B is the respective included titles. For the convenience of calculation, the maximum value of n is taken as 3, and the similarity values when n = 1, n = 2, and n = 3 are calculated respectively, and the minimum value of the three is taken as the final similarity value.
[0089] Embodiment 2
[0090] Figure 2 It is a system schematic diagram of an entity event library automatic construction device shown according to another exemplary embodiment, including:
[0091] Data acquisition module 1: used to set a time interval according to a time span, and segmentally acquire data from the message queue at the preset time interval. The data includes a data title, a data resource locator, and a release time;
[0092] Title clustering module 2: used to cluster according to the data titles acquired within each time interval through the DBSCAN algorithm, select the top X clustering results in the clustering results, and regard each selected clustering result as a clustering result cluster;
[0093] Entity selection module 3: used to extract the entities of each title in each clustering result cluster through LAC. If the number of entities of the title is greater than 1, then select the first entity that appears as the entity of the title. After obtaining the entities of all titles in the clustering result cluster, take the entity with the largest number of occurrences as the entity of the clustering result cluster;
[0094] Entity merging module 4: used to merge the entities within each time interval through the DBSCAN algorithm respectively, merge the titles within the same entity, and then merge all entities within the time span; sort the data resource locators of all titles within the same entity according to the release time to construct an entity event library;
[0095] It can be understood that in this application, the data acquisition module 1 sets the time interval according to the time span, and segments and acquires data from the message queue at the preset time interval. The data includes the data title, the data resource locator, and the release time. The title clustering module 2 clusters the data titles obtained within each time interval through the DBSCAN algorithm, selects the top X clustering results in the clustering results, and regards each selected clustering result as a clustering result cluster. The main body selection module 3 extracts the entities of each title in each clustering result cluster through LAC. If the number of entities of the title is greater than 1, the first entity that appears is selected as the entity of the title. After obtaining the entities of all titles in the clustering result cluster, the entity with the largest number of occurrences is used as the main body of the clustering result cluster. The main body merging module 4 merges the main bodies within each time interval through the DBSCAN algorithm, merges the titles within the same main body, and then merges all the main bodies within the time span. Sort the data resource locators of all titles within the same main body according to the release time to construct the main body event library. In this application, in the case of a large time span, a method of segmented processing and then fusion is adopted, which changes the drawback of time-consuming processing on the overall data in the past and greatly improves the data processing efficiency. Through the main body event library automatically constructed by this application, it is convenient for personnel to quickly obtain information related to the main body event. When people are concerned about the development context of an event, they can conveniently extract relevant main body information from the library without being affected by the time span and have the function of easy backtracking. Compared with the existing method of manually searching for main body keywords, the possibility of data omission is reduced.
[0096] Embodiment 3:
[0097] This embodiment provides a storage medium that stores a computer program. When the computer program is executed by the main controller, each step in the above method is implemented;
[0098] It can be understood that the above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.
[0099] It can be understood that the same or similar parts in the above embodiments can be referred to each other. For the content not detailed in some embodiments, reference can be made to the same or similar content in other embodiments.
[0100] It should be noted that in the description of the present invention, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "plurality" refers to at least two.
[0101] Any process or method description depicted in the flowchart or described otherwise herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations where functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner according to the relevant functions or in the reverse order, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0102] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0103] Those of ordinary skill in the art of the present technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the said program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0104] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0105] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like.
[0106] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0107] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for automatically constructing a main event library, characterized in that The method includes: Setting a time interval according to a time span, and obtaining data in segments from a message queue at the preset time interval. The data includes a data title, a data resource locator, and a publishing time; According to the data titles obtained within each time interval, clustering is performed through the DBSCAN algorithm, and the top X clustering results in the clustering results are selected. Each selected clustering result is used as a clustering result cluster; Clustering different titles respectively through the DBSCAN algorithm includes: Dividing all the obtained data titles into several segments, performing clustering on each segment of data titles through the DBSCAN algorithm, and selecting the top X clustering results in the clustering results of this segment of data titles; Combining the X clustering results of each segment of data titles, and then performing secondary clustering through the DBSCAN algorithm, and selecting the top X clustering results to form a clustering result cluster; Extracting entities of each title in each clustering result cluster through LAC. If the number of entities of a title is greater than 1, then select the first entity that appears as the entity of this title. After obtaining the entities of all titles in this clustering result cluster, use the entity with the largest number of occurrences as the main body of this clustering result cluster; Merging the main bodies within each time interval respectively through the DBSCAN algorithm, merging the titles within the same main body, and then merging all the main bodies within the time span; sorting the data resource locators according to the publishing time for all titles within the same main body to construct a main body event library; The merging the main bodies within each time interval respectively through the DBSCAN algorithm includes: Selecting a group of main bodies and obtaining the similarity between two main bodies in this group; Presetting the neighborhood radius E for defining density and the threshold M for defining a core point in the DBSCAN algorithm; Substituting the similarity between the two main bodies into the DBSCAN algorithm, and calculating the respective labels of the two main bodies according to the preset neighborhood radius E for defining density and the threshold M for defining a core point; Repeating the above steps until all the main bodies are labeled with corresponding labels; Merging the main bodies with the same label.
2. The method according to claim 1, wherein The clustering through the DBSCAN algorithm includes: Selecting a group of titles in the data titles, vectorizing this group of titles, and calculating the similarity matrix of two vectors; Presetting the neighborhood radius E for defining density and the threshold M for defining a core point in the DBSCAN algorithm; Substituting the similarity matrix into the DBSCAN algorithm, and calculating the respective labels of the two titles according to the preset neighborhood radius E for defining density and the threshold M for defining a core point; Repeating the above steps until all the titles are labeled with corresponding labels; Grouping the titles with the same label into a class. After all the titles are classified, ranking according to the number of titles in each class, and selecting the top X classes. Each class is a clustering result cluster.
3. The method according to claim 2, wherein The selecting a group of main bodies and obtaining the similarity between two main bodies in this group includes: Search for related words of one of the entities through the trained word vector model, and take the top Y related words with the highest occurrence frequency; If the other entity to be compared appears in the top Y related words with the highest occurrence frequency, the similarity value of the two groups of entities is 1. If the other entity to be compared does not appear in the top Y related words with the highest occurrence frequency; Then calculate the similarity between the two entities through the Jaccard similarity calculation formula.
4. The method according to claim 3, wherein The Jaccard similarity calculation formula is shown as follows: In the formula, A and B respectively represent two entities.
5. An automatic construction device for a main event library, characterized in that The device includes: Data acquisition module: used to set the time interval according to the time span, and segment and acquire data from the message queue at the preset time interval. The data includes data titles, data resource locators, and release times; Title clustering module: used to cluster the data titles obtained within each time interval through the DBSCAN algorithm, select the top X clustering results in the clustering results, and regard each selected clustering result as a clustering result cluster; Clustering different titles through the DBSCAN algorithm respectively includes: Split all the obtained data titles into several segments, cluster each segment of data titles through the DBSCAN algorithm, and select the top X clustering results in the clustering results of this segment of data titles; Combine the X clustering results of each segment of data titles, and then perform secondary clustering through the DBSCAN algorithm, and select the top X clustering results to form a clustering result cluster; Entity selection module: used to extract the entities of each title in each clustering result cluster through LAC. If the number of entities of the title is greater than 1, select the first entity that appears as the entity of the title. After obtaining the entities of all titles in the clustering result cluster, take the entity with the largest number of occurrences as the entity of the clustering result cluster; Entity merging module: used to merge the entities within each time interval through the DBSCAN algorithm respectively, merge the titles within the same entity, and then merge all the entities within the time span; sort the data resource locators according to the release time for all titles within the same entity to construct an entity event library; The merging of the entities within each time interval through the DBSCAN algorithm respectively includes: Select a group of entities and obtain the similarity between two entities in the group; Preset the neighborhood radius E for defining density and the threshold M for defining core points in the DBSCAN algorithm; Substitute the similarity between the two entities into the DBSCAN algorithm, and calculate the respective labels of the two entities according to the preset neighborhood radius E for defining density and the threshold M for defining core points; Repeat the above steps until all entities are labeled with corresponding labels; Merge the entities with the same label.
6. A storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by the main controller, it realizes each step in an automatic construction method of an entity event library as described in any one of claims 1-4.
Citation Information
Patent Citations
Automatic mining system and method of news events in large-scale data
CN103020251A
Public opinion analysis method, system, apparatus and storage medium
CN109408804A