Information event extraction method and system
By receiving seed event sentences and calculating similarity based on vectorized representations, the problem of low efficiency and accuracy of information event extraction in the prior art is solved, and efficient and accurate information event extraction is achieved, reducing cost and complexity.
Patent Information
- Application Number
- CN202411418759.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-17
- Filing Date
- 2024-10-11
- Publication Date
- 2025-05-06
AI Technical Summary
When the prior art extracts information events from massive information, the accuracy and recall rate are difficult to meet the needs of commercialization, and requires a large amount of corpus training and high cost annotation, the model is complex and misjudgment is prone to occur.
By receiving seed event sentences, the similarity between seed event sentences and information event sentences is calculated based on vectorized representations, and the results of extracted information event sentences are determined, avoiding the need for pre-defined templates and large-scale corpus training.
It improves the efficiency and accuracy of information event extraction, reduces labor costs and model complexity, and expands the scope of application to massive network data.
Smart Images

Figure CN119939000A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of massive information extraction, and in particular to an information event extraction method and system. Background Art
[0002] With the rapid development of science and technology, the Internet plays an increasingly important role in people's daily life and work, and has become the main channel for people to obtain information. However, as a borderless "big network", the Internet covers a wide range of information and has a particularly large amount of data. How to accurately find information (events) with targeted value from these massive amounts of information has become a research hotspot in the field of network retrieval technology. Taking news events as an example, daily news events emerge in an endless stream around the world, and a large number of reports appear on major news websites in the form of unstructured text, recording various events that occur in the real world. By fully exploring the value behind news event data, scientific and effective support can be provided for content production, user personalized needs, upper-level decision-making and other matters. In order to achieve this goal, people have conducted continuous exploration.
[0003] One feasible method in the prior art is the "pattern matching method". This method uses professional knowledge to predefine event extraction templates and extract news events based on the templates. Although this method can realize event extraction from massive information events to a certain extent, since the predefined templates cannot cover all the rules, the non-templated sentence features cannot be considered, resulting in the accuracy and recall of event extraction being difficult to meet commercial requirements. An improved method is to adopt the "machine learning method" based on deep learning theory. This method regards each task as a classification problem, first collects training corpus, annotates the corpus, and then uses natural language processing tools to construct trigger word features, identifies events based on trigger words, and finally completes the extraction of news events through the training of machine learning models. However, this method requires a large amount of corpus training, the corpus annotation cost is high, and the model training difficulty and complexity are large, resulting in low event extraction efficiency; since the judgment of news events mainly depends on trigger words, when the trigger words are polysemous and ambiguous, the judgment of news events is prone to misjudgment, thereby reducing the accuracy of event extraction. Summary of the invention
[0004] The embodiments of the present application provide an information event extraction method and system for improving the efficiency of extracting information events from massive data and increasing the accuracy of information event extraction.
[0005] On the one hand, the information event extraction method provided in the embodiment of the present application includes:
[0006] receiving at least one seed event sentence, wherein the seed event sentence can reflect information event extraction requirements;
[0007] Calculating the similarity between each seed event sentence and information event sentences in an information event sentence library based on the vectorized representation results of the seed event sentences, wherein the information event sentence library is a vector database, and the vector database contains vectorized representation results of information event sentences captured from the Internet;
[0008] The extracted information event sentence conclusion is determined according to the calculated similarity between the seed event sentence and the information event sentence.
[0009] Preferably, a post-processing rule for an event corresponding to the seed event sentence is determined, wherein the post-processing rule is a knowledge rule based on expert knowledge, and the method further comprises:
[0010] Matching and processing the information event sentence results according to the post-processing rules;
[0011] Adjusting the similarity of information event sentences that meet the post-processing rules;
[0012] The information event sentence result after adjusting the similarity is taken as the final information event sentence result.
[0013] Preferably, the method further comprises:
[0014] The information event results are grouped and the similarities of the information event sentences in the groups are summed up;
[0015] The set of information event sentences whose similarities and values meet the predetermined conditions is taken as the information event result.
[0016] Preferably, the information event sentence library is constructed according to the following steps:
[0017] Crawl information event data from the Internet and process it in different bureaus to obtain information event sentences;
[0018] Determine the vector representation of the information event sentence;
[0019] The vector representation of the information event sentence and the address of the information event sentence corresponding to the information event sentence are stored in the information event sentence library.
[0020] Preferably, receiving at least one seed event sentence includes: receiving one or more seed event sentences input in a seed event sentence dedicated interface, and / or receiving one or more seed event sentences selected from a seed event sentence library, wherein the seed event sentence library is a seed event sentence candidate template established in advance according to seed event sentence generation rules.
[0021] Preferably, the seed event sentence generation rule includes: using event sentences in hot events that reach a hot threshold within a predetermined period as seed event sentences; and / or using event sentences formed after standardizing information event extraction requirements as seed event sentences.
[0022] Preferably, the method of determining the extracted information event sentence results based on the calculated similarity between the seed event sentence and the information event sentence includes: sorting the seed event sentence and the information event sentence based on the calculated similarity; determining a predetermined number of information event sentences ranked at the top / bottom as the extracted information event sentence results, or determining the information event sentences whose similarity values between the seed event sentence and the information event sentence are greater than a predetermined threshold as the extracted information event sentence results.
[0023] On the other hand, the information event extraction system provided by the embodiment of the present application includes: a seed event sentence receiving module, a similarity calculation module, and an event sentence result determination module, wherein:
[0024] The seed event sentence receiving module is used to receive at least one seed event sentence, and the seed event sentence can reflect the extraction requirements of the information event;
[0025] The similarity settlement module is used to calculate the similarity between each seed event sentence and the information event sentences in the information event sentence library based on the vectorized representation results of the seed event sentences, wherein the information event sentence library is a vector database, and the vector database contains the vectorized representation results of the information event sentences captured from the network;
[0026] The event sentence result determination module is used to determine the extracted information event sentence result according to the calculated similarity between the seed event sentence and the information event sentence.
[0027] Preferably, the system also includes: a grouping module, used to group the information event sentence results and sum the similarity values of the information event sentences in the group; the event sentence result determination module, specifically used to take the set of information event sentences whose similarities and values meet predetermined conditions as the information event result.
[0028] Preferably, the system includes a seed event sentence construction module and an information event sentence library.
[0029] The seed event sentence construction module is used to establish a seed event sentence candidate template according to a seed event sentence generation rule, wherein the seed event sentence generation rule includes: using event sentences in hot events that reach a hot threshold within a predetermined period as seed event sentences; and / or using event sentences formed after standardizing information event extraction requirements as seed event sentences;
[0030] The information event sentence library is constructed according to the following steps: crawling information events from the Internet and performing sentence segmentation processing to obtain information event sentences; determining the vector representation of the information event sentence; storing the vector representation of the information event sentence and the address of the information event corresponding to the information event sentence in the information event sentence library.
[0031] In the embodiment of the present application, the seed event sentence and the information event sentence are vectorized respectively, and the similarity is calculated, and the result is selected based on the calculated similarity. Compared with the prior art, on the one hand, it is no longer necessary to predefine the template that frames the information event sentence, which not only saves the definition process for the information event sentence with template features and improves efficiency, but also for the information event sentence outside the template features, it is included in the information extraction scope, and it is entirely possible to be extracted as an information event in the end, thereby avoiding the limitation of the template itself and expanding the scope of application for massive network data. On the other hand, there is no need to collect and annotate training corpus and construct trigger word features in large quantities like machine learning methods, thereby greatly saving labor costs and reducing complexity. On the other hand, since the embodiment of the present application performs similarity calculation on two sets of vectors, the vector representation contains the context and semantic relationship of the same information report, and each information event sentence is no longer viewed in isolation, which is conducive to improving the accuracy of information extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0033] Figure 1 A flowchart of an embodiment of the information event extraction method provided by the present application;
[0034] Figure 2 A scene diagram of an embodiment of the information event extraction method provided in this application;
[0035] Figure 3 This is a framework diagram of the information event extraction system provided in this application. DETAILED DESCRIPTION
[0036] In order to better understand the technical solution of the present application, some basic terms are properly explained before introducing the complete embodiment. In the embodiment of the present application, basic elements such as seed event sentences, information event sentences, news event sentences, and basic operations such as information extraction and mining will be mentioned. In the network world, there are all kinds of massive information. The existence form of this information can be text, table, picture, video, etc., and the content of the information can be to express events, convey knowledge, spread emotions, etc. However, in the embodiment of the present application, the focus is mainly on "event sentences". The "event" of the "event sentence" frames the characteristics of the data object processed by the present application, and the "sentence" of the "event sentence" expresses the existence form of the data object processed by the present application. Of course, the "event sentence" here is also just an input for the processing object. Before the input, for data objects in other forms and other contents, if the "event sentence" can be converted by preprocessing, it is naturally also in the discussion of the embodiment of the present application. There are seed event sentences, information event sentences, and news event sentences in the news field. These are the differences in the application of event sentences, and do not mean that they are essentially different. The seed event sentence is a "rule" used to express the information event extraction requirements. The information event sentence is a corpus representing the information of a certain event formed after processing from the Internet, through which the information event itself can be traced back. The relevant concepts are further explained in conjunction with the embodiments below.
[0037] Reference Figure 1 and Figure 2 ,in Figure 1 The following is a flow chart of an information event extraction method provided by an embodiment of the present application. Figure 2 The scene diagram of this embodiment is shown. This embodiment includes:
[0038] Step S11: receiving at least one seed event sentence, wherein the seed event sentence reflects information event extraction requirements.
[0039] The seed event sentence is a key factor in information event extraction. Different from the "keywords" in common retrieval systems, it is actually a sentence-based description of the information event that the user of the present application method wants to extract, showing two important characteristics: First, it is "event-based" in content, not a fragmented "point", with time, place, event, etc. as the elements of the "event". Of course, in the actual application process, the user is not required to provide complete event information, as long as it has event-based features and contains event-based elements. For example, "eating" cannot be regarded as an event. It is just a word that represents a universal meaning. However, "a well-known actor eating in the China World Trade Center" is an event. The information has been concretized and specified due to the increase of event elements. However, it still needs to be explained that sometimes due to the existence of background conditions, the "event" is not presented in the form of a "complete event". It is possible that a word also reflects an event, and an event can also be expressed by a word. This "overlap" does not hinder the "event" feature of the seed event sentence. As a technician in this field, you can understand the substantial event nature under the representational form. Second, it is "sentence-based" in form, not a single word, word or file. According to the needs of information event extractors, seed event sentences reflect the topics, contents, types, time periods, popularity, etc. that extractors need to pay attention to from massive data, but when presented, they are described in a sentence-like manner and have the characteristics of a "sentence", regardless of whether the sentence is long or short, or whether any grammatical components are omitted.
[0040] After explaining the "seed event sentence", the construction of the seed event sentence is further explained below. In fact, the seed event sentence can be constructed in a variety of ways, simple shortcuts are possible, and complex customized methods are also possible. For example, a relatively simple way is to provide a dedicated interface for seed event sentence construction, and the information event extractor directly inputs it in the interface. Another feasible way is to establish multiple seed event sentence templates in advance according to the seed event sentence generation rules based on possible topics, and provide them to the information event extractor in different categories, so that he can choose from these candidate preset seed event sentences. The seed event sentence generation rules can be to use the event sentences in the hot events that reach the hot threshold within a predetermined period as seed event sentences, or to use the event sentences formed after standardizing the information event extraction requirements as seed event sentences. Of course, in addition to direct input, pre-selection methods, etc., a customized interface can also be opened, and personnel with professional knowledge can construct seed event sentences according to the rules.
[0041] The reason why the seed event sentence is called a "seed" is that it has the characteristics of typicality, representativeness, and source. A seed event sentence will be associated with a class, a group of information event sentences or information events. The information event extractor can choose to construct a seed event sentence that is sufficient to cover all the required information events, or construct multiple seed event sentences from different angles as needed. When multiple seed event sentences are constructed, they can be stored in an array, for example, original_sentence_seeds = [seed event sentence 1, seed event sentence 2, seed event sentence 3, ...]. In order to extract information event sentences more accurately, multiple seed event sentences can be divided into themes, categories, and fields according to different standards, that is, the type of seed event sentence event_type is determined, and placed under different arrays or different fields of the same array according to the type. It should be noted that each seed event sentence can be a relatively independent event sentence under the same theme, category, field, etc., or it can be multiple seed event sentences derived from a seed event sentence. These seed event sentences have the same essential meaning, but there are differences in expression. For example, the dialect expression, foreign language expression, elegant expression, popular expression, etc. (derived seed event sentences) of the same seed event sentence (standard seed event sentence) can be regarded as multiple independent seed event sentences.
[0042] Step S12: calculating the similarity between the seed event sentence and the information event sentences in the information event sentence library based on the vectorized representation result of the seed event sentence, wherein the information event sentence library is a vector database, and the vector database contains the vectorized representation results of the information event sentences captured from the Internet;
[0043] In order to calculate the similarity, both the seed event sentence and the information event sentence captured from the Internet must be vectorized in the embodiments of the present application. There are many ways to vectorize the seed event sentence. For example, the open source Sentence-BERT model can be used to calculate the vector representation of each seed event sentence, so that the original seed event sentence is changed to sentence_seeds = [vector representation of seed event sentence 1, vector representation of seed event sentence 2, vector representation of seed event sentence 3, ...]. Similarly, taking news events as an example, a large amount of news event data crawled from the Internet, after the natural language processing tool is processed by the news corpus bureau to obtain the news event sentence, can also use the Sentence-BERT model to calculate the vector representation of each news event sentence, and then store the calculation results in the vector database. Since the news event sentences captured from the network are derived from news events, in order to ensure the traceability of the event, when storing the vector representation of the news event sentence, the news event from which the news event sentence comes, as well as the storage address of the news event on the network can also be stored. Therefore, the data structure of the news event sentence storage can be reflected in the following example:
[0044] {"sentence_id":5678,
[0045] "news_id":1234,
[0046] "news_url":"https: / / udn.com / news / story / 9213 / 7800548?from=udn-catelistnews_ch2",
[0047] “sentence”: “(news event original sentence)”,
[0048] "sentence_embedding":[-0.17820226,0.6134248,0.23683488,-0.15526013,-0.3248349,0.02237497,……]}
[0049] Among them: sentence_id is the unique number of the news event sentence, news_id and news_url are the unique number and url of the news where the news event sentence is located, sentence is the original text of the news event sentence, and sentence_embedding is the vector representation of the news event sentence.
[0050] After the seed event sentence and the information event sentence are vectorized, some algorithms can also be used to establish a corresponding vector index for the vector representation. For example, in the news event sentence library, the HNSW (Hierarchical Navigable Small World graphs) algorithm is used to establish the index of the news event sentence representation vector. Then, the vector representation of the seed event sentence is used to search for event sentences similar to the seed event sentence in the news event sentence library, that is, similarity calculation is performed, and the cosine similarity calculation method can be specifically adopted. After the similarity is calculated, subsequent processing is performed according to the similarity result. During the specific calculation, each seed event sentence can be calculated for similarity with each information event sentence in the information event sentence library. If there are multiple seed event sentences, multiple similarity values corresponding to each information event sentence can be obtained. These multiple similarity values can be directly used as independent results, or the result after weighted average can be used as the final similarity value. For example, there are three seed event sentences: seed event sentence 1, seed event sentence 2, and seed event sentence 3, and there are 100 news event sentences. Then, for seed event sentence 1, 100 similarity values can be obtained, for seed event sentence 2, 100 similarity values can be obtained, and for seed event sentence 3, another 100 similarity values can be obtained.
[0051] Step S13: Determine the extracted information event sentence result according to the calculated similarity between the seed event sentence and the information event sentence.
[0052] After obtaining the similarity between the seed event sentence and the information event sentence through similarity calculation, the information event sentences can be sorted according to their similarity values. After sorting, according to the needs of the information event extractor, a certain number of information event sentences (for example, the top K news event sentences with the highest similarity) that are ranked first (in descending order) or last (in ascending order) can be selected as the corresponding output results, or a threshold can be set in advance, and the information event sentences greater than the threshold can be used as the output results corresponding to the seed event sentences. For example, the TopK algorithm can be used to select the top K news event sentences with the highest similarity to the seed event sentence and output them as the result set. Taking the aforementioned three seed event sentences and 100 news event sentences as examples, all 300 results can be directly sorted, and the first K that meet the conditions can be selected as the information event sentence output results. It is also possible to select K1 of the seed event sentences 1 as the information event sentence output result set 1 by sorting, select K2 of the seed event sentences 2 as the information event sentence output result set 2 by sorting, and select K3 of the seed event sentences 3 as the information event sentence output result set 3 by sorting, and then use these three result sets as the final output results. It is also possible to use the news event sentences as a benchmark, perform weighted synthesis of the three similarity values of each news event sentence to obtain a numerical value, and then select K of them as the information output results by sorting. In short, the embodiments of the present application have no special restrictions on the sorting algorithm, the number of layers of sorting, the selection of sorting results, etc., depending on how to better meet the needs of the information event extractor.
[0053] In the embodiment of the present application, the seed event sentence and the information event sentence are vectorized respectively, and the similarity is calculated, and the result is selected according to the calculated similarity. On the one hand, compared with the existing pattern matching method, it is no longer necessary to predefine the template for framing the information event sentence, which not only saves the definition process for the information event sentence with template features and improves efficiency, but also includes the information event sentence with template features outside the template features into the information extraction scope, and it is entirely possible to be extracted as an information event in the end, thereby avoiding the limitation of the template itself and expanding the scope of application for massive network data. On the other hand, compared with the machine learning method, it is not necessary to collect and annotate a large amount of training corpus, nor to construct trigger word features, thereby greatly saving labor costs and reducing training complexity. On the other hand, since the embodiment of the present application is to calculate the similarity of two sets of vectors, the vector representation contains the context and semantic relationship of the same information report, and each information event sentence is not viewed in isolation, which is conducive to improving the accuracy of information extraction. It should be noted that the embodiment of the present application is also different from the syntactic analysis method. The syntactic analysis method divides the text into sentences and words, performs lexical analysis and syntactic analysis on the sentences and words respectively, constructs a syntactic tree, and then obtains the information event element information by parsing the syntactic tree, thereby completing the extraction of information events. This technology directly relies on syntactic analysis of sentences for the extraction of information events, and has high requirements on the accuracy of syntactic analysis tools. In addition, the coverage of syntactic analysis tools on information events is low, and it is difficult to achieve the same satisfactory accuracy as the embodiment of the present application.
[0054] The above examples show the basic process of the embodiments of the present application. In the practical process, all or individual steps in the above embodiment process can be optimized as needed to obtain better technical effects and further improve the efficiency, accuracy and recovery rate of information event extraction. The following is an exemplary description.
[0055] One of the exemplary optimization directions: After calculating the similarity result between the seed event sentence and the information event sentence in the aforementioned step S12, the similarity calculation result can be post-processed (Post Processing), and the pre-defined "accuracy improvement rules" can be fully utilized to fine-tune the calculation result to improve the accuracy range of the information event sentence similarity calculation result. This optimization direction requires the pre-determination of "post-processing rules", which are established based on expert knowledge and reflect the posterior knowledge formed by long-term experience and with strong objectivity. It can achieve the elevation of the position of the target information event sentence in the similarity ranking table, so as to be more conducive to locking the truly needed, mined, and valuable news event sentences in the result set.
[0056] For example, the post-processing rules for the events corresponding to the seed event sentences are defined as follows:
[0057] post_processing={“co_occurren”:[“×× / ××”,“××× / ×××”],“score”:0.1};
[0058] The post-processing rules can be included in the data structure of the seed event sentence and are expressed as follows:
[0059] {"event_id":"1",
[0060] "event_type":"Meet / Meet",
[0061] “original_sentence_seeds”: [seed event sentence 1, seed event sentence 2, seed event sentence 3, …],
[0062] “sentence_seeds”: [vector representation of seed event sentence 1, vector representation of seed event sentence 2, vector representation of seed event sentence 3, …],
[0063] "post_processing":{"co_occurren":["×× / ××","××× / ×××"],"score":0.1}}
[0064] Among them: event_id is the seed event sentence number, event_type is the time type, original_sentence_seeds is the original text of the seed event sentence, sentence_seeds is the vector representation of the seed event sentence, post_processing is the post-processing rule, co_occurrence is the keyword that appears in each news event sentence in the calculated similarity result, and score is the corresponding news event sentence bonus value when the key co-occurrence rule is met.
[0065] Based on the aforementioned post-processing rules, the post-processing rules contained in the post_processing field are used to process each news event sentence: if the news event sentence currently being processed meets the co-occurrence rule defined in co_occurrence, that is, the keywords defined in co_occurrence appear simultaneously in one news event sentence or multiple news event sentences, then the similarity value of the news event sentence is added with the score value. Through this processing, the similarity of news event sentences that are highly relevant to the needs of information event extractors is increased, so that the news events that the information event extractors really want are more likely to be presented in the results, which is conducive to improving the accuracy of finding news events.
[0066] Exemplary optimization direction 2: Aggregate information event sentences based on similarity calculation results. For example, still taking news event sentences as an example, group them according to the news_id attribute of the news event sentences (group by news_id), and actually group the news event sentences belonging to the same news report into one place, and calculate the sum of the similarity values of all news event sentences in each group. The calculated sum value represents the similarity value between the news report (article) and the seed event sentence, which is recorded as: news_similarity, to form a news event result set. The news event result set is arranged in descending (or ascending) order according to the news_similarity field, and a similarity ranking list of news events and seed event sentences is obtained, and then a certain number of news reports (news event sentences) that are ahead or behind are selected from the similarity ranking list as the result output according to actual needs. It should be noted that only the selected news event sentences can be presented, and the news report (news article) where the news event sentence is located can also be presented, or both can be presented. When clicking on the news event sentence, it can be linked to the news report, so that the event process can be more fully understood.
[0067] This optimization method aggregates information event sentences and sorts them by similarity at the information event level (paragraph level), which reflects the association between different information event sentences and the context and semantic relationship of information event sentences in the same paragraph, thereby making the output information events more accurate.
[0068] The above embodiment describes in detail the method of the embodiment of the present application. Corresponding to the method, the embodiment of the present application also provides an information event extraction system. Figure 3 , which shows a structural diagram of an embodiment of the system of the present application. The information event extraction system comprises: a seed event sentence receiving module U31, a similarity calculation module U32, and an event sentence result determination module U33, wherein:
[0069] A seed event sentence receiving module U31 is used to receive at least one seed event sentence, where the seed event sentence can reflect the extraction requirements of the information event;
[0070] The similarity calculation module U32 uses the HNSW (Hierarchical Navigable Small World graphs) algorithm to establish a vector index of the news event sentence representation vector. At the same time, the news event sentences with similarity to the seed event sentences are retrieved from the vector database in combination with the seed data sentence records to be matched, and are arranged in descending order according to the similarity value. Then, the TopK algorithm or the threshold segmentation algorithm is used to filter and extract the result set that meets the conditions;
[0071] The event sentence result determination module U33 is used to determine the extracted information event sentence result according to the calculated similarity between the seed event sentence and the information event sentence.
[0072] The above system embodiment can also achieve the technical effects of the above method embodiment of the present application. To avoid repetition, it is not repeated here. In addition, the system embodiment of the present application can also include a seed event sentence construction module U34, an information event library U35, and a post-processing module U36, etc. Adding these modules can further optimize the technical effects of the embodiment of the present application.
[0073] The seed event sentence construction module U34 is used to construct the seed event sentence and save it to the database. Taking news events as an example, the attributes that can be recorded in the constructed seed event sentence include: event_id, event_type, original_sentence_seeds, sentence_seeds, post_processing, and post_processing includes co_occurren and score. Among them, event_id is the event number, event_type is the time type, original_sentence_seeds is the original text of the seed event sentence, sentence_seeds is the vector representation of the seed event sentence, post_processing is the post-processing rule, co_occurren is the keyword that co-appears in the target event sentence, and score is the bonus value of the target event sentence when the keyword co-occurrence rule is met.
[0074] The information event sentence library U35 is used to collect a large amount of news data from the Internet through Internet crawler technology, and clean and parse it to extract the news text content that does not contain HTML tags. When constructing the information event library, each piece of news collected can be processed by natural language processing tools to obtain news event sentences. The vector representation of each news event sentence is calculated using the Sentence-BERT model and stored in the vector database. The stored data structure contains the original source of the event sentence, making the event sentence traceable. The attributes of the news event sentence record may include: sentence_id, news_id, news_url, sentence, sentence_embedding. Among them: sentence_id is the unique number of the news event sentence, news_id and news_url are the unique number and url of the news where the news event sentence is located, sentence is the original text of the news event sentence, and sentence_embedding is the vector representation of the news event sentence.
[0075] The post-processing module U36 applies the pre-defined relevance and accuracy improvement rules (co_occurrence and score in post_processing) in the seed event sentence to recalculate and sort the search results to improve the similarity and accuracy of the news event sentence search results, which is reflected in the sentence-level similarity. At the same time, the news event sentences in the search results are grouped (group by news_id) and aggregated according to the news_id attribute, and the sum of the similarity values of all news event sentences in each group is calculated as the similarity value between the news article and the seed event sentence, which is reflected in the chapter-level similarity.
[0076] The embodiments of the present application can utilize the aforementioned hardware to realize all the invention purposes, and thus mainly introduce the hardware part (module part). It should be understood by those skilled in the art that the embodiments of the present application can be provided as modules, devices, systems, or related computer program products. Therefore, the present application can be implemented using a complete hardware embodiment, or can be implemented in the form of a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0077] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0078] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0079] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0080] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0081] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0082] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0083] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0084] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for extracting information events, characterized in that: The method comprises: receiving at least one seed event sentence, wherein the seed event sentence can reflect the extraction requirements of the information event; Calculating the similarity between each seed event sentence and an information event sentence in an information event sentence library based on the vectorized representation result of the seed event sentence, wherein the information event sentence library is a vector database, and the vector database contains the vectorized representation results of the information event sentences captured from the Internet; The extracted information event sentence result is determined according to the calculated similarity between the seed event sentence and the information event sentence.
2. The method according to claim 1, characterized in that The method further includes: determining a post-processing rule for an event corresponding to the seed event sentence, wherein the post-processing rule is a knowledge rule based on expert knowledge. After determining the extracted information event sentence result, the method further includes: Matching and processing the information event sentence results according to the post-processing rules; Adjusting the similarity of information event sentences that meet the post-processing rules; The information event sentence result after adjusting the similarity is used as the final extracted information event sentence result.
3. The method according to claim 1, characterized in that The method further comprises: The information event sentence results are grouped, and the similarity values of the information event sentences in the groups are summed up; The set of information event sentences whose similarities and values meet the predetermined conditions is taken as the information event result.
4. The method according to claim 1, characterized in that The information event sentence library is constructed according to the following steps: Crawl information events from the Internet and perform sentence processing to obtain information event sentences; Determine the vector representation of the information event sentence; The vector representation of the information event sentence and the address of the information event corresponding to the information event sentence are stored in the information event sentence library.
5. The method according to claim 1, characterized in that: The receiving at least one seed event sentence comprises: One or more seed event sentences inputted in the seed event sentence dedicated interface are received, and / or one or more seed event sentences selected from a seed event sentence library are received, wherein the seed event sentence library is a seed event sentence candidate template established in advance according to a seed event sentence generation rule.
6. The method according to claim 5, characterized in that The seed event sentence generation rules include: The event sentences in the hot events reaching the hot threshold within a predetermined period are used as seed event sentences; and / or the event sentences formed after standardizing the information event extraction requirements are used as seed event sentences.
7. The method according to claim 1, characterized in that The step of determining the extracted information event sentence result according to the calculated similarity between the seed event sentence and the information event sentence includes: Sort the seed event sentences and information event sentences according to the calculated similarity; A predetermined number of information event sentences ranked at the top / bottom are determined as the extracted information event sentence results, or information event sentences whose similarity values between the seed event sentences and the information event sentences are greater than a predetermined threshold are determined as the extracted information event sentence results.
8. An information event extraction system, characterized in that: The system comprises: a seed event sentence receiving module, a similarity calculation module, and an event sentence result determination module, wherein: The seed event sentence receiving module is used to receive at least one seed event sentence, and the seed event sentence can reflect the extraction requirements of the information event; The similarity calculation module is used to calculate the similarity between each seed event sentence and the information event sentences in the information event sentence library based on the vectorized representation results of the seed event sentences, wherein the information event sentence library is a vector database, and the vector database contains the vectorized representation results of the information event sentences captured from the network; The event sentence result determination module is used to determine the extracted information event sentence result according to the calculated similarity between the seed event sentence and the information event sentence.
9. The system according to claim 8, further comprising: A grouping module, used to group the information event sentence results and sum the similarity values of the information event sentences in the group; The event sentence result determination module is specifically used to take the information event sentence set whose similarity and value meet the predetermined conditions as the information event result.
10. The system according to claim 8, comprising a seed event sentence building module and an information event sentence library, The seed event sentence construction module is used to establish a seed event sentence candidate template according to a seed event sentence generation rule, and the seed event sentence generation rule includes: Taking event sentences in hot events that reach the hot threshold within a predetermined period as seed event sentences; and / or, using the event sentence formed after standardizing the information event extraction requirements as the seed event sentence; The information event sentence library is constructed according to the following steps: crawling information events from the Internet and performing sentence segmentation processing to obtain information event sentences; determining the vector representation of the information event sentence; storing the vector representation of the information event sentence and the address of the information event corresponding to the information event sentence in the information event sentence library.