Method and system for constructing a Chinese event library and analyzing and predicting meta-events based on the meta-event library

The Chinese meta-event library uses deep learning to address GDELT's limitations, enhancing detection and prediction accuracy through semantic analysis and visualization.

CN116383331BActive Publication Date: 2025-07-15TRS INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310001827.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-03
Publication Date
2025-07-15
Estimated Expiration
2043-01-03

AI Technical Summary

Technical Problem

The existing GDELT meta-event library has problems such as severe data bloat, insufficient processing accuracy, high false positive rate, inability to perform semantic analysis and meta-event relationship reasoning, and lack of Chinese event library, which cannot meet the processing and prediction needs of Chinese news and intelligence data.

Method used

A Chinese event library is constructed, through meta-event extraction, co-reference, association and aggregation processing, the core elements are extracted using the BERT+BiLSTM+CRF model, semantic analysis and visual prediction are used using deep learning, and situational awareness analysis is performed in combination with GIS.

Benefits of technology

It realizes accurate identification and visualization of Chinese meta events, can predict the next development trend, improves processing accuracy and efficiency, and meets the analysis needs of Chinese news and intelligence data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116383331B_ABST
    Figure CN116383331B_ABST
Patent Text Reader

Abstract

The present invention proposes a method and system for constructing a Chinese event library and analyzing and predicting meta-events based on the meta-event library. The specific steps of the method for constructing the Chinese event library include: S1: Meta-event extraction; S2: Meta-event coreference; S3: Meta-event association; S4: Meta-event aggregation; S5: Finally, through S1-S4, a meta-event extraction library, a meta-event coreference library, a meta-event association library, and a meta-event topic library are formed, which together constitute the Chinese event library. A method for visual analysis and prediction of meta-events based on the meta-event library specifically includes: S1: Meta-event library retrieval; S2: Meta-event topic analysis; S3: Meta-event prediction analysis. The present invention constructs a Chinese event library suitable for processing, analyzing, and predicting Chinese news and intelligence data, which is not limited to data statistics, realizes semantic analysis of events, and through this Chinese event library, visualizes the meta-event context, makes the recognition of Chinese meta-events more accurate, and can predict the next development trend of meta-events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method and system for constructing a Chinese event library and analyzing and predicting meta-events based on the meta-event library. Background Art

[0002] Currently, the number of conflict meta-events worldwide is on the rise, and the actors and interests involved in conflict meta-events are becoming increasingly diverse. In the face of the continuous occurrence of conflict meta-events, how to accurately detect the occurrence of meta-events such as violence and conflict and make response decisions has become a research hotspot.

[0003] Therefore, effectively constructing a meta-event library has become the premise and foundation for emergency decision-making. Through the meta-event library, human real-world activities can be comprehensively recorded, and a large amount of meta-event data is recorded in the meta-event library, realizing a comprehensive mapping of the real world and human activities. Among them, the GDELT meta-event library is the most famous. GDELT discovers and records the main meta-events that have occurred in human society since 1979 from global news media data in more than 100 languages. GDELT has become the most widely used public dataset in meta-event analysis, meta-event graph, trend prediction, meta-event early warning, and computational sociology research.

[0004] However, GDELT itself also has many deficiencies:

[0005] 1. Serious data expansion and extremely large scale. After 2002, the data has grown explosively, and the data volume exceeded 2.5TB in 2018.

[0006] 1) Perform text analysis on each collected page, and extract meta-events and related people, organizations, geographical locations, etc. sentence by sentence;

[0007] 2) Each page may generate multiple meta-event records;

[0008] 3) Do not perform duplicate removal on similar pages; do not perform normalization or duplicate removal on the same meta-events.

[0009] 2. Insufficient processing accuracy and very high false alarm rate;

[0010] Adopting a software processing method based on dictionaries and rules, misreporting and missing reporting are relatively serious, and the entries cannot be updated in a timely manner.

[0011] 3. Unable to perform semantic-level analysis

[0012] Only the meta-event type, meta-event subject, and meta-event source URL are extracted, and the description information is not extracted. It is impossible to determine the specific meta-events that occurred, and it is impossible to obtain the development context of the meta-events. Only statistical information can be obtained, and text semantic analysis cannot be performed.

[0013] 4. Unable to perform meta - event relationship reasoning and specific meta - event prediction

[0014] The GDELT meta - event library does not save the association relationships between meta - events, making it impossible to analyze the development laws of meta - events and predict the specific progress of the next meta - event.

[0015] 5. Lack of Chinese event library

[0016] The data collected by GDELT is all translated into English for processing, and most of the dictionaries and classifications it uses are biased towards the United States, which cannot meet the processing of Chinese news and intelligence data.

[0017] In particular, for meta - event analysis, it is basically mainly about public opinion analysis at present, targeting the text - level. And the content included at the text - level is numerous and miscellaneous, mainly based on statistical analysis. It is very difficult to analyze the semantic information of meta - events, making it difficult to conduct meta - event prediction, and there is a lack of a meta - event visualization analysis and prediction platform. Summary of the Invention

[0018] Aiming at the problem that there is currently no Chinese event library that can process, analyze, and predict Chinese news and intelligence data, especially for semantic analysis, the present invention provides a method and system for constructing a Chinese event library and analyzing and predicting meta - events based on this meta - event library, constructs a Chinese event library suitable for processing, analyzing, and predicting Chinese news and intelligence data, is not limited to data statistics, realizes semantic analysis of events, visualizes the meta - event context through this Chinese event library, makes Chinese meta - event recognition more accurate, and predicts the next development trend of meta - events through deep fully - connected.

[0019] The specific solution is as follows:

[0020] A method for constructing a Chinese event library, the specific steps are as follows:

[0021] S1: Meta - event extraction: A meta - event refers to an event that is extracted and analyzed. The meta - event extraction is to identify and extract meta - event elements from Chinese news or intelligence texts to form meta - event element texts. The meta - event elements include the time, location, participating roles, and changes in related actions or states when the meta - event occurs. The meta - event extraction includes core meta - event element extraction and meta - event type extraction;

[0022] S2: Meta - event co - reference: The meta - event co - reference is to merge the extracted meta - events and normalize the meta - events based on the common characteristics of the meta - event elements, including: formatting of meta - event time and location elements, resolution of meta - event subject entity references, and meta - event merging;

[0023] S3: Meta-event Association: The meta-event association is to identify the association relationships between meta-events, including: intra-sentence association, intra-discourse association, cross-discourse association, and frequent set mining; the types of association relationships include: condition, causality, concession, sequence, coordination, and co-occurrence.

[0024] S4: Meta-event Aggregation: Further aggregate and analyze the meta-events after the co-reference processing of S2 to form meta-event topics. The aggregation analysis includes: vector representation of the text, calculation of text vector similarity, and automatic clustering to form meta-event topics.

[0025] S5: Constructing a Chinese Event Library: Save the meta-event element data after the meta-event extraction process in S1 in the meta-event extraction library; save the data after the co-reference processing of meta-events in S2 in the meta-event co-reference library; save the relationship data between two meta-events after the meta-event association processing in S3 in the meta-event association library; save the meta-event topic data formed after the meta-event aggregation processing in S4 in the meta-event topic library. The meta-event extraction library, meta-event co-reference library, meta-event association library, and meta-event topic library together constitute the Chinese event library.

[0026] Preferably, the specific method for extracting the core elements of the meta-events in S1 is as follows:

[0027] S111: Use BIO sequence labeling to label the core elements of the meta-events, including: labeling time as B-TIME, I-TIME, labeling location as B-LOC, I-LOC, labeling the subject as B-A0, I-A0, and labeling the object as B-A1, I-A1.

[0028] S112: Use the BERT+BiLSTM+CRF model to extract the core elements of the meta-events labeled by the BIO sequence in S111 and output the results of the core elements of the meta-events.

[0029] S1121: First, use the BERT model to perform word embedding on the text tokens in the core elements of the meta-events labeled by the BIO sequence.

[0030] S1122: Use the memory network model BiLSTM to perform further sequence prediction on the core elements of the meta-events after word embedding in S1121. The sequence prediction method includes: performing sequence prediction by globally modeling the features between trigger words, between core elements of meta-events, and between trigger words and core elements of meta-events.

[0031] S1123: Use the conditional random field (CRF) to calculate the feature scores of the core elements of the meta-events in the sequence prediction results described in S1122 at the current position, the feature scores of label transitions, and the sum of the feature scores of the current position and label transitions. Then, select the core element of the meta-event with the highest sum of feature scores in the entire sequence prediction result as the result of the core element of the meta-event;

[0032] S1124: Output the result of the core element of the meta-event.

[0033] Preferably, the extraction of the meta-event type described in S1 is based on the BERT neural network model. The specific steps are as follows:

[0034] S121: Tokenize, perform part-of-speech tagging, syntactic analysis, extract core elements of meta-events, and extract central words of meta-events for the meta-event sentences of the meta-events, and form corresponding texts;

[0035] S122: Perform vector representation on the text formed by the tokenization results in S121, the syntactic structures obtained after syntactic analysis, the core elements of the meta-events extracted, and the central words of the meta-events extracted respectively;

[0036] S123: Perform deep learning on the vectors obtained in S122 based on the BERT neural network model to obtain the vector features of the meta-event;

[0037] S124: Correspond the major categories in the TRS meta-event classification standard with the vector features of the meta-event obtained in S123 to achieve the initial classification of the meta-event;

[0038] S125: After defining the major category to which the meta-event belongs, further correspond the vector features under the major category of the meta-event with the minor categories in the TRS meta-event classification standard to achieve the secondary classification of the meta-event.

[0039] Preferably, the major categories in the TRS meta-event classification standard include speech meta-events, accident disasters, public health meta-events, political meta-events, judicial meta-events, military meta-events, diplomatic meta-events, work meta-events, life meta-events, sports meta-events, economic meta-events, and security meta-events.

[0040] Preferably, the specific method for normalizing the meta-event in S2 is as follows:

[0041] S211: Standardize the time elements extracted from the meta-event and normalize them to the standard date format: YYYY.MM.DD;

[0042] S212: Normalize the location elements extracted from the meta-event and normalize them to location coordinates: [latitude, longitude], where the north latitude is positive, the south latitude is negative, the east longitude is positive, and the west longitude is negative.

[0043] Preferably, the coreference resolution of the core event main entity in S2 utilizes the BERT-ENE model, and the specific steps are as follows:

[0044] S221: Construct an entity name dictionary using the entity names and alias information of the knowledge base;

[0045] S222: Through the entity description text of the knowledge base, use the BERT pre-trained model to select the vector output at the CLS position as the entity name embedding;

[0046] S223: Match the core event with the entity name dictionary to obtain candidate entities in the short text of different core events;

[0047] S224: Aggregate all pronouns in the candidate entities in the short text that point to the same named entity into a reference chain through the BERT-ENE model.

[0048] Preferably, the core event merging in S2 merges the same or similar core events by similarity comparison and merges the relevant core event element texts. The specific method is as follows:

[0049] S231: Sort according to the occurrence time of the core event, and perform similarity judgment on the core events occurring on the same day;

[0050] S232: Confirm the main body of the core event, and perform similarity judgment on the core events with the same main body;

[0051] S233: Calculate the similarity of the core event types and the similarity of the core event description sentences based on the short text semantic similarity model of the deep neural network;

[0052] S234: Sort the sentences according to the similarity, and merge the core events with a similarity greater than the similarity threshold.

[0053] Preferably, the method for calculating the similarity by the short text semantic similarity model based on the deep neural network in S233 is as follows:

[0054] S2331: Segment the core event text and calculate the word vectors of individual words;

[0055] S2332: Use the central averaging method to superimpose the word vectors in the text and calculate their average value to obtain the text vector of the core event;

[0056] S2333: Calculate the cosine of the angle between the text vectors of two core events to calculate the text vector similarity of the two core events.

[0057] Preferably, the method for annotating the association relationship between core events by intra-sentence association, intra-discourse association, cross-discourse association, and frequent set mining in S3 is as follows:

[0058] S31: Intra-sentence association: The intra-sentence association is to judge the relationship between two meta-events based on intra-sentence correlative words, including adversative, conditional, causal, parallel, and sequential relationships;

[0059] S32: Intra-discourse association: The intra-discourse association is to judge the relationship between two meta-events according to the agent of the meta-event and the occurrence time of the meta-event, including sequential and causal relationships;

[0060] S33: Cross-discourse association: The cross-discourse association is to establish the association between meta-events in different discourses through the co-reference operation of meta-events;

[0061] S34: Frequent set mining: The frequent set mining is to mine the frequently co-occurring meta-events in the co-reference library and establish the co-occurrence relationship.

[0062] Preferably, for the meta-events with the labeled association relationship, the relationship extraction of meta-events is carried out based on the relationship extraction model of LSTM+Attention, and the specific steps are as follows:

[0063] S311: At the word level, add the position relationship and extract features using bidirectional LSTM+Attention;

[0064] S312: At the sentence level, after extracting features, adopt the Attention mechanism for the sentence features;

[0065] S313: Use softmax classification to obtain the meta-event relationship type.

[0066] Preferably, the method of automatic clustering in S4 is:

[0067] S41: Select K meta-events as the centers of the initial classes by using heuristic rules;

[0068] S42: Assign all meta-events to the nearest classes;

[0069] S43: Recalculate the center of each class;

[0070] S44: Repeat steps S2 and S3 for t times;

[0071] S45: Extract the meta-event descriptions from each class to obtain the clustering result.

[0072] Preferably, the meta-event extraction library field is compatible with the GDELT field, and the newly added meta-event is represented by the TRS_ field, and the TRS_ field includes: event description, event source sentence, characters in the event sentence, place in the event sentence, organization in the event sentence, event viewpoint type, event tense, related topics, position of the sentence where the event is located, content of the sentence where the event is located, event field, whether there are similar events, similar event ids, title of the document where the event is located, detailed JSON tags of event elements, original words of the event occurrence time, and storage time.

[0073] Preferably, the MentionNum value in the meta-event extraction library is set to 1, and the MentionNum value in the meta-event co-reference library is set to the co-reference quantity.

[0074] Preferably, the meta-event association library stores the relationship between meta-events, and its fields include: meta-event IDs and basic information of associated relationships, event relationships, event relationship types, storage time, and event relationship weights.

[0075] Preferably, the meta-event topic library stores basic information of the meta-event topic and the meta-event IDs contained therein, and its fields include: ID unique value, topic name, topic content, topic picture, topic map location coordinates, and event list.

[0076] A method for visual analysis and prediction of meta-events based on the Chinese event database, the specific steps are as follows:

[0077] S1: Meta-event database search: Users search for meta-events in the visual interface of the Chinese event database according to their needs. The search methods include: search by meta-event occurrence time, search by meta-event classification, search by meta-event description, search by meta-event participant, search by meta-event occurrence location, and search by meta-event source;

[0078] S2: Meta-event topic analysis: The GIS map is used to customize the topic in the visual interface of the Chinese event library. The topic analysis can be conducted on a certain meta-event or a certain type of meta-event and displayed in the form of charts. The types of topic analysis include: meta-event heat analysis, hot entity analysis, meta-event type analysis, meta-event context analysis, meta-event public opinion field analysis, event map analysis, and meta-event geographic network analysis. The hot entity analysis displays the entity information involved in the meta-event topic in the form of a word cloud, including people, institutions, and places. The meta-event context analysis displays the occurrence and development of the topic meta-event in the form of a timeline. The meta-event geographic network analysis displays the relationship between meta-event participants through the connection of the geographical locations of the meta-event participants.

[0079] S3: Meta-event prediction and analysis: For conflicting meta-events, predict the next development trend of the meta-events. First, quantitatively represent the conflicting meta-events, and then use a prediction model to predict the quantity and intensity of the conflicting meta-events, the Global Conflict Index (GCI), the Local Conflict Index (LCI), and the conflict stage QuadClass.

[0080] Preferably, the content of the quantitative representation of the conflicting meta-events in S3 includes: meta-event coding and quantity, time and location coding, and meta-event attribute coding.

[0081] Preferably, the specific method of the prediction model in S3 is as follows: Use the ActionGeo_CountryCode of the meta-event location information recorded in the Chinese event library as the location where the conflicting meta-event occurs for analysis, calculate the conflict index using the meta-event attribute coding recorded in the Chinese event library, classify the meta-event types using the conflict stage QuadClass, measure the potential impact of the meta-event on regional stability using the GoldsteinScale, and evaluate the importance of a certain meta-event using the number of occurrences NumMentions of a certain meta-event in the Chinese event library.

[0082] Preferably, the calculation method of the Global Conflict Index (GCI) is as follows:

[0083]

[0084] In the formula: ABS(GS i ) is the absolute value of the GoldsteinScale score of each conflicting meta-event; N represents the total number of conflicting news meta-events in a certain country / region; MT i represents the number of mentions of the i-th meta-event; M represents the total number of all news meta-events in a certain country / region.

[0085] Preferably, the calculation method of the Local Conflict Index (LCI) is as follows:

[0086]

[0087] In the formula: ABS(GS i ) represents the absolute value of the GoldsteinScale score of each conflicting meta-event; N represents the total number of conflicting news meta-events in a certain country / region; MT i represents the number of mentions of the i-th meta-event.

[0088] Preferably, the prediction model adopts a deep fully connected method for the model architecture, and the specific steps are as follows:

[0089] S1: Use the meta-events extracted and analyzed from Chinese news or intelligence texts as meta-events, quantitatively represent the meta-events, and perform feature processing;

[0090] S2: Input the primitive event and its vector features at the input layer;

[0091] S3: Use the hidden layer to separate the vector features of the primitive event, predict the primitive event, and evaluate the closeness between the predicted value and the true value to improve the model.

[0092] A system for constructing a Chinese event library, comprising:

[0093] Primitive event extraction module: The primitive event extraction module identifies and extracts primitive event elements from Chinese news or intelligence texts, including a primitive event core element extraction unit for extracting primitive event core elements and a primitive event type extraction unit for extracting primitive event types; The primitive event refers to the event that is extracted and analyzed.

[0094] Primitive event co-reference module: The primitive event co-reference module includes a primitive event time and location element formatting unit for normalizing the primitive event, a primitive event subject entity co-reference resolution unit for aggregating pronouns pointing to the same named entity in different primitive events into a co-reference chain, and a primitive event merging unit for merging the extracted primitive events;

[0095] Primitive event association module: The primitive event association module identifies the association relationships between primitive events, including: intra-sentence association unit, intra-discourse association unit, cross-discourse association unit, frequent set mining unit; The types of association relationships include: condition, causality, transition, succession, parallelism, co-occurrence;

[0096] Primitive event aggregation module: The primitive event aggregation module further aggregates and analyzes the results of the primitive event co-reference module to form primitive event topics, including: a text vector representation unit for representing the text vector using text features, a text vector similarity calculation unit for calculating the similarity of the text vector using the text vector, and an automatic clustering unit for classifying the primitive events into primitive event topics using heuristic rules;

[0097] Chinese event library: Save the primitive event element data processed by the primitive event extraction module in the primitive event extraction library; Save the data processed by the primitive event co-reference module in the primitive event co-reference library; Save the relationship data between two primitive events processed by the primitive event association module in the primitive event association library; Save the primitive event topic data formed by the primitive event aggregation module in the primitive event topic library, and the primitive event extraction library, primitive event co-reference library, primitive event association library, and primitive event topic library together constitute the Chinese event library.

[0098] The present invention provides a method and system for constructing a Chinese event library and analyzing and predicting meta-events based on the meta-event library. The Chinese event library is constructed based on deep learning to make the recognition of Chinese meta-events more accurate; the meta-event extraction library, meta-event coreference library, meta-event association library, and meta-event topic library together constitute the Chinese event library. When performing coreference analysis on meta-events, it includes time and location formatting, meta-event subject reference resolution, meta-event subject entity linking, and meta-event coreference. When performing meta-event coreference, the comparison of meta-event elements includes the similarity of the meta-event occurrence time, meta-event occurrence location, meta-event type, meta-event description, and meta-event verb. When performing meta-event association analysis, it includes intra-sentence association, intra-discourse association, cross-discourse association, and frequent set mining association; the association categories include: causal relationship, parallel relationship, conditional relationship, turning relationship, and sequential relationship; the association analysis method adopts a method based on deep learning, considering both the word level and the sentence level. The word level adopts the method of LSTM+Attention, and the sentence level also adds the method of Attention. The present invention defines the extraction formula of meta-event types, defining 12 major categories and 305 minor categories of meta-event types; the extraction of meta-event types adopts a method based on deep learning, and classification features introduce meta-event core element features, syntactic features, central word features, etc. The present invention constructs a Chinese event library suitable for processing, analyzing, and predicting Chinese news and intelligence data. In this meta-event library, it is not limited to data statistics, realizes semantic analysis of events, and through this Chinese event library, visualizes the meta-event context, makes the recognition of Chinese meta-events more accurate, and predicts the next development trend of meta-events through deep fully connected prediction.

[0099] Meanwhile, the present invention provides a method for visual analysis and prediction of meta-events, adopting situation awareness analysis based on GIS: analyzing from all angles such as time, space, and association, including meta-event heat analysis, meta-event entity analysis, meta-event type analysis, meta-event element analysis, meta-event context analysis, meta-event geographical network visualization analysis, and meta-event public opinion field analysis; on the GIS map, various visualization graphs such as bar charts, pie charts, time axes, and geographical networks are used for analysis. When performing meta-event prediction analysis, a linear fully connected method is adopted for meta-event prediction, which can predict the number of meta-events occurring in the next stage, the meta-event conflict index, and the meta-event development stage. Brief Description of the Drawings

[0100] Figure 1 It is a flowchart of a method for constructing a Chinese event library.

[0101] Figure 2 It is an architecture diagram of a meta-event type recognition model based on the BERT neural network model.

[0102] Figure 3It is an architecture diagram of a relation extraction model based on LSTM+Attention.

[0103] Figure 4 It is a flowchart of a method for visual analysis and prediction of meta-events in the embodiment.

[0104] Figure 5 It is an architecture diagram of a deep fully-connected model in the embodiment.

[0105] Figure 6 It is a system structure diagram of a system for constructing a Chinese event library in the embodiment.

[0106] Figure 7 It is a diagram showing the extraction results of meta-events of "a certain event" in the embodiment.

[0107] Figure 8 It is a line chart of the GoldsteinScale values of the meta-events of "a certain event" per day in the embodiment.

[0108] Figure 9 It is a diagram showing the number of meta-events of "a certain event" from XX date to XX date in the embodiment.

[0109] Figure 10 It is a diagram showing the LCI distribution of "a certain event" from XX date to XX date in the embodiment.

[0110] Figure 11 It is a diagram showing the stage QuadClass of "a certain event" from XX date to XX date in the embodiment.

[0111] Figure 12 It is a diagram showing the prediction results of the meta-event model of "a certain event" in the embodiment. Detailed implementation manners

[0112] The present invention will be further described below in conjunction with the embodiments and the drawings.

[0113] Embodiment 1:

[0114] A method for constructing a Chinese event library is as follows:

[0115] S1: Meta-event extraction: A meta-event refers to an event that is extracted and analyzed. The meta-event extraction is to identify and extract meta-event elements from Chinese news or intelligence texts to form meta-event element texts. The meta-event elements include the time, place, participating roles, and changes in related actions or states of the meta-event. The meta-event extraction includes core meta-event element extraction and meta-event type extraction;

[0116] S2: Coreference of Meta-events: The coreference of meta-events is to merge the extracted meta-events based on the common features of the meta-event elements and to normalize the meta-events, including: formatting of meta-event time and location elements, resolution of entity references of meta-event subjects, and merging of meta-events.

[0117] S3: Association of Meta-events: The association of meta-events is to identify the association relationships between meta-events, including: intra-sentence association, intra-discourse association, cross-discourse association, and frequent set mining; the types of the association relationships include: condition, causality, transition, succession, parallelism, and co-occurrence.

[0118] S4: Aggregation of Meta-events: Further aggregative analysis is performed on the meta-events processed by S2 coreference of meta-events to form meta-event topics, and the aggregative analysis includes: vector representation of texts, calculation of similarity of text vectors, and automatic clustering to form meta-event topics.

[0119] S5: Construction of Chinese Event Library: The meta-event element data after the meta-event extraction processing in S1 is saved in the meta-event extraction library; the data after the coreference processing of meta-events in S2 is saved in the meta-event coreference library; the relationship data between two meta-events after the association processing of meta-events in S3 is saved in the meta-event association library; the meta-event topic data formed after the aggregation processing of meta-events in S4 is saved in the meta-event topic library, and the meta-event extraction library, the meta-event coreference library, the meta-event association library, and the meta-event topic library jointly constitute the Chinese event library.

[0120] Preferably, the specific method for extracting the core elements of meta-events in S1 is as follows:

[0121] S111: Use BIO sequence labeling to label the core elements of meta-events, including: labeling time as B-TIME, I-TIME, labeling location as B-LOC, I-LOC, labeling the subject as B-A0, I-A0, and labeling the object as B-A1, I-A1.

[0122] S112: Use the BERT+BiLSTM+CRF model to extract the core elements of meta-events labeled by the BIO sequence in S111 and output the results of the core elements of meta-events.

[0123] S1121: First, use the BERT model to perform word embedding on the text tokens in the core elements of meta-events labeled by the BIO sequence.

[0124] S1122: Use the memory network model BiLSTM to perform further sequence prediction on the core elements of meta-events after word embedding in S1121, and the sequence prediction method includes: performing sequence prediction by globally modeling the features between trigger words, between core elements of meta-events, and between trigger words and core elements of meta-events.

[0125] S1123: Use the conditional random field (CRF) to calculate the feature scores of the core elements of the meta-events in the sequence prediction result described in S1122 at the current position, the feature scores of label transitions, and the sum of the feature scores of the current position and label transitions. Then, select the core element of the meta-event with the highest sum of feature scores in the entire sequence prediction result as the result of the core element of the meta-event;

[0126] S1124: Output the result of the core element of the meta-event.

[0127] Preferably, the extraction of the meta-event type described in S1 is based on the BERT neural network model. The specific steps are as follows:

[0128] S121: Perform word segmentation, part-of-speech tagging, syntactic analysis, extraction of core elements of the meta-event, and extraction of the central word of the meta-event on the meta-event sentence of the meta-event, and form the corresponding text;

[0129] S122: Perform vector representation on the text formed by the word segmentation result in S121, the syntactic structure obtained after syntactic analysis, the extracted core elements of the meta-event, and the extracted central word of the meta-event respectively;

[0130] S123: Perform deep learning on the vectors obtained in S122 based on the BERT neural network model to obtain the vector features of the meta-event;

[0131] S124: Correspond the major categories in the TRS meta-event classification standard with the vector features of the meta-event obtained in S123 to achieve the initial classification of the meta-event;

[0132] S125: After defining the major category to which the meta-event belongs, further correspond the vector features under the major category to which the meta-event belongs with the minor categories in the TRS meta-event classification standard to achieve the secondary classification of the meta-event.

[0133] The reference for the meta-event type is sorted out to form the TRS meta-event classification standard with reference to "CAMEO Coding", "GB / T 20093-2013 Chinese News Information Classification and Code", and "GBT 35561-2017 Classification and Coding of Sudden Meta-Events", including 12 major categories and 305 minor categories.

[0134] Preferably, the major categories in the TRS meta-event classification standard include speech meta-events, accident disasters, public health meta-events, political meta-events, judicial meta-events, military meta-events, diplomatic meta-events, work meta-events, life meta-events, sports meta-events, economic meta-events, and security meta-events.

[0135] Preferably, the specific method for normalizing the meta-event in S2 is as follows:

[0136] S211: Standardize the time elements extracted from the meta-events and normalize them into the standard date format: YYYY.MM.DD;

[0137] S212: Normalize the location elements extracted from the meta-events and normalize them into location coordinates: [latitude, longitude], where the north latitude is positive, the south latitude is negative, the east longitude is positive, and the west longitude is negative.

[0138] Preferably, the coreference resolution of the meta-event subject entity described in S2 utilizes the BERT-ENE model, and the specific steps are as follows:

[0139] S221: Construct an entity name dictionary using the entity names in the knowledge base and the alias information of the entities;

[0140] S222: Through the entity description text in the knowledge base, use the BERT pre-trained model to select the vector output at the CLS position as the entity name embedding;

[0141] S223: Match the meta-events with the entity name dictionary to obtain candidate entities in the short texts of different meta-events;

[0142] S224: Aggregate all the pronouns in the candidate entities in the short text that point to the same named entity into a coreference chain through the BERT-ENE model.

[0143] Preferably, the meta-event merging in S2 is to merge the same or similar meta-events by similarity comparison and merge the relevant meta-event element texts. The specific method is as follows:

[0144] S231: Sort according to the occurrence time of the meta-events and perform similarity judgment on the meta-events that occur on the same day;

[0145] S232: Confirm the subject of the meta-event and perform similarity judgment on the meta-events with the same subject;

[0146] S233: Calculate the similarity of the meta-event types and the similarity of the meta-event description sentences based on the short text semantic similarity model of the deep neural network;

[0147] S234: Sort the sentences according to the similarity and merge the meta-events with a similarity greater than the similarity threshold.

[0148] Preferably, the method for calculating the similarity by the short text semantic similarity model of the deep neural network described in S233 is as follows:

[0149] S2331: Segment the meta-event text and calculate the word vectors of individual words;

[0150] S2332: Use the central averaging method to stack the word vectors in the text and calculate their average value to obtain the text vector of the meta-event;

[0151] S2333: Calculate the text vector similarity of two meta-events by calculating the cosine of the angle between the text vectors of the two meta-events.

[0152] Preferably, the method for annotating the association relationship between meta-events by the intra-sentence association, the intra-discourse association, the cross-discourse association, and the frequent set mining in S3 is as follows:

[0153] S31: Intra-sentence association: The intra-sentence association is to judge the relationship between two meta-events based on the intra-sentence correlative words, including the relationship of transition, condition, causality, parallelism, and succession;

[0154] S32: Intra-discourse association: The intra-discourse association is to judge the relationship between two meta-events according to the agent of the meta-event and the occurrence time of the meta-event, including the relationship of succession and causality;

[0155] S33: Cross-discourse association: The cross-discourse association is to establish the association between the meta-events between the discourses through the co-reference operation of the meta-events;

[0156] S34: Frequent set mining: The frequent set mining is to mine the frequently co-occurring meta-events in the co-reference library and establish the co-occurrence relationship.

[0157] Preferably, for the meta-events with the annotated association relationship, perform the meta-event relationship extraction based on the relationship extraction model of LSTM+Attention. The specific steps are as follows:

[0158] S311: At the word level, add the positional relationship and extract features using bidirectional LSTM+Attention;

[0159] S312: At the sentence level, after extracting the features, adopt the Attention mechanism for the sentence features;

[0160] S313: Use softmax classification to obtain the meta-event relationship type.

[0161] The structural framework of the relationship extraction model of LSTM+Attention is as Figure 3 shown.

[0162] Preferably, the method of automatic clustering in S4 is:

[0163] S41: Select K meta-events as the centers of the initial classes by using heuristic rules;

[0164] S42: Assign all the meta-events to the nearest classes;

[0165] S43: Recalculate the center of each class;

[0166] S44: Repeat steps S2 and S3 for t times;

[0167] S45: Extract the meta-event descriptions from each class to obtain the clustering result.

[0168] Preferably, the fields of the meta-event extraction library are compatible with the GDELT fields, and the new meta-events are represented by TRS_fields. As shown in Table 1, the TRS_fields include: event description, event source sentence, person in the event sentence, location in the event sentence, organization in the event sentence, event view type, event tense, related topic, position of the event in the sentence, content of the event in the sentence, event field, whether there are similar events, similar event id, title of the document where the event is located, detailed JSON label of the event elements, original word of the event occurrence time, and storage time.

[0169] Table 1

[0170] TRS_EventDescription phrase Event description TRS_EventRootName char None TRS_EventVerb char Event verb TRS_EventVerbDes char Event verb description TRS_EventRelaOrg char Related organization TRS_EventRelaPerson char Related person TRS_EventRelaLoc char Related location TRS_EventRelaCountry char Related country TRS_EventRelaSheng char Related province TRS_EventRelaShi char Related city TRS_EventLevel number Event level TRS_EventPoint char Event perspective TRS_EventTense char Event tense TRS_EventTrigger char Event trigger word TRS_EventValue number Event weight TRS_EventThemes char Related theme TRS_EventSenlD int Position of the sentence where the event is located TRS_EventSentence document Content of the sentence where the event is located TRS_EventFields char Event field TRS_EventKeywords char Event keywords TRS_EventTitle char Title of the document where the event is located TRS_EventSimlD char Similar event id TRS_EventSimTag int Whether there is a similar event TRS_lnserttime date Warehousing time TRS_EventTags phrase Detailed JSON tag of event elements TRS_EventTimeOriWord char Original word of the event occurrence time

[0171] Preferably, the value of MentionNum in the meta-event extraction library is 1, and the value of MentionNum in the meta-event coreference library is the coreference quantity.

[0172] Preferably, the meta-event association library stores the relationships between meta-events. As shown in Table 2, its fields include: ID and basic information of the meta-events with an association relationship, event relationship, event relationship type, storage time, and event relationship weight.

[0173] Table 2

[0174] Field name Type Description Id char Unique ID value eventId1 char Event 1 ID eventId2 char Event 2 ID eventId1_des Phrase Event 1 description eventId2_des Phrase Event 2 description eventId1_actor char Subject of event 1 eventId2_actor char Subject of event 2 eventId1_country char Country of event 1 eventId2_country char Country of event 2 eventIdl_time char Time of event 1 eventId2_time char Time of event 2 eventId1_type char Type of event 1 eventId2_type char Type of event 2 eventId1_theme char Event 1 theme ID eventId2_theme char Event 2 Theme ID eventRela char Event Relationship iRela number Event Relationship Type inserttime date Warehousing Time dRelaValue number Event Relationship Weight

[0175] Preferably, the meta-event topic library stores the basic information of the meta-event topics and the meta-event IDs included therein. As shown in Table 3, its fields include: unique ID value, topic name, topic content, topic picture, topic map location coordinates, and event list.

[0176] Table 3

[0177] Field Name Type Description Zt_id char Unique ID Value Zt_name char Theme Name Zt_content char Theme Content Zt_pic Phrase Theme Picture Zt_pos Phrase Theme Map Location Coordinates eventIds document Event List

[0178] As Figure 4 shown, a method for visual analysis and prediction of meta-events based on a Chinese event library is as follows:

[0179] S1: Meta-event library retrieval: The user retrieves meta-events on the visualization interface of the Chinese event library according to requirements. The retrieval methods include: retrieval by meta-event occurrence time, retrieval by meta-event classification, retrieval by meta-event description, retrieval by meta-event participants, retrieval by meta-event occurrence location, and retrieval by meta-event source;

[0180] S2: Meta-event topic analysis: The GIS map is used to customize the topic in the visual interface of the Chinese event library. The topic analysis can be conducted on a certain meta-event or a certain type of meta-event and displayed in the form of charts. The types of topic analysis include: meta-event heat analysis, hot entity analysis, meta-event type analysis, meta-event context analysis, meta-event public opinion field analysis, event map analysis, and meta-event geographic network analysis. The hot entity analysis displays the entity information involved in the meta-event topic in the form of a word cloud, including people, institutions, and places. The meta-event context analysis displays the occurrence and development of the topic meta-event in the form of a timeline. The meta-event geographic network analysis displays the relationship between meta-event participants through the connection of the geographical locations of the meta-event participants.

[0181] S3: Meta-event prediction analysis: For conflict meta-events, predict the next meta-event development trend of the meta-event; first quantify the conflict meta-event, and then use the prediction model to predict the number and intensity of conflict meta-events, the global conflict index GCI, the local conflict index LCI and the conflict stage QuadClass.

[0182] Preferably, the content of the quantitative representation of the conflict meta-event in S3 includes: meta-event coding and quantity, time and place coding, and meta-event attribute coding.

[0183] Preferably, the specific method of the prediction model in S3 is: using the meta-event location information ActionGeo_CountryCode recorded in the Chinese event library as the location of the conflict meta-event for analysis, using the meta-event attribute coding recorded in the Chinese event library to calculate the conflict index, using the conflict stage QuadClass to divide the meta-event types, using GoldsteinScale to measure the potential impact of the meta-event on regional stability, and using the number of occurrences of a meta-event in the Chinese event library NumMentions to evaluate the importance of a meta-event.

[0184] Preferably, the calculation method of the global conflict index GCI is:

[0185]

[0186] Where: ABS(GS i ) is the absolute value of the GoldsteinScale score of each conflict meta-event; N represents the total number of conflict news meta-events in a country / region; MT i represents the number of mentions of meta-event i; M represents the total number of news meta-events in a country / region.

[0187] Preferably, the calculation method of the local conflict index LCI is:

[0188]

[0189] where: ABS(GS i ) represents the absolute value of the Goldstein Scale score of each conflict meta-event; N represents the total number of conflict news meta-events in a certain country / region; MT i represents the number of mentions of the i-th meta-event.

[0190] Preferably, as Figure 5 shown, the prediction model adopts a fully connected deep method for the model architecture, and the specific steps are as follows:

[0191] S1: Use the meta-events extracted and analyzed from Chinese news or intelligence texts as meta-events, quantitatively represent the meta-events, and perform feature processing;

[0192] S2: Input the meta-events and their vector features into the input layer;

[0193] S3: Use the hidden layer to separate the vector features of the meta-events, predict the meta-events, and evaluate the closeness between the predicted value and the true value to improve the model.

[0194] As Figure 6 shown, a system for constructing a Chinese event library includes:

[0195] Meta-event extraction module: The meta-event extraction module identifies and extracts meta-event elements from Chinese news or intelligence texts, including a meta-event core element extraction unit for extracting meta-event core elements and a meta-event type extraction unit for extracting meta-event types; the meta-event refers to the event extracted and analyzed;

[0196] Meta-event co-reference module: The meta-event co-reference module includes a meta-event time and location element formatting unit for normalizing meta-events, a meta-event subject entity reference resolution unit for aggregating pronouns pointing to the same named entity in different meta-events into a reference chain, and a meta-event merging unit for merging the extracted meta-events;

[0197] Meta-event association module: The meta-event association module identifies the association relationships between meta-events, including: in-sentence association unit, intra-discourse association unit, cross-discourse association unit, frequent set mining unit; the types of association relationships include: condition, causality, transition, succession, parallelism, co-occurrence;

[0198] Meta-event aggregation module: The meta-event aggregation module further aggregates and analyzes the results of the meta-event co-reference module to form meta-event topics, including: a text vector representation unit for representing text vectors using text features, a text vector similarity calculation unit for calculating the similarity of text vectors using the text vectors, and an automatic clustering unit for classifying meta-events into meta-event topics using heuristic rules;

[0199] Chinese Event Library: The meta-event element data processed by the meta-event extraction module is saved in the meta-event extraction library; the data processed by the meta-event coreference module is saved in the meta-event coreference library; the relationship data between two meta-events processed by the meta-event association module is saved in the meta-event association library; the meta-event topic data formed after the meta-event aggregation module processes is saved in the meta-event topic library. The meta-event extraction library, the meta-event coreference library, the meta-event association library, and the meta-event topic library together constitute the Chinese Event Library.

[0200] The knowledge base is the Chinese knowledge graph data publicly disclosed by Baidu.

[0201] QuadClass divides meta-events into 4 major types: 1 represents oral cooperation (Q1), 2 represents substantial cooperation (Q2), 3 represents oral conflict (Q3), and 4 represents substantial conflict (Q4).

[0202] An abnormal increase in the proportion of Q4x and Q4x+1 can indicate a drastic social upheaval; an increase or decrease in the total number of meta-events mentioned in Q1x, Q2x, Q3x, or Q4x can indicate a change in media pressure. Here, x represents the time point.

[0203] GoldsteinScale is a numerical value used to measure the potential impact of meta-events on regional stability, and its value range is between -10 and +10.

[0204] Example 2:

[0205] Taking "a certain event" as an example, relevant Chinese news media data and intelligence data are obtained from the TRS data center, and meta-events are extracted from the data and meta-event aggregation is performed to form "a certain event".

[0206] 1. Meta-event extraction and coreference

[0207] In the "a certain event" topic, a total of 22,622 meta-events are extracted as meta-events using the Chinese Event Library, and the extraction results are as Figure 7 shown.

[0208] 2. Meta-event heat analysis

[0209] In the "a certain event" topic, the number of meta-events occurring in each region is counted and then presented as a heat map on the map. The hottest region of the meta-event is the "meta-event occurrence location"; the number of meta-events occurring daily is counted and the meta-event distribution is presented as a bar chart, and the meta-event heat shows a gradual fade; as Figure 8 shown, the GoldsteinScale values of the meta-events are counted daily and presented as a line chart. The regional stability of the meta-events is basically below 0, indicating that the region has been in a state of conflict.

[0210] 3. Meta-event Hot Entity Analysis

[0211] In the "certain event" topic, count the participant fields of the meta-events and display them in the form of a word cloud.

[0212] 4. Meta-event Context Analysis

[0213] In the "certain event" topic, analyze the key meta-event nodes and display the results through a timeline. For non-key meta-event nodes, extended display can be performed to obtain the overall development context of the meta-events. For example: The meta-event occurred on XX, XX, XX, and the starting meta-event was "XX".

[0214] 5. Meta-event Type Analysis

[0215] In the "certain event" topic, count the meta-event types and display them in the form of a pie chart.

[0216] 6. Meta-event Public Opinion Field Analysis

[0217] In the "certain event" topic, in the public opinion field, count the view types of the speech meta-events and display them in the form of a pie chart, and display the emotional degree in the form of a bar chart. In the public opinion field, neutral accounts for 58.54%, conflict accounts for 26.12%, and cooperation accounts for 15.34%; among them: the emotional value on XX, XX, XX is the highest at 521, and the emotional value on XX, XX, XX is the lowest at -164.

[0218] 7. Meta-event Association Analysis

[0219] In the "certain event" topic, display the relationships between the extracted meta-events on an association graph and analyze the association relationships between the meta-events.

[0220] 8. Meta-event Geographic Network Analysis

[0221] In the "certain event" topic, display the location coordinates of Participant 1 and Participant 2 of the meta-events on a map and connect them to form a meta-event geographic network, including classifying and connecting conflict meta-events and cooperation meta-events.

[0222] Example 3:

[0223] Taking "a certain event" as an example, the main steps for meta-event prediction are as follows:

[0224] Collect relevant conflict meta-event data since XX.

[0225] The retrieval conditions are: (Actor1CountryCode: Country 1 AND Actor2CountryCode: Country 2) OR (Actor2CountryCode: Country 1 AND Actor1CountryCode: Country 2) AND EventRootCode: (X yuan event)

[0226] Nearly 7,000 pieces of data were obtained.

[0227] First, group and statistically analyze the data: including grouping and statistical analysis in two dimensions of time and space: in terms of time, group by year, month, week, etc.; in terms of space, first group by the country Actor1CountryCode, and then continue to group by fields such as Actor1Geo_FeatureID.

[0228] Then, construct the data: the construction form of the data set is: <time, region, statistical value of a certain yuan event (number of yuan events, LCI, QuadClass3, QuadClass4)>, and the number of yuan events of "a certain event", the LCI distribution of "a certain event", and QuadClass3 / QuadClass4 data from XX date to XX date are obtained, and the peak of this event is identified. For example Figures 9 - 11 It can be observed that one month before the outbreak of the X event, there is an obvious upward trend in the number of relevant yuan events, LCI, and the value of QuadClass4. This has an excellent indicative effect for early warning of conflicts.

[0229] Finally, conduct yuan event prediction: use the data of consecutive N months as input, and test with a variety of regression methods to predict the number of yuan events of relevant yuan events in the next month, such as Figure 12 As shown, line 1 represents the answer, line 2 represents linear regression, line 3 represents polynomial regression, line 4 represents decision tree, line 5 represents random forest, line 6 represents ElasticNet, line 7 represents LSTM, and line 8 represents linear fully connected. The results show that the test results of the linear fully connected method are the best.

[0230] It should be noted that the above - described specific embodiments can enable those skilled in the art to understand the present invention - creation more comprehensively, but do not limit the present invention - creation in any way. Therefore, although this specification has described the present invention - creation in detail with reference to the drawings and embodiments, those skilled in the art should understand that the present invention - creation can still be modified or equivalently replaced. In short, all technical solutions and their improvements that do not depart from the spirit and scope of the present invention - creation should be covered by the protection scope of the patent of the present invention - creation.

Claims

1. A method for constructing a Chinese event library, characterized in that, The specific steps are as follows: S1: Meta-event extraction: A meta-event refers to an event that is extracted and analyzed. The meta-event extraction is to identify and extract meta-event elements from Chinese news or intelligence texts to form meta-event element texts. The meta-event elements include the time, location, participating roles of the meta-event, and the changes in related actions or states. The meta-event extraction includes meta-event core element extraction and meta-event type extraction; S2: Meta-event co-reference: The meta-event co-reference is to merge the extracted meta-events based on the common features of the meta-event elements and to normalize the meta-events, including: formatting of meta-event time and location elements, resolution of meta-event subject entity references, and meta-event merging; S3: Meta-event association: The meta-event association is to identify the association relationships between meta-events, including: intra-sentence association, intra-discourse association, cross-discourse association, frequent set mining; The types of the association relationships include: condition, causality, transition, succession, parallelism, co-occurrence; S4: Meta-event aggregation: Further aggregating and analyzing the meta-events processed by S2 meta-event co-reference to form meta-event topics. The aggregation analysis includes: vector representation of texts, calculation of text vector similarity, and automatic clustering to form meta-event topics; S5: Constructing a Chinese event library: Saving the metadata element data after the meta-event extraction in S1 in the meta-event extraction library; Saving the data after the meta-event co-reference in S2 in the meta-event co-reference library; Saving the relationship data between two meta-events after the meta-event association in S3 in the meta-event association library; Saving the meta-event topic data formed after the meta-event aggregation in S4 in the meta-event topic library. The meta-event extraction library, the meta-event co-reference library, the meta-event association library, and the meta-event topic library together constitute the Chinese event library.

2. The method for constructing a Chinese event library according to claim 1, characterized in that The specific method for the meta-event core element extraction in S1 is: S111: Using BIO sequence labeling to label the meta-event core elements, including: labeling the time as B-TIME, I-TIME, labeling the location as B-LOC, I-LOC, labeling the subject as B-A0, I-A0, and labeling the object as B-A1, I-A1; S112: Using the BERT+BiLSTM+CRF model to extract the meta-event core elements labeled by the BIO sequence in S111 and output the meta-event core element results; S1121: First, using the BERT model to perform word embedding on the text tokens in the meta-event core elements labeled by the BIO sequence; S1122: Using the memory network model BiLSTM to perform further sequence prediction on the meta-event core elements after word embedding in S1121. The sequence prediction method includes: performing sequence prediction by globally modeling the features between trigger words, between meta-event core elements, and between trigger words and meta-event core elements; S1123: Use the conditional random field (CRF) to calculate the feature scores of the core elements of the meta-event at the current position, the feature scores of label transitions, and the sum of the feature scores of the current position and label transitions in the sequence prediction result described in S1122. Then, select the core element of the meta-event with the highest sum of feature scores in the entire sequence prediction result as the result of the core element of the meta-event; S1124: Output the result of the core element of the meta-event.

3. A method for constructing a Chinese event library according to claim 1, characterized in that, The extraction of the meta-event type described in S1 is based on the BERT neural network model. The specific steps are as follows: S121: Tokenize, perform part-of-speech tagging, syntactic analysis, extract the core elements of the meta-event, and extract the central word of the meta-event for the meta-event sentence of the meta-event, and form the corresponding text; S122: Respectively perform vector representation on the text formed by the tokenization result in S121, the syntactic structure obtained after syntactic analysis, the extracted core elements of the meta-event, and the extracted central word of the meta-event; S123: Perform deep learning on the vectors obtained in S122 based on the BERT neural network model to obtain the vector features of the meta-event; S124: Correlate the major categories in the TRS meta-event classification standard with the vector features of the meta-event obtained in S123 to achieve the initial classification of the meta-event; S125: After defining the major category to which the meta-event belongs, further correlate the vector features under the major category to which the meta-event belongs with the minor categories of the TRS meta-event classification standard to achieve the secondary classification of the meta-event.

4. A method for constructing a Chinese event library according to claim 3, characterized in that, The major categories in the TRS meta-event classification standard include speech meta-events, accident disasters, public health meta-events, political meta-events, judicial meta-events, military meta-events, diplomatic meta-events, work meta-events, life meta-events, sports meta-events, economic meta-events, and security meta-events.

5. A method for constructing a Chinese event library according to claim 1, characterized in that, The specific method for normalizing the meta-event in S2 is as follows: S211: Standardize the time elements extracted from the meta-event and normalize them to the standard date format: YYYY.MM.DD; S212: Normalize the location elements extracted from the meta-event and normalize them to location coordinates: [latitude, longitude], where the north latitude is positive, the south latitude is negative, the east longitude is positive, and the west longitude is negative.

6. A method for constructing a Chinese event library according to claim 1, characterized in that, The anaphora resolution of the meta-event subject entity in S2 utilizes the BERT-ENE model. The specific steps are as follows: S221: Construct an entity name dictionary using the entity names and alias information of the entities in the knowledge base; S222: Through the entity description text of the knowledge base, use the BERT pre-trained model to select the vector output at the CLS position as the entity name embedding; S223: Match the meta-event with the entity name dictionary to obtain candidate entities in the short text of different meta-events; S224: Aggregate all pronouns in the candidate entities in the short text that point to the same named entity into a reference chain through the BERT-ENE model.

7. A method for constructing a Chinese event library according to claim 1, characterized in that, The meta-event merging in S2 is to merge the same or similar meta-events by similarity comparison and merge the relevant meta-event element texts. The specific method is as follows: S231: Sort according to the occurrence time of the meta-event and perform similarity judgment on the meta-events that occur on the same day; S232: Confirm the subject of the meta-event, and perform similarity judgment on meta-events with the same subject; S233: Calculate the similarity of meta-event types and the similarity of meta-event description sentences based on the short text semantic similarity model of the deep neural network; S234: Sort the sentences according to the similarity, and merge the meta-events with a similarity greater than the similarity threshold.

8. A method for constructing a Chinese event library according to claim 7, characterized in that The method for calculating similarity by the short text semantic similarity model of the deep neural network described in S233 is as follows: S2331: Segment the meta-event text and calculate the word vectors of individual words; S2332: Use the central averaging method to stack the word vectors in the text and calculate their average value to obtain the text vector of the meta-event; S2333: Calculate the similarity of the text vectors of two meta-events by calculating the cosine of the angle between the text vectors of the two meta-events.

9. A method for constructing a Chinese event library according to claim 1, characterized in that, The method for annotating the association relationship between meta-events by the intra-sentence association, intra-discourse association, cross-discourse association, and frequent set mining described in S3 is as follows: S31: Intra-sentence association: The intra-sentence association is based on the intra-sentence correlative words to judge the relationship between two meta-events, including the relationship of turning, condition, causality, parallelism, and succession; S32: Intra-discourse association: The intra-discourse association is to judge the relationship between two meta-events according to the agent of the meta-event and the occurrence time of the meta-event, including the relationship of succession and causality; S33: Cross-discourse association: The cross-discourse association is to establish the association between meta-events between discourses through the co-reference operation of meta-events; S34: Frequent set mining: The frequent set mining is to mine the meta-events that frequently co-occur in the co-reference library and establish the co-occurrence relationship.

10. A method for constructing a Chinese event library according to claim 9, characterized in that, For the meta-events with the annotated association relationship, perform meta-event relationship extraction based on the relationship extraction model of LSTM+Attention. The specific steps are as follows: S311: At the word level, add the positional relationship and extract features with bidirectional LSTM+Attention; S312: At the sentence level, after extracting features, adopt the Attention mechanism for the sentence features; S313: Use softmax classification to obtain the meta-event relationship type.

11. A method for constructing a Chinese event library according to claim 1, characterized in that The method for automatic clustering described in S4 is as follows: S41: Select K meta-events as the centers of the initial classes by using heuristic rules; S42: Assign all meta-events to the nearest classes; S43: Recalculate the center of each class; S44: Repeat steps S2 and S3 for t times; S45: Extract the meta-event descriptions from each class to obtain the clustering result.

12. A method for constructing a Chinese event library according to claim 1 or 2, characterized in that, The fields of the meta-event extraction library are compatible with the GDELT fields, and at the same time, the TRS_ fields are used to represent the newly added meta-events. The TRS_ fields include: event description, event source sentence, person in the event sentence, location in the event sentence, organization in the event sentence, event view type, event tense, related topic, position of the event in the sentence, content of the event in the sentence, event field, whether there are similar events, similar event id, title of the document where the event is located, detailed JSON label of the event elements, original word of the event occurrence time, and storage time.

13. A method for constructing a Chinese event library according to claim 1, characterized in that Set the MentionNum value in the meta-event extraction library to 1, and the MentionNum value in the meta-event co-reference library to the co-reference quantity.

14. A method for constructing a Chinese event library according to claim 1, characterized in that, The meta-event association library stores the relationships between meta-events, and its fields include: meta-event IDs and basic information of associated relationships, event relationships, event relationship types, storage time, and event relationship weights.

15. A method for constructing a Chinese event library according to claim 1, characterized in that, The meta-event topic library stores basic information of meta-event topics and meta-event IDs contained therein, and its fields include: ID unique value, topic name, topic content, topic picture, topic map location coordinates, and event list.

16. A method for visual analysis and prediction of meta-events of a method for constructing a Chinese event library according to any one of claims 1-15, characterized in that, The specific steps are as follows: S1: Meta-event database search: Users search for meta-events in the visual interface of the Chinese event database according to their needs. The search methods include: search by meta-event occurrence time, search by meta-event classification, search by meta-event description, search by meta-event participant, search by meta-event occurrence location, and search by meta-event source; S2: Meta-event topic analysis: The GIS map is used to customize the topic in the visual interface of the Chinese event library. Meta-events can be extracted for a certain meta-event or a certain type of meta-event and topic analysis can be performed, and the topic analysis can be displayed in the form of charts. The types of topic analysis include: meta-event heat analysis, hot entity analysis, meta-event type analysis, meta-event context analysis, meta-event public opinion field analysis, event map analysis, and meta-event geographic network analysis; the hot entity analysis displays the entity information involved in the meta-event topic in the form of a word cloud, including people, institutions, and places; the meta-event context analysis displays the occurrence and development of topic meta-events in the form of a timeline; the meta-event geographic network analysis displays the relationship between meta-event participants through the connection of the geographical locations of meta-event participants; S3: Meta-event prediction analysis: For conflict meta-events, predict the next development trend of the meta-events; first quantify the conflict meta-events, and then use the prediction model to predict the number and intensity of conflict meta-events, the global conflict index GCI, the local conflict index LCI and the conflict stage QuadClass.

17. A method for visual analysis and prediction of meta-events according to claim 16, characterized in that The content of the quantitative representation of the conflict meta-event in S3 includes: meta-event coding and quantity, time and place coding, and meta-event attribute coding.

18. A method for visual analysis and prediction of meta-events according to claim 16, characterized in that, The specific method of the prediction model described in S3 is: using the meta-event location information ActionGeo_CountryCode recorded in the Chinese event library as the location of the conflict meta-event for analysis, using the meta-event attribute coding recorded in the Chinese event library to calculate the conflict index, using the conflict stage QuadClass to divide the meta-event types, using GoldsteinScale to measure the potential impact of the meta-event on regional stability, and using the number of occurrences of a meta-event in the Chinese meta-event library NumMentions to evaluate the importance of a meta-event.

19. A method for visual analysis and prediction of meta-events according to claim 16, characterized in that, The calculation method of the Global Conflict Index GCI is: Where: ABS(GS i ) is the absolute value of the Goldstein Scale score of each conflict meta-event; N represents the total number of conflict news meta-events in a certain country / region; MT i represents the number of mentions of the i-th event; M represents the total number of news meta-events in a country / region.

20. A method for visual analysis and prediction of meta-events according to claim 16, characterized in that, The calculation method of the local conflict index LCI is: Where: ABS(GS i ) represents the absolute value of the Goldstein Scale score of each conflict meta-event; N represents the total number of conflict news meta-events in a country / region; MT i Represents the number of mentions of the i-th event.

21. A method for visual analysis and prediction of meta-events according to claim 16, characterized in that, The prediction model adopts a deep fully connected model architecture, and the specific steps are as follows: S1: Extract and analyze meta-events from Chinese news or intelligence texts as meta-events, quantify meta-events and perform feature processing; S2: Input the primitive event and its vector features at the input layer; S3: Use the hidden layer to separate the vector features of the primitive event, predict the primitive event, and evaluate the closeness between the predicted value and the true value to improve the model.

22. A system for constructing a Chinese event library, characterized in that, Including: Primitive event extraction module: The primitive event extraction module identifies and extracts primitive event elements from Chinese news or intelligence texts, including a primitive event core element extraction unit for extracting primitive event core elements and a primitive event type extraction unit for extracting primitive event types; The primitive event refers to the event that is extracted and analyzed. Primitive event co-reference module: The primitive event co-reference module includes a primitive event time and location element formatting unit for normalizing the primitive event, a primitive event subject entity co-reference resolution unit for aggregating pronouns pointing to the same named entity in different primitive events into a reference chain, and a primitive event merging unit for merging the extracted primitive events. Primitive event association module: The primitive event association module identifies the association relationships between primitive events, including: an intra-sentence association unit, an intra-discourse association unit, an inter-discourse association unit, and a frequent set mining unit; The types of association relationships include: condition, causality, transition, succession, parallelism, co-occurrence. Primitive event aggregation module: The primitive event aggregation module further aggregates and analyzes the results of the primitive event co-reference module to form primitive event topics, including: a text vector representation unit for representing the text vector using text features, a text vector similarity calculation unit for calculating the similarity of the text vector using the text vector, and an automatic clustering unit for classifying the primitive events into primitive event topics using heuristic rules. Chinese event library: Save the primitive event element data processed by the primitive event extraction module in the primitive event extraction library; Save the data processed by the primitive event co-reference module in the primitive event co-reference library; Save the relationship data between two primitive events processed by the primitive event association module in the primitive event association library; Save the primitive event topic data formed by the primitive event aggregation module in the primitive event topic library. The primitive event extraction library, the primitive event co-reference library, the primitive event association library, and the primitive event topic library together constitute the Chinese event library.

Citation Information

Patent Citations

  • Chinese syntax rule-based event extraction method and system

    CN106959944A

  • Event graph construction method based on social media

    CN108763333A