Border event aggregation classification algorithm based on space-time correlation analysis
By constructing event classification and named entity models, calculating association rules and building a rational map, the problem of border event aggregation and classification is solved, effective aggregation and classification of events is achieved, and computing efficiency and interpretability of rules are improved.
Patent Information
- Application Number
- CN202510059729.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
In the border field, it is difficult for the existing technology to effectively use spatiotemporal correlation analysis to aggregate and categorize border events, resulting in the failure to fully explore the event correlation relationship.
A border event aggregation classification algorithm based on spatiotemporal correlation analysis is proposed, including building event classification and named entity models, calculating correlation rules between border events, building a matter-of-map, and classifying and sorting unclassified events through density clustering and matter-of-map.
It realizes effective aggregation and classification of border events, improves the interpretability and use of event association rules, simplifies the frequent item set calculation process, and improves the computing efficiency.
Smart Images

Figure CN119989044A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and more specifically to a border event aggregation classification algorithm based on spatiotemporal correlation analysis. Background Art
[0002] In the border area, using existing border emergency data to conduct correlation analysis between events, explore the sequence and causal relationship between events, and establish a rule model for border events, which will be of great help in revealing the evolution law and logic of events. Traditional correlation mining algorithms are used in social media, medical and public transportation, but are rarely used in the border area.
[0003] Therefore, how to provide a border event aggregation classification algorithm based on spatiotemporal correlation analysis is an urgent problem that technical personnel in this field need to solve. Summary of the invention
[0004] In view of this, the present invention provides a border event aggregation classification algorithm based on spatiotemporal correlation analysis.
[0005] In order to achieve the above object, the present invention adopts the following technical solution:
[0006] An algorithm for clustering and classifying border events based on spatiotemporal correlation analysis, including:
[0007] Build event classification and named entity models;
[0008] Calculate association rules between border events based on historical event thematic collection data;
[0009] Construct a reasoning graph based on association rules;
[0010] Based on event classification and named entity model, border events in the unclassified event set are classified and named entity recognized;
[0011] Based on event categories and named entities, border events in the unclassified event set are classified and sorted through the event graph to obtain the border event aggregation results.
[0012] Preferably, constructing an event classification and a named entity model specifically includes:
[0013] Step 101: pre-processing border events to obtain a text dataset that only includes Chinese characters;
[0014] Step 102: Construct an event classification and named entity model, including a Bert structure, a TextCNN model, a bi-lstm layer, and a crf layer. The specific processing process is as follows:
[0015] The text dataset extracts word features through the Bert structure to obtain text feature vectors;
[0016] The text feature vector is classified into text event categories through the TextCNN model, and the probability corresponding to each category is output. The text feature vector is sequentially passed through the bi-lstm layer and the crf layer for named entity recognition to obtain the entity identifier corresponding to each word;
[0017] Step 103: Train event classification and named entity models based on historical event classification data and historical event entity identification data.
[0018] Preferably, the TextCNN model includes a convolutional layer, a maximum pooling layer, a fully connected layer, and a softmax classification layer;
[0019] After the text feature vector passes through three convolution kernels and three maximum pooling layers in parallel, the features are connected in parallel and classified through a full connection and a softmax classification layer to output the probability corresponding to each category.
[0020] Preferably, the association rules between border events are calculated based on the historical event subject collection data, specifically including:
[0021] Step 201: Assemble the historical event themes T = {T1, T2, ..., T n}, count the unique items and common items in each historical event theme, n represents the total number of historical event themes;
[0022] Step 202: For a single historical event topic set T j ={t1, t2, ..., t m}, t i Indicates T j An event collection in e e is the last event in the event collection, then {e e} is T j Frequent 1-itemset L k , k = 1;
[0023] Step 203: Traverse T j Each event is collected and the statistics include e e Frequent 2-itemsets, check each frequent 2-itemsets and remove the items in the frequent 1-itemsets, determine whether the newly added event item is a unique item or a common item, calculate the support of the new event frequent 2-itemsets, delete the items that are less than the minimum support threshold, and then all the remaining items form the frequent 2-itemsets L2;
[0024] Step 204: Traverse T jAfter each event is grouped, the last k items in the frequent k-1-item set are counted as candidate item sets. After checking each candidate k-item set and removing the items in the frequent k-1-item set, it is determined whether the newly added event item is a unique item or a common item. The support of the newly added event k-item set is calculated, and the items less than the minimum support threshold are deleted to obtain the frequent k-item set L. k ;
[0025] Step 205: let k=k+1, and repeat the above step 204 until no new candidate item set is generated;
[0026] Step 206: L1 to L k Take the union to obtain all frequent item sets L;
[0027] Step 207: Assemble all historical events into a topic set T = {T1, T2, ..., T n}Perform steps 202 to 206 to obtain a frequent item set for each historical event topic.
[0028] Preferably, the judgment formula for special items and common items is:
[0029]
[0030] in, Represents event S about the historical event topic set T j Probability value, Represents event S in the historical event subject set T j The number of occurrences in , if θ is the preset value, and the S event is not T j The end node of , then the event S is considered to be the historical event subject set T j The support is a unique item, otherwise it is a common item; the support calculation formula is:
[0031]
[0032] in, It represents item sets A and B about the historical event topic set T j The support of Indicates that item sets AB are simultaneously in the historical event topic set T j The number of times it appears in Indicates T j The number of event clusters that include item set A or item set B, Indicates T j The number of all events in the collection.
[0033] Preferably, constructing a reasoning graph based on association rules specifically includes:
[0034] Step 301: Assume that the historical event topic set T j The frequent itemset is L = {l1, l2, ..., l g}, there are g frequent item sets in total, traverse them, and take l1 as a link in the event graph, l q Represents the qth frequent set, let q = 2;
[0035] Step 302: Select frequent set l q , if frequent set l q There are p nodes, and the node set is Traverse the node set in reverse order and take the end node of the event graph as the frequent set l q The current node of
[0036] Step 303: q Move the current node forward to get As the current node of the frequent item, obtain the predecessor node set of the current node of the event graph, and determine the current node of the frequent item Whether it is in the set of predecessor nodes of the current node in the event graph;
[0037] If the frequent item current node If it exists in the predecessor node set of the current node in the event graph, then set the current node in the event graph to the predecessor node set and The same node is repeated in step 303 to continue the iteration. If it does not exist, a new link is created in the current node of the event graph. The new link is connected to l q same;
[0038] Step 304: Let q = q + 1, and repeat steps 302 to 303 until all frequent item sets L are traversed to obtain the updated event graph M. j .
[0039] Preferably, it also includes:
[0040] Step 305: Traverse the event graph M in sequence j For each node of the event graph M j A node m in h With the rear drive node m h+1 If the types are the same, delete the subsequent node m. h+1 , m h The following node becomes m h+2 , and m h The node adds a loop to itself, indicating that m h Nodes may recur.
[0041] Preferably, based on event categories and named entities, border events in the unclassified event set are classified and sorted through the event graph to obtain border event aggregation results, specifically including:
[0042] Step 401: Set a distance threshold ε, which consists of three parts: geographic location, time interval δ, and subject unit. Determine whether the border events are density-reachable based on the named entity. If two border events have the same geographic location, the same subject unit, and the time interval between occurrences is less than δ, then the density is reachable.
[0043] Step 402: Create a new cluster, select an unlabeled event y from the unclassified event set, traverse other events, add y and events with density reachable to y to the cluster, and label the y event;
[0044] Step 403: select an unmarked event u from the cluster in turn, traverse other events in the unclassified event set that are not assigned to the cluster, add events that are density-reachable to u to the cluster, and mark the u event;
[0045] Step 404: Repeat step 403 until all points in the cluster are marked;
[0046] Step 405: Repeat steps 402 to 404 until all events in the unclassified event set are marked and have their own clusters, and obtain the event set cluster R = {R1, R2, ..., R c}, through the event graph, the events in the event cluster are classified and collected, so that border events can be aggregated according to the event theme.
[0047] Preferably, the events in the event cluster are classified and collected using the event graph, specifically including:
[0048] Step 4051: Select an event cluster, sort the events in it in chronological order, traverse the events in it, and select one of the terminal events;
[0049] Step 4052: Select a corresponding event graph according to the event category attribute of the terminal event, and select a link of the event graph; traverse the event cluster, and fill the events therein into the nodes of the corresponding category on the link in sequence according to the event category of each event;
[0050] Step 4053: loop step 4052 to fill all links of the event graph with events;
[0051] Step 4054: Score each link in turn, select the link with the highest score and take all events on the link as a set, and divide the topics according to the end nodes of the set to complete the aggregation of border events, and exclude all the above events from the event cluster. If there are other end nodes in the event cluster, repeat steps 4052 to 4054 until there are no end events in the event cluster.
[0052] Preferably, the scoring rules for each link are:
[0053] F=v*v / z
[0054] Among them, F represents the score, z represents the number of nodes in a link, and v represents the filled node.
[0055] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a border event aggregation classification algorithm based on spatiotemporal correlation analysis, which has the following advantages:
[0056] (1) A creative method for border event classification and entity recognition is proposed, which can simultaneously complete the classification task and entity recognition task through a multi-task approach;
[0057] (2) In terms of mining association rules for border events, a method for calculating the support of common items and special items of events is proposed. In the process of calculating frequent item sets, the uniqueness of special events is taken into account. The frequent item set analysis method of events is optimized according to the characteristics of border events, the frequent item calculation process is simplified, and the calculation efficiency is improved;
[0058] (3) Introducing the construction of the event graph into the border event correlation analysis process makes the event association rule modeling more explainable and usable.
[0059] (4) By integrating spatiotemporal information, event categories, and text features, events are aggregated into clusters based on the idea of density clustering, and the event graph is used to realize the theme collection of events. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0061] Figure 1 A flow chart of a border event aggregation classification algorithm based on spatiotemporal correlation analysis provided by the present invention.
[0062] Figure 2This is the event classification and named entity model structure diagram provided by the present invention.
[0063] Figure 3 This is a flow chart of the association rule algorithm provided by the present invention.
[0064] Figure 4 A flow chart for generating a logical diagram provided by the present invention. DETAILED DESCRIPTION
[0065] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0066] The embodiment of the present invention discloses a border event aggregation classification algorithm based on spatiotemporal correlation analysis, such as Figure 1 As shown, including:
[0067] Build event classification and named entity models;
[0068] Calculate association rules between border events based on historical event thematic collection data;
[0069] Construct a reasoning graph based on association rules;
[0070] Based on event classification and named entity model, border events in the unclassified event set are classified and named entity recognized;
[0071] Based on event categories and named entities, border events in the unclassified event set are classified and sorted through the event graph to obtain the border event aggregation results.
[0072] Aiming at the characteristics of border event data, the present invention proposes an association relationship mining algorithm based on border events, obtains the cause-and-effect graph of border events, and uses it for the association aggregation of events. It can effectively extract border event rules and group events according to themes.
[0073] The specific implementation process of each step of the present invention is described in detail below
[0074] (1) Constructing event classification and named entity model
[0075] In border management, various events may occur, and there may be certain associations and accompaniment between events. These related events can be grouped together to represent a certain type of specific theme. By analyzing the sequence and causal relationship between emergencies and other events related to emergencies, mining association rules, and establishing a reasoning map, it has certain reference and practical value for border control and prevention. In order to mine the association rules of events, it is necessary to conduct category analysis on the events.
[0076] For an event, a text recognition algorithm is needed to obtain the time, location, participants, etc. of the event, and to be able to classify the event. Here, an event text recognition algorithm based on deep learning is proposed. While identifying the key entities in the event, it also classifies the event and provides category information and entity information for subsequent association rule mining. The model is constructed as follows:
[0077] Step 101: Preprocess the border event text. Use a semi-manual method to convert the numbers in the document into Chinese characters, then use the JieBa word segmentation tool to segment the data and obtain its part of speech, and finally use the stop word list to filter the stop words in the text to obtain a text dataset containing only Chinese characters. The border event is the text part of the historical event classification data and entity identification data.
[0078] Step 102: Construct an event classification and named entity model. The structure of the model is as follows: Figure 2 As shown:
[0079] First, we extract word features through Bert to get the context vector representation of each word. Then we divide it into two branches. One is to classify text events with the help of text convolutional neural network (TextCNN), and the other is to use bidirectional long short-term memory neural network bi-lstm and conditional random field crf algorithm to perform named entity recognition, mainly to identify the time, place and participants of the event.
[0080] The TextCNN model mainly consists of four layers: convolutional layer, maximum pooling layer, fully connected layer and the final softmax classification layer.
[0081] The convolution layer includes three convolution kernels with sizes of (3,256), (4,256), and (5,256). After the border event text is represented by Bert through the bidirectional encoder, a text feature vector can be obtained. After the feature vector passes through the three convolution kernels and the maximum pooling layer in parallel, the features are connected in parallel and passed through a fully connected softmax layer for category classification, and the probability corresponding to each category is output.
[0082] The named entity recognition circuit is completed through a bi-lstm layer and a crf (Conditional Random Fields) layer. The bi-lstm is based on an LSTM model with 128 internal units. After the text passes through the Bert feature extraction model, the text feature vector is further extracted through the bi-lstm and input into a crf conditional random field model to obtain the corresponding entity identifier of each word. The BIO entity identifier method is used here, B represents the beginning of the entity noun, I represents the middle part of the entity noun, and o represents the non-entity vocabulary. The crf will get the corresponding identifier of each word.
[0083] The model is an end-to-end, multi-task model. Compared with separate entity recognition and text classification, it uses a common feature extraction layer to reduce the overall computational complexity. The multi-task mode also enables the model to have better feature expression capabilities.
[0084] Step 103: Use historical event classification data and historical event entity identification data to train the model. The Bert layer uses the Bert-base-chinese model released by Google; freeze the parameters of the Bert layer during training, and only train the subsequent TextCNN feature extraction layer. Stop training when the model loss converges. Then use historical event entity identification data to train the model. For example, freeze the Bert layer parameters during training, and only train the BiLSTM-crf layer. Stop training when the model loss value converges, and you can get the model parameters.
[0085] (2) Calculate association rules between border events based on historical event thematic collection data
[0086] Event collections are usually composed of multiple related events, which contain certain temporal associations. Each type of event collection usually ends with a fixed event, and this type of event will only be the end node of this type of event collection and will not appear elsewhere. Here, an event association mining algorithm is provided to mine the association rules of events.
[0087] In the border event data set, for n topics, each type of topic has a sup value of AB, because for some events, the probability of their occurrence in a specific topic i will be very high, but the number of occurrences in other events will be small. It is necessary to calculate the probability distribution of each type of event A in all topics separately to determine whether the event is a topic-specific event. The formula is as follows:
[0088]
[0089] in, Represents event S about the historical event topic set Tj Probability value, Represents event S in the historical event subject set T j The number of occurrences in , if θ is a preset value, generally not greater than 0.7, and the S event is not T j The end node of the event collection, then the event S is considered to be about the topic T i It is a unique event, called a unique item, otherwise it is called a common item.
[0090] In the correlation calculation process of border events, it is assumed that the number of items in AB is The calculation method is as follows:
[0091]
[0092] in, It represents the topic set T of item sets A and B about historical events j The support of Indicates that AB item sets are simultaneously grouped in the historical event theme T j The number of times it appears in Indicates T j The number of event clusters that include item set A or item set B, Indicates T j The number of all events in the collection.
[0093] like Figure 3 As shown, the calculation steps of the association rule model algorithm are as follows:
[0094] Step 201: Assemble the historical event themes T = {T1, T2, ..., T n}, and count the unique items and common items in each topic according to formula (1).
[0095] Step 202: For a single historical event topic set T j ={t1, t2, ..., t m}, t i Indicates T j An event collection in e e is the last event in the event collection, then {e e} is T j Frequent 1-itemset L k , k=1.
[0096] Step 203: Traverse T j Each event is collected and the statistics include e eFrequent 2-item sets, after checking each frequent 2-item set and removing the items in the frequent 1-item set, determine whether the newly added event item is a unique item or a common item. In the support calculation formula, the calculation methods of unique items and common itemsets are different, and the support of the frequent 2-item set is calculated according to formula (2). Items with less than the minimum support threshold are deleted, and then all the remaining items form the frequent 2-item set L2. For example, if a 2-item set is {e w , e e}, the item set removes {e e}, then {e w}, check {e w} is a unique item, and calculate {e w , e e}'s sup support At this time, the frequent k-item set L k In the example, k=2, let k=k+1=3, and then calculate the frequent 3-itemsets.
[0097] Step 204: Similar to step 203, traverse T j For each event set, count the last k items in the frequent k-1-item set as the candidate item set. After checking each candidate k-item set and removing the items in the frequent k-1-item set, determine whether the newly added event item is a unique item or a common item. In the support calculation formula, the calculation methods of unique items and common item sets are different. The support of the k-item set is calculated according to formula (2), and the items less than the minimum support threshold are deleted to obtain the frequent k-item set L. k .
[0098] Step 205: Let k = k + 1, and repeat the above step 204 until no new candidate item set is generated.
[0099] Step 206: L1 to L k Take the union to get all frequent itemsets L.
[0100] Step 207: Assemble all event topics into a set T = {T1, T2, ..., T n}Perform the above calculation to obtain the frequent item set of each historical event topic.
[0101] (3) Constructing a causal graph based on association rules
[0102] When the subject cluster T is obtained j When we have a frequent item set, we need to combine it into a factual graph M j , used as a rule model for the subject collection.
[0103] like Figure 4 As shown in the figure, the algorithm for converting frequent item sets into event graphs is as follows:
[0104] Step 301: Assume that the historical event topic set T j The frequent itemset is L = {l1, l2, ..., l g}, there are g frequent item sets in total, traverse them, and take l1 as a link in the event graph, l q Represents the qth frequent set, let q=2.
[0105] Step 302: Select frequent set l q , if frequent set l q There are p nodes, and the node set is Traverse the node set in reverse order. First, compare With M j The terminal nodes of the event graph are the same, and the terminal nodes of the event graph are taken as the frequent set l q The current node.
[0106] Step 303: q Move the current node forward to get As the current node of the frequent item, obtain the predecessor node set of the current node of the event graph. If If it exists in the predecessor node set of the current node in the event graph, then set the current node in the event graph to the predecessor node set and If the same node does not exist, a new link is created at the current node of the event graph, and the new link is connected to l q same.
[0107] Step 304: Let q = q + 1, and repeat steps 302 to 303 until all frequent event sets L are traversed. At this time, the event graph M can be obtained. j .
[0108] Step 305: Traverse the event graph M in sequence j For each node of the event graph M j A node m in h With the rear drive node m h+1 If the types are the same, delete m h+1 Node, m h The following node becomes m h+2 , and m h The node adds a loop to itself, indicating that m h Nodes may recur.
[0109] (4) Event type and named entity recognition
[0110] For unclassified event sets, it is necessary to identify the event types and named entities to provide support for the subsequent construction of border event aggregation.
[0111] By using event classification and named entity models to predict the content of the event to be identified, we can obtain the category information of the event. At the same time, we can obtain the key entities of the event, including the time of occurrence, geographical location and main unit of the event, the spatiotemporal features of the event, and the text features.
[0112] (5) Border event aggregation method based on spatiotemporal correlation analysis
[0113] For unclassified event sets, we can aggregate the events into multiple event sets by combining the event's spatiotemporal features, text features, and association rule model (event graph). First, we refer to the idea of density clustering and merge the potentially related events together to form event clusters. The specific steps are as follows:
[0114] Step 401: Set the distance threshold ε, which consists of three parts: geographic location, time interval δ, and subject unit. If two events have the same geographic location, the same subject unit, and the time interval between occurrences is less than δ, then their density is considered to be reachable; use classification and named entity recognition for unclassified events in turn to obtain each event category and each event's named entity (including event location, event time, and event subject unit), add event categories and event named entities to event attributes, and subsequent calculations of density reachability are based on comparing the named entity attributes of the two events; event category attributes are mainly used to compare event types with node types within the event graph.
[0115] Step 402: Create a new cluster, select an unlabeled event y from the unclassified event set, traverse other events, add y and events with density reachable to y to the cluster, and label the y event;
[0116] Step 403: select an unmarked event u from the cluster in turn, traverse other events in the unclassified event set that are not assigned to the cluster, add events that are density-reachable to u to the cluster, and mark the u event;
[0117] Step 404: Repeat step 403 until all points in the cluster are marked;
[0118] Step 405: Repeat steps 402 to 404 until all events in the unclassified event set are marked and have their own clusters, and obtain the event set cluster R = {R1, R2, ..., R c}, by using the event graph to classify and aggregate the events in the event cluster, the effect of aggregating border events according to event themes can be achieved. The specific algorithm is as follows:
[0119] Step 4051: Select an event cluster, sort the events therein in chronological order, traverse the events therein, and select the terminal event therein, so as to clarify the subject event set contained in the event cluster.
[0120] Step 4052: Select the corresponding event graph according to the event category attribute of the terminal event, and select a link of the event graph; traverse the event cluster, and fill the events therein into the nodes of the corresponding category on the link in sequence according to the event category of each event.
[0121] Step 4053: Loop step 4052 to fill all links of the event graph with events.
[0122] Step 4054: Score each link in turn. The scoring rules are as follows: if a link has z nodes, of which v nodes are filled, its score is:
[0123] F=v*v / z (3)
[0124] The link with the highest score is selected and all events on the link are taken as a set. The topics are divided according to the terminal nodes of the set to complete the aggregation of border events and exclude these events from the event cluster. If there are other terminal nodes in the event cluster, steps 4052 to 4054 are repeated until there are no terminal events in the event cluster.
[0125] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0126] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A border event aggregation classification algorithm based on spatiotemporal correlation analysis, characterized in that: include: Build event classification and named entity models; Calculate association rules between border events based on historical event thematic collection data; Construct a reasoning graph based on association rules; Based on event classification and named entity model, border events in the unclassified event set are classified and named entity recognized; Based on event categories and named entities, border events in the unclassified event set are classified and sorted through the event graph to obtain the border event aggregation results.
2. According to claim 1, a border event aggregation classification algorithm based on spatiotemporal correlation analysis is characterized in that: Build event classification and named entity models, including: Step 101: pre-processing border events to obtain a text dataset that only includes Chinese characters; Step 102: Construct an event classification and named entity model, including a Bert structure, a TextCNN model, a bi-lstm layer, and a crf layer. The specific processing process is as follows: The text dataset extracts word features through the Bert structure to obtain text feature vectors; The text feature vector is classified into text event categories through the TextCNN model, and the probability corresponding to each category is output. The text feature vector is sequentially passed through the bi-lstm layer and the crf layer for named entity recognition to obtain the entity identifier corresponding to each word; Step 103: Train event classification and named entity models based on historical event classification data and historical event entity identification data.
3. The border event aggregation classification algorithm based on spatiotemporal correlation analysis according to claim 2 is characterized in that: The TextCNN model includes convolutional layers, maximum pooling layers, fully connected layers, and softmax classification layers; After the text feature vector passes through three convolution kernels and three maximum pooling layers in parallel, the features are connected in parallel and classified through a full connection and a softmax classification layer to output the probability corresponding to each category.
4. The border event aggregation classification algorithm based on spatiotemporal correlation analysis according to claim 1 is characterized in that: The association rules between border events are calculated based on the historical event subject collection data, including: Step 201: Assemble the historical event themes T = {T1, T2, ..., T n }, count the unique items and common items in each historical event theme, n represents the total number of historical event themes; Step 202: For a single historical event topic set T j ={t1, t2, ..., t m }, t i Indicates T j An event collection in e e is the last event in the event collection, then {e e } is T j Frequent 1-itemset L k , k = 1; Step 203: Traverse T j Each event is collected and the statistics include e e Frequent 2-itemsets, check each frequent 2-itemsets and remove the items in the frequent 1-itemsets, determine whether the newly added event item is a unique item or a common item, calculate the support of the new event frequent 2-itemsets, delete the items that are less than the minimum support threshold, and then all the remaining items form the frequent 2-itemsets L2; Step 204: Traverse T j After each event is grouped, the last k items in the frequent k-1-item set are counted as candidate item sets. After checking each candidate k-item set and removing the items in the frequent k-1-item set, it is determined whether the newly added event item is a unique item or a common item. The support of the newly added event k-item set is calculated, and the items less than the minimum support threshold are deleted to obtain the frequent k-item set L. k ; Step 205: let k=k+1, and repeat the above step 204 until no new candidate item set is generated; Step 206: L1 to L k Take the union to obtain all frequent item sets L; Step 207: Assemble all historical events into a topic set T = {T1, T2, ..., T n }Perform steps 202 to 206 to obtain a frequent item set for each historical event topic.
5. The border event aggregation classification algorithm based on spatiotemporal correlation analysis according to claim 4 is characterized in that: The judgment formula for special items and common items is: in, Represents event S about the historical event topic set T j Probability value, Represents event S in the historical event subject set T j The number of occurrences in , if θ is the preset value, and the S event is not T j The end node of , then the event S is considered to be the historical event subject set T j The support is a unique item, otherwise it is a common item; the support calculation formula is: in, It represents item sets A and B about the historical event topic set T j The support of Indicates that item sets AB are simultaneously in the historical event topic set T j The number of times it appears in Indicates T j The number of event clusters that include item set A or item set B, Indicates T j The number of all events in the collection.
6. The border event aggregation classification algorithm based on spatiotemporal correlation analysis according to claim 4 is characterized in that: Constructing a reasoning graph based on association rules includes: Step 301: Assume that the historical event topic set T j The frequent itemset is L = {l1, l2, ..., l g }, there are g frequent item sets in total, traverse them, and take l1 as a link in the event graph, l q Represents the qth frequent set, let q = 2; Step 302: Select frequent set l q , if frequent set l q There are p nodes, and the node set is Traverse the node set in reverse order and take the end node of the event graph as the frequent set l q The current node of Step 303: q Move the current node forward to get As the current node of the frequent item, obtain the predecessor node set of the current node of the event graph, and determine the current node of the frequent item Whether it is in the set of predecessor nodes of the current node in the event graph; If the frequent item current node If it exists in the predecessor node set of the current node in the event graph, then set the current node in the event graph to the predecessor node set and The same node is repeated in step 303 to continue the iteration. If it does not exist, a new link is created in the current node of the event graph. The new link is connected to l q same; Step 304: Let q = q + 1, and repeat steps 302 to 303 until all frequent item sets L are traversed to obtain the updated event graph M. j .
7. The border event aggregation classification algorithm based on spatiotemporal correlation analysis according to claim 6 is characterized in that: Also includes: Step 305: Traverse the event graph M in sequence j For each node of the event graph M j A node m in h With the rear drive node m h+1 If the types are the same, delete the subsequent node m. h+1 , m h The following node becomes m h+2 , and m h The node adds a loop to itself, indicating that m h Nodes may recur.
8. The border event aggregation classification algorithm based on spatiotemporal correlation analysis according to claim 6 or 7 is characterized in that: Based on event categories and named entities, border events in the unclassified event set are classified and sorted through the event graph to obtain the aggregation results of border events, including: Step 401: Set a distance threshold ε, which consists of three parts: geographic location, time interval δ, and subject unit. Determine whether the border events are density-reachable based on the named entity. If two border events have the same geographic location, the same subject unit, and the time interval between occurrences is less than δ, then the density is reachable. Step 402: Create a new cluster, select an unlabeled event y from the unclassified event set, traverse other events, add y and events with density reachable to y to the cluster, and label the y event; Step 403: select an unmarked event u from the cluster in turn, traverse other events in the unclassified event set that are not assigned to the cluster, add events that are density-reachable to u to the cluster, and mark the u event; Step 404: Repeat step 403 until all points in the cluster are marked; Step 405: Repeat steps 402 to 404 until all events in the unclassified event set are marked and have their own clusters, and obtain the event set cluster R = {R1, R2, ..., R c }, through the event graph, the events in the event cluster are classified and collected, so that border events can be aggregated according to the event theme.
9. The border event aggregation classification algorithm based on spatiotemporal correlation analysis according to claim 8 is characterized in that: The event graph is used to classify and collect the events in the event cluster, including: Step 4051: Select an event cluster, sort the events in it in chronological order, traverse the events in it, and select one of the terminal events; Step 4052: Select a corresponding event graph according to the event category attribute of the terminal event, and select a link of the event graph; traverse the event cluster, and fill the events therein into the nodes of the corresponding category on the link in sequence according to the event category of each event; Step 4053: loop step 4052 to fill all links of the event graph with events; Step 4054: Score each link in turn, select the link with the highest score and take all events on the link as a set, and divide the topics according to the end nodes of the set to complete the aggregation of border events, and exclude all the above events from the event cluster. If there are other end nodes in the event cluster, repeat steps 4052 to 4054 until there are no end events in the event cluster.
10. The border event aggregation classification algorithm based on spatiotemporal correlation analysis according to claim 9 is characterized in that: The scoring rules for each link are: F=v*v / z Among them, F represents the score, z represents the number of nodes in a link, and v represents the filled node.
Citation Information
Cited By
Event cognitive reasoning method, device, equipment, medium and program product
CN121146055A