Hot event prediction method based on time-varying knowledge reasoning

By constructing a hot-spot event prediction model based on time-varying knowledge inference, using graph convolutional neural network to learn time-varying knowledge of events, the accuracy problem of existing methods in complex event prediction is solved, and efficient analysis and prediction of hot-spot events is achieved.

CN120235299APending Publication Date: 2025-07-01UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510325634.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

When existing event prediction methods deal with complex, dynamic, and cross-domain event data, they cannot effectively capture the rich knowledge and time dependence of the data, resulting in limited accuracy of prediction results. Especially when facing events with multiple factors intertwined, it is difficult to achieve efficient hot-spot event analysis and prediction.

Method used

A time-varying knowledge inference method is used to construct a circular evolution network based on graph convolutional neural network. Through the time-varying knowledge graph sequence learning evolutionary representation of entities and relationships, combining event time-varying knowledge subgraph extraction, evolutionary reasoning unit and event prediction unit, decision label analysis and event occurrence probability prediction are carried out, and decision label tendency trend chart and event occurrence probability chart are output.

Benefits of technology

It improves the accuracy of hot-spot event prediction, can more accurately link hot-spot events with timing knowledge, capture the time and space evolution process of events, and provide more efficient information processing methods, which are suitable for decision-making support systems in multiple fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235299A_ABST
    Figure CN120235299A_ABST
Patent Text Reader

Abstract

The invention discloses a hot event prediction method based on time-varying knowledge reasoning, which comprises the following steps: firstly, constructing a hot event prediction model based on time-varying knowledge reasoning, inputting a plurality of decision labels and hot event names, and carrying out event data acquisition and event time-varying knowledge sub-graph extraction; and then selecting different model processing modes according to historical event processing conditions, carrying out decision label analysis and event occurrence probability prediction until all time-varying sub-graphs of the event are fitted by the event prediction model, outputting a decision label tendency graph and an event occurrence probability graph, and completing hot event prediction. According to the method, association between the hot events and time sequence knowledge can be more accurately established, the time-varying association sub-graph of the events can be accurately extracted, analysis of complex events can be more visual and clearer, a more efficient information processing mode is provided for decision makers, and by automatically extracting, analyzing and predicting the events, the efficiency of the decision makers is improved. Tiny changes in the hot events are effectively captured, and the event prediction accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of time-varying knowledge graphs, and particularly relates to a hot event prediction method based on time-varying knowledge reasoning. Background Art

[0002] Event prediction refers to making pre-inferences and estimations about events that have not occurred or are occurring but with uncertain results based on various information, knowledge, and methods, so as to speculate on their possible future development trends, occurrence probabilities, etc. In recent years, the amount of information on global hot events on the Internet has shown an explosive growth. The relevant information of these events spreads rapidly and is widely shared in the form of news, social media, government announcements, and public reports, forming a large amount of, heterogeneous, and dynamically changing data resources. How to use these resources to monitor and predict the evolution of hot events in real time has become the focus of attention in fields such as international relations, foreign policies, and security prevention. Therefore, in order to meet the needs of a large number of users and organizers for the effective management and situation analysis of emergencies, the research on emergency detection and search for the network has important value and significance. However, in the face of a large amount of heterogeneous network information, how to effectively extract valuable knowledge from it and reveal the internal connections between events, especially the complex relationships across time and space dimensions, remains an important research problem.

[0003] Existing event prediction methods usually directly mine the logical relationships between long text data using natural language processing methods, and then rely on statistical analysis, time series prediction, or machine learning models to fit historical data to predict the occurrence possibility of future events. With the proposal of graph computing methods, many studies also believe that the occurrence of events will bring obvious changes to users' behaviors. In the early stage of event occurrence, by analyzing the changes in users' behaviors, classifying the behavioral subjects, and performing sentiment analysis, the detection of events can be achieved. However, the limitations of these methods are also relatively obvious, especially when dealing with complex, dynamic, and cross-domain event data. Most existing methods cannot effectively capture the rich knowledge information and time dependence of the data. Especially when an event is the product of multiple factors intertwined, the accuracy of the prediction results is often limited. Currently, there are mainly three categories of existing hot event prediction methods, namely rule-based, statistical and machine learning-based, and knowledge graph-based hot event analysis and prediction methods.

[0004] The rule-based hot event analysis and prediction method analyzes the laws and patterns of event occurrence by combining historical data with established rules. The rule-based analysis method can help analyze the response of decision makers under certain conditions, but in practical applications, the model relies on a large amount of prior knowledge and rules, and the definition of rules for different types of events requires a lot of manual analysis and annotation, which is a huge workload and the performance is limited by the quality of parsing rules in historical scenarios, making it difficult to cope with new scenarios; the hot event prediction method based on statistics and machine learning models a large amount of historical event data to mine potential laws and patterns. This type of method can automatically learn from data and optimize, but it is very dependent on the contextual feature relationship of long text data, especially when facing ambiguous or ambiguous text, the context of some events may cause large errors in the prediction results. And for new, unseen events or changes, models based on statistics and machine learning may not be able to effectively capture new patterns or laws, and their adaptability is poor; knowledge graphs usually refer to a relational graph constructed based on existing information and knowledge, which contains various entities (such as people, events, places, organizations, etc.) and their relationships with each other. In the field of event prediction, static knowledge graphs are used to predict possible future events or trends by analyzing the relationships between entities and their historical patterns. However, existing knowledge graphs do not have the concept of event stamps and cannot reflect the order of events in the information in the graph. The reasoning ability is mainly based on existing static knowledge and rules and lacks time-based dynamic reasoning. It is unable to handle complex and rapidly changing situations and capture the dynamic evolution of complex events.

[0005] In summary, existing event prediction methods often only consider the statistical relationship between events and establish prediction models through supervised learning methods. However, these methods only consider the relationship between events themselves, and pay little attention to the factors that cause the events to be associated. There is a problem of low event prediction accuracy, so it is difficult to apply them to actual event analysis work. Summary of the invention

[0006] To solve the above technical problems, the present invention provides a hot event prediction method based on time-varying knowledge reasoning, which can more accurately associate hot events with temporal knowledge and take into account the spatiotemporal evolution of events, thereby effectively improving the accuracy of event prediction.

[0007] The technical solution adopted by the present invention is: a hot event prediction method based on time-varying knowledge reasoning, and the specific steps are as follows:

[0008] S1. Construct a hot event prediction model based on time-varying knowledge reasoning;

[0009] The time-varying knowledge graph is regarded as a sequence of knowledge graphs, and the entire sequence of knowledge graphs is uniformly modeled. All historical facts are encoded into entity and relationship representations, and a hot event prediction model based on time-varying knowledge reasoning is constructed. The model uses a recurrent evolutionary network based on the graph convolutional neural network (GCN) to learn the evolutionary representations of entities and relationships at each timestamp through recurrently modeling the knowledge graph sequence.

[0010] Among them, the event prediction model includes: 1 event time-varying knowledge subgraph extraction module, 2 evolutionary reasoning units, and 1 event prediction unit. The 2 evolutionary reasoning units are the fitting hot event model and the initializing hot event model respectively.

[0011] S2. Based on step S1, through the event time-varying knowledge subgraph extraction module, event data collection and event time-varying knowledge subgraph extraction are performed.

[0012] S3. Based on the time-varying knowledge subgraph obtained in step S2, it is input into the evolutionary reasoning unit for temporal knowledge graph evolutionary reasoning.

[0013] For different event time-varying knowledge subgraphs, different evolutionary reasoning units are used for two modes of fitting and initialization. The fitting hot event model adopts the fitting mode to continuously fit new temporal knowledge information features based on the learned historical entity and relationship features. The initializing hot event model adopts the initialization mode to train the original model for new events and learn the structural dependencies and temporal feature information of entities and relationships.

[0014] After the event time-varying knowledge subgraph is sliced and divided at a set time interval, the sliced data undergoes model structure dependency learning and temporal feature learning in chronological order.

[0015] S4. Based on step S3, through the event prediction unit, decision label analysis and event occurrence probability prediction are performed until all time-varying subgraphs of the event are completely fitted by the event prediction model, and a decision label tendency trend graph and an event occurrence probability graph are output to complete the hot event prediction.

[0016] Among them, after updating the entity and relationship embedding information and fitting the graph network structure, the evolutionary score of the decision label at each time slice is calculated. After the model finishes fitting all the sliced data, a decision label tendency trend graph and an event occurrence probability graph are output.

[0017] Furthermore, step S2 is specifically as follows:

[0018] According to the constraint conditions of various factors affecting the decision of the person on entity and relationship types, the schema of the time-varying knowledge graph is defined, and the schema includes entity types and relationship types.

[0019] Among them, entity types include: person, organization, geographical location, and event, and relationship types include: positive emotional attitude, negative emotional attitude, general association relationship, and event participation method.

[0020] Then, crawl text data from social media platforms, news release platforms, and current affairs comment platforms, extract quadruple data that conforms to the above-mentioned pattern, and store it in the total database to form an original dataset.

[0021] Then, input the name of the hot event in the original dataset and the preset set of decision labels, retrieve the target entities from the total database according to the decision labels, and extract a set of quadruples centered on the entity with the association relationship extended to the two-degree neighborhood. Denoise and clean the set of quadruples, including: removing data with missing or incorrectly formatted timestamps, and filtering low-reliability relationships based on a confidence threshold to generate a time-varying knowledge subgraph, that is, sort the cleaned quadruples in timestamp order to construct a time-varying knowledge subgraph, where the nodes represent entities, the edges represent relationship types, and the edge attributes include timestamps and confidence levels.

[0022] Among them, the decision label is the key political figure entity pair of the event.

[0023] Furthermore, the specific steps of step S3 are as follows:

[0024] Based on the event time-varying knowledge subgraph obtained in step S2, select the evolutionary reasoning unit to be used, that is, judge whether the event model already exists. If so, select the fitting processing hot event model; if not, select the initialized hot event model to learn the structural dependencies and temporal features.

[0025] First, slice the time-varying knowledge graph in the time dimension and represent it in the form of a knowledge graph sequence, that is, {G1, G2, …, G t , …}. For each knowledge graph G t = (V, R, E t ), represent it in the form of a quadruple (s, r, o, t).

[0026] Among them, V represents the entity set, R represents the relationship set, and E t represents the fact set at the discrete timestamp t. s, r, and o represent the head entity, relationship, and tail entity respectively.

[0027] Then, use a recurrent neural network (RNN) to model the event information, encode all historical facts into the embedding representations of entities and relationships into the same vector space. And use the stacking of multi-layer relationship-aware graph convolutional neural networks to obtain the structural features in the knowledge graph and capture the dependency relationships between entities in the knowledge graph.

[0028] Then, a knowledge graph sequence of a specific length is randomly initialized and mapped into an entity sequence and a relationship sequence. For the static knowledge graph slice at a specified timestamp, the tail entity of the l-th layer obtains information from the head entity in the message passing framework of the graph convolutional neural network, sets the embedding relationship at the l-th layer, and obtains the structure-dependent embedding representation at the l+1-th layer. The expression is as follows:

[0029]

[0030] Among them, represents the entity o at time t, and s is the embedding representation learned according to the relevant triple ε t at the l-th layer, represents the evolving representation of the relationship r at time t, represents the real number field, and d represents the dimension of the vector. respectively represent the weight matrix for aggregating information around nodes and the weight matrix for passing the information of the node itself. c o represents the normalization constant, which is equal to the in-degree of the entity o, and f(·) represents the ReLU activation function.

[0031] For entities that do not involve any facts, only a self-loop operation with an additional parameter is performed. The relationship-aware graph convolutional neural network obtains the evolving representation of the entity embedding according to the facts occurring at each timestamp, and realizes self-evolution with isolated entities through the self-loop operation, aggregating all its static attribute information.

[0032] The entity embedding is modeled by stacking the sequential patterns of the relationship-aware graph convolutional neural network. The evolutionary inference unit uses a gated unit to learn temporal features. The specific expression is as follows:

[0033]

[0034] Among them, represents the Hadamard product, and H t represents the cumulative representation of the model for the entity embedding matrix at time step t, represents the preliminary response of the model to the input at the current time step t.

[0035] The candidate hidden state is calculated as follows:

[0036]

[0037] Among them, W h represents the weight matrix, r t represents the reset gate, X t represents the input at the current time step, b h represents the bias term, and ⊙ represents element-wise multiplication.

[0038] Time gate U t The expression for performing the non - linear transformation is as follows:

[0039] U t = σ(W4H t-1 + b)

[0040] Where, σ(·) represents the sigmoid function, W4 represents the weight matrix of the time gate, and b represents the bias term.

[0041] The sequential pattern of the relationship captures the information of the entities involved in the corresponding fact. The relationship The embedding at timestamp t is affected by the relevant entities Embedded at timestamps t and t - 1. Then, a gated unit component is used to model the sequential pattern of the relationship. Through the average pooling operation of the relevant entity embedding matrix, the relationship The expression for the input of the gated unit at timestamp t is as follows:

[0042]

[0043] Where, Represents the embedding representation of the relationship r in R, and is vector - concatenated with the output of the pooling layer. For the relationship that has no corresponding fact at timestamp t The relationship embedding matrix R t-1 Is updated to R t , and the expression is as follows:

[0044] R t = GRU(R t-1 , R t ′ )

[0045] Then, for the static knowledge graph which is a multi - relational graph, each entity aggregates all its static attribute information using a single - layer graph convolutional neural network with self - loops. The update rule definition expression of the static graph is as follows:

[0046]

[0047] Where, And Represent the corresponding rows in H s And H ′s , H s And H ′s Represent the output and randomly - initialized input embedding matrices respectively. Represents the relationship matrix in the graph convolutional neural network unit. Represents the relationship matrix of r s In R - GCN, Υ(·) represents the ReLu activation function, ci is a normalized constant representing the number of entities connected to entity i, ε s represents the set of all relationships.

[0048] After obtaining the relationship matrix, the measurement method of the cosine similarity of the embedding is adopted, and the angle between the evolutionary embedding and the static embedding of the same entity is limited to 90°. The specific definition expression is as follows:

[0049] θ x = min(γx, 90°), x ∈ [0, 1, …, v]

[0050] where the γ parameter represents the step size of the angle increase and controls the rising speed of the angle between the embeddings. θ x represents the angle between the evolutionary embedding and the static embedding of entity x, and v represents the total number of entities. Then the definition expression of the static constraint component loss at timestamp t is as follows:

[0051]

[0052] Then, by optimizing the loss make and gradually approach in space. If the angle between the two is less than θ x , the loss is 0. Finally, the expression of the total static constraint loss is as follows:

[0053]

[0054] Finally, through multiple rounds of iteration of the evolutionary reasoning unit until L st obtains a minimum value, so that the static constraint finally converges, and the final embedding vectors of entities and relationships at the current timestamp are obtained.

[0055] Furthermore, the specific steps of step S4 are as follows:

[0056] Decision label analysis is used to predict the tendency relationship probability of entity pairs corresponding to decision labels at several timestamps of the event time-varying knowledge subgraph. Event occurrence probability prediction is used to comprehensively predict the probability of an event occurring at several timestamps based on the tendency relationships of several decision labels. The expressions are as follows:

[0057]

[0058]

[0059] where the entity pair corresponding to the i-th decision label of the event is defined as s i and o i , and when the event time-varying knowledge graph is divided into k different time-period subgraphs {TG1, TG2, …, TGk}, each TG i includes the events that occur during this time period and their related entities and relationships.

[0060] For TG i The prediction of the tendency relationship of the decision label depends on the information of the knowledge graph slices {G t-Δt+1 , …, G t} at the previous Δt moments. The information of {G t-Δt+1 , …, G t} is modeled as the entity embedding matrix and the relationship embedding matrix

[0061] Then the event prediction unit uses a strategy of step-by-step iterative training for decision label analysis and event occurrence probability prediction, as follows:

[0062] First, use TG1 as the initial training set of the model, use TG2 as the validation set, and optimize the model parameter learning through three types of features, namely structural dependence, temporal features, and static attributes. After training is completed, adjust the model based on the performance in the validation set TG2, and update the entity and relationship embeddings of the TG1 subgraph for predicting the decision label tendency relationship and event occurrence probability corresponding to the t moment of the subgraph.

[0063] After the training of TG1 is completed and verified, use TG2 as the new training set to further train the model, and use TG3 as the validation set, and iterate in this way until the last subgraph TG k .

[0064] Regard entity embedding fitting and relationship embedding fitting as multi-label learning problems, and let and represent the label vectors of entity and relationship fitting tasks at time t + 1 respectively. Each element in the vector takes the value of 1 for the occurrence of a fact, otherwise it takes the value of 0. The loss function expressions of entity embedding fitting and relationship embedding fitting are as follows:

[0065]

[0066] where T represents the total number of time slices in the subgraph training set, that is, the length of the divided knowledge graph sequence. and represent the probability scores of entities and relationships respectively.

[0067] By setting the learnable parameters λ1 and λ2, the final total loss expression of the model is as follows:

[0068] L = λ1L e + λ2L r + Lst

[0069] By optimizing the total loss of the model, the dynamic embedding fitting results of different entity and relationship elements in the time-varying knowledge graph under the decision-making mode constraint are obtained.

[0070] Then, the dynamic embeddings H t , R t of entities and relationships are input into the scoring module. Given a decision label entity pair (head entity s, tail entity o), let the head entity s traverse all possible valid relationships r that can occur. Among them, the valid relationships are divided into beneficial relationships and non-beneficial relationships. Using the default prediction triple set of these two types of relationships, the conditional probability vector is modeled. Finally, the conditional probability score of the tail entity o at time t under each relationship r i is simulated through a ConvTransE decoder containing a one-dimensional convolutional layer and a fully connected layer, and the calculation formula is as follows:

[0071]

[0072] Then, according to the score ranking, the scores under all relationships are weighted and summed to obtain the comprehensive score of the decision label at time t, which reflects the possibility of forming the target label in the current relationship environment (s, o). The expression is as follows:

[0073]

[0074] Among them, w i represents the weight of the entity pair (s, o) in the relationship r i ; att i represents the weighted coefficient of the relationship type of the relationship r i , which is defined as: +1 for beneficial relationships, -1 for non-beneficial relationships, and 0 for others; n r represents the number of possible relationships.

[0075] After the scores of all decision label entity pairs are calculated, the scores of all labels are weighted and summed and normalized to obtain the event occurrence probability at time t. The expression is as follows:

[0076]

[0077] Finally, it is judged whether there is new data in the current time-varying subgraph. If so, return to step S3 until all time-varying subgraphs of the event are all fitted by the model, and the decision label tendency trend graph and the event occurrence probability graph are output to complete the hot event prediction.

[0078] Advantages of the present invention: The method of the present invention first constructs a hot event prediction model based on time-varying knowledge reasoning, inputs a number of decision labels and the names of hot events, collects event data and extracts time-varying knowledge subgraphs of the events, and then selects different model processing methods according to the historical processing situations of the events to analyze decision labels and predict the occurrence probabilities of the events until all time-varying subgraphs of the events are all fitted by the event prediction model, outputs a decision label tendency trend graph and an event occurrence probability graph, and completes the prediction of hot events. By defining decision label entity pairs, the method of the present invention decouples complex events into the interest game relationships of several key entity pairs, can more accurately establish the association between hot events and temporal knowledge, accurately extracts the time-varying association subgraphs of the events, makes the analysis of complex events more intuitive and clear, and provides a more efficient information processing method for decision-makers; the method of the present invention designs and implements a continuous tracking and prediction scheme for hot events for the time-varying knowledge graph, solves the problems that the existing hot event prediction methods usually rely on expert experience and manual rules, are inefficient and have serious subjective biases, effectively captures the subtle changes in hot events through automatic event extraction, analysis and prediction, and improves the accuracy of event prediction; the method of the present invention has good cross-domain applicability in practical applications. The abstraction of decision labels and the dynamic update method of the time-varying knowledge graph can be widely applied to the prediction and analysis in multiple fields such as economy, international trade, and politics. Through transfer learning and the generality of the model, it can be applied to the decision support systems in other fields to provide richer and more diverse prediction functions. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 It is a flowchart of a method for predicting hot events based on time-varying knowledge reasoning according to the present invention.

[0080] Figure 2 It is a structural diagram of a hot event prediction model based on time-varying knowledge reasoning in an embodiment of the present invention.

[0081] Figure 3 It is a schematic structural diagram of an evolutionary reasoning unit in an embodiment of the present invention.

[0082] Figure 4 It is a schematic structural diagram of an event prediction unit in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0083] The method of the present invention will be further described below in conjunction with the drawings and embodiments.

[0084] As Figure 1 shown, a flowchart of a method for predicting hot events based on time-varying knowledge reasoning according to the present invention is as follows:

[0085] S1. Construct a hot event prediction model based on time-varying knowledge reasoning;

[0086] As shown Figure 2 in the figure, the time-varying knowledge graph is regarded as a sequence of knowledge graphs, and the entire knowledge graph sequence is uniformly modeled. All historical facts are encoded into entity and relationship representations to facilitate event prediction tasks. A hot event prediction model based on time-varying knowledge reasoning is constructed. The model adopts a recurrent evolutionary network based on the graph convolutional neural network (GCN), and learns the evolutionary representations of entities and relationships at each timestamp by recurrently modeling the knowledge graph sequence. Through the learning of three types of features (structural dependencies, temporal features, static attributes), the model can relatively clearly understand the potential knowledge in the knowledge graph, excavate deeper decision-making elements of characters under each slice, and then continuously iterate and update to form a more perfect and more accurately embedded fitting matrix of entity and relationship evolution representations, which is used to calculate the tendency relationship probability and event occurrence probability of subsequent decision labels.

[0087] Among them, the event prediction model includes: 1 event time-varying knowledge subgraph extraction module, 2 evolutionary reasoning units, and 1 event prediction unit. The 2 evolutionary reasoning units are respectively a fitting hot event model and an initializing hot event model.

[0088] S2. Based on step S1, through the event time-varying knowledge subgraph extraction module, event data collection and event time-varying knowledge subgraph extraction are carried out;

[0089] S3. Based on the time-varying knowledge subgraph obtained in step S2, it is input into the evolutionary reasoning unit for temporal knowledge graph evolutionary reasoning;

[0090] For different event time-varying knowledge subgraphs, different evolutionary reasoning units are used for two modes of fitting and initialization. The fitting hot event model adopts the fitting mode and continues to fit new temporal knowledge information features based on the learned historical entity and relationship features. The initializing hot event model adopts the initialization mode to train the original model for new events and learn the structural dependencies and temporal feature information of entities and relationships.

[0091] After the event time-varying knowledge subgraph is sliced and divided at a set time interval, the sliced data undergoes model structure dependency learning and temporal feature learning in chronological order.

[0092] S4. Based on step S3, through the event prediction unit, decision label analysis and event occurrence probability prediction are carried out until all time-varying subgraphs of the event are completely fitted by the event prediction model, and a decision label tendency trend graph and an event occurrence probability graph are output to complete the hot event prediction;

[0093] Among them, after updating the entity and relationship embedding information and fitting the graph network structure, the evolution score of the decision label at each time slice is calculated. After the model finishes fitting all the slice data, the decision label tendency trend graph and the event occurrence probability graph are output.

[0094] In this embodiment, the step S2 is specifically as follows:

[0095] The occurrence and development of hot events are often affected by many factors, including: historical events, geopolitics and culture, public sentiment, and behind-the-scenes forces. Based on the above factors, the decisions of the group of people have a catalytic effect on the occurrence and development of hot events, and to a large extent will promote the change of the event situation. Therefore, in this embodiment, according to the constraints of each factor affecting people's decisions on entities and relationship types, the schema of the time-varying knowledge graph is defined, and the schema includes entity types and relationship types.

[0096] Among them, entity types include: people, organizations, geographical locations, and events, and relationship types include: positive emotional attitude, negative emotional attitude, general association relationship, and event participation method.

[0097] Then, text data is crawled from social media platforms, news release platforms, and current affairs comment platforms, and quadruple data that conforms to the schema is extracted and stored in the total database to form an original data set. Hot events are usually triggered by the behaviors and decisions of various entities in complex interest games. Although the specific forms and contents of events vary widely, from a macro perspective, all events can be attributed to the interest influence relationships between different entities.

[0098] Then, input the name of the hot event in the original data set and the preset set of decision labels, retrieve the target entities from the total database according to the decision labels, and extract the quadruple set centered on the entity with the association relationship extended to the two-degree neighborhood. Denoising and cleaning are performed on the quadruple set, including: removing data with missing timestamps or incorrect formats, and filtering low-reliability relationships based on the confidence threshold to generate a time-varying knowledge subgraph, that is, sorting the cleaned quadruples in timestamp order to construct a time-varying knowledge subgraph, where the nodes represent entities, the edges represent relationship types, and the edge attributes include timestamps and confidence levels.

[0099] Among them, the decision label is the key political figure entity pair of the event.

[0100] In this embodiment, the step S3 is specifically as follows:

[0101] Based on the event time-varying knowledge subgraph obtained in step S2, select the evolution inference unit to be used, that is, determine whether the event model already exists (whether the event is an event that has been processed). If so, select the model for fitting and processing hot events, and if not, select the model for initializing hot events to learn the structural dependencies and temporal features.

[0102] As shown Figure 3 in the figure, first, the time-varying knowledge graph is sliced in the time dimension and represented in the form of a knowledge graph sequence, i.e., {G1, G2, …, G t , …}. For each knowledge graph G t = (V, R, E t ), it is represented in the form of a quadruple (s, r, o, t).

[0103] Among them, V represents the set of entities, R represents the set of relationships, and E t represents the set of facts at the discrete timestamp t. s, r, and o represent the head entity, relationship, and tail entity respectively.

[0104] Then, a recurrent neural network (RNN) is used to model the event information, encoding all historical facts into the embedding representations of entities and relationships in the same vector space as the input for the subsequent feature learning process.

[0105] The structural dependency relationships between concurrent facts capture the associations between entities through facts and the associations between relationships through shared entities. Since each knowledge graph slice is a multi-relational graph, a relation-aware graph convolutional neural network can aggregate the neighbor nodes and relationship information of a specified node. Therefore, a stack of multi-layer relation-aware graph convolutional neural networks is used to obtain the structural features in the knowledge graph and capture the dependency relationships between various entities in the knowledge graph.

[0106] Next, a knowledge graph sequence of a specific length is randomly initialized and mapped into an entity sequence and a relationship sequence. For the static knowledge graph slice at a specified timestamp, the tail entity at the l-th layer obtains information from the head entity under the message passing framework of the graph convolutional neural network, sets the embedding relationship at the l-th layer, and obtains the embedding representation based on the structural dependency at the l + 1-th layer. The expression is as follows:

[0107]

[0108] Among them, represents the entity o at time t, and s is the embedding representation learned by s at the l-th layer according to the relevant triple ε t , represents the evolutionary representation of the relationship r at time t, represents the real number field, and d represents the dimension of the vector. respectively represent the weight matrix for aggregating the information around the node and the weight matrix for passing the information of the node itself. c o represents the normalization constant, which is equal to the in-degree of the entity o, and f(·) represents the ReLU activation function.

[0109] For entities that do not involve any facts, only the additional parameter W3 is executedl Self-loop operation. The relation-aware graph convolutional neural network obtains the evolutionary representation of entity embeddings based on the facts occurring at each timestamp and realizes self-evolution with isolated entities through self-loop operation, aggregating all its static attribute information.

[0110] For any entity, the temporal patterns contained in its historical facts reflect its behavioral trends and preferences. To capture as many historical facts as possible, entity embeddings can be directly modeled by stacking the sequential patterns of relation-aware graph convolutional neural networks. However, there is an over-smoothing problem when the same entity pair has repeated relations at adjacent timestamps, that is, the embeddings of entities tend to converge to the same value. And when the historical knowledge graph sequence gradually becomes longer, the large stacking of graph convolutional neural networks may lead to vanishing gradients.

[0111] Then, in this embodiment, entity embeddings are modeled by stacking the sequential patterns of relation-aware graph convolutional neural networks, and the evolutionary inference unit uses a gated unit to learn temporal features. The specific expression is as follows:

[0112]

[0113] Among them, denotes the Hadamard product, and H t represents the cumulative representation of the model for the entity embedding matrix at time step t, represents the initial response of the model to the input at the current time step t.

[0114] Candidate hidden state The calculation formula is as follows:

[0115]

[0116] Among them, W h denotes the weight matrix, r t denotes the reset gate, X t represents the input at the current time step, b h denotes the bias term, and ⊙ denotes element-wise multiplication.

[0117] Time gate U t The expression for performing nonlinear transformation is as follows:

[0118] U t = σ(W4H t-1 + b)

[0119] Among them, σ(·) represents the sigmoid function, W4 represents the weight matrix of the time gate, and b represents the bias term.

[0120] The sequential pattern of the relation captures the information of the entities involved in the corresponding facts. The embedding of the relation at timestamp t is affected by the relevant entities For the influence embedded at timestamps t and t-1, a gated unit component is used to model the sequential pattern of the relationship. Through the average pooling operation of the relevant entity embedding matrix, the relationship The expression for the gated unit input at timestamp t is as follows:

[0121]

[0122] where, represents the embedding representation of relationship r in R, and is vector-connected with the output of the pooling layer. For a relationship that has no corresponding fact at timestamp t The relationship embedding matrix R is updated to R t-1 through a gated recurrent unit t , and the expression is as follows:

[0123] R t = GRU(R t-1 , R t ′ )

[0124] Since each hidden unit has a separate reset and update gate, each hidden unit will learn to capture dependencies within different time ranges, and finally learn the global and evolving features of the relationship.

[0125] Then, in addition to the information contained in the knowledge graph sequence in the time dimension, some static attributes of the entity can be regarded as background knowledge of the time-varying knowledge graph, and this part of the information helps the model learn a more accurate entity evolution representation. Since the static knowledge graph is a multi-relational graph, each entity aggregates all its static attribute information using a single-layer graph convolutional neural network with self-loops. The update rule definition expression of the static graph is as follows:

[0126]

[0127] where, and represent the corresponding rows in H s and H ′s , H s and H ′s represent the output and randomly initialized input embedding matrices respectively. represents the relationship matrix in the graph convolutional neural network unit. represents the relationship matrix of r s in R-GCN, Υ(·) represents the ReLu activation function, c i is the normalization constant, representing the number of entities connected to entity i, and ε s represents the set of all relationships.

[0128] After obtaining the relational matrix, the measurement method of embedded cosine similarity is adopted, and the angle between the evolutionary embedding and the static embedding of the same entity is limited to 90°. The specific definition expression is as follows:

[0129] θ x =min(γx,90°),x∈[0,1,…,v]

[0130] Among them, the γ parameter represents the step size of angle growth and controls the rising speed of the angle between embeddings. θ x represents the angle between the evolutionary embedding and the static embedding of entity x, and v represents the total number of entities. Then the definition expression of the static constraint component loss at timestamp t is as follows:

[0131]

[0132] Then, by optimizing the loss make and gradually approach in space. If the angle between the two is less than θ x , the loss is 0. Then, the closer to t, the larger θ x is, and the smaller the binding force of the static graph spectrum is. Finally, the expression of the total static constraint loss is as follows:

[0133]

[0134] Finally, through multiple rounds of iteration of the evolutionary reasoning unit until L st obtains a minimum value, so that the static constraint finally converges, and the final embedding vectors of entities and relationships at the current timestamp are obtained.

[0135] In this embodiment, the specific steps of step S4 are as follows:

[0136] Decision label analysis is used to predict the tendency relationship probability of entity pairs corresponding to decision labels at several timestamps of the event time-varying knowledge subgraph. Event occurrence probability prediction is used to comprehensively predict the probability of an event occurring at several timestamps based on the tendency relationships of several decision labels. The expressions are as follows:

[0137]

[0138] Among them, the entity pair corresponding to the i-th decision label of the event is defined as s i and o i . When dividing the event time-varying knowledge graph into k different time-period subgraphs {TG1, TG2, …, TG k} at time interval Δt, each TG i includes the events that occurred in this time period and their related entities and relationships.

[0139] For TGi The prediction of the tendency relationship of decision labels depends on the information of the knowledge graph slices {G t-Δt+1 , …, G t} at the previous Δt moments. The information of {G t-Δt+1 , …, G t} is modeled as the entity embedding matrix and the relationship embedding matrix

[0140] at time t. Figure 4 As shown in

[0141] , the event prediction unit uses a strategy of step-by-step iterative training for decision label analysis and event occurrence probability prediction, as follows:

[0142] First, use TG1 as the initial training set of the model, use TG2 as the validation set, and optimize the model parameter learning through three types of features, namely structural dependence, temporal features, and static attributes. After the training is completed, adjust the model based on the performance in the validation set TG2, and update the entity and relationship embeddings of the TG1 subgraph for predicting the tendency relationship of decision labels and the event occurrence probability corresponding to the subgraph at time t. k After the training of TG1 is completed and verified, use TG2 as the new training set to further train the model, and use TG3 as the validation set, and iterate in this way until the last subgraph TG

[0143] k is processed. Continuously update the model to adapt to the changes in the new time period. In this way, during each round of iterative training, the latest knowledge graph subgraph is used to further optimize the entity and relationship embeddings to ensure that the model can continuously adapt to the dynamic changes of the knowledge graph, so as to effectively capture the regularity and trend of time-varying events.

[0143] Regard the entity embedding fitting and the relationship embedding fitting as multi-label learning problems, and let and represent the label vectors of the entity and relationship fitting tasks at time t + 1 respectively. Each element in the vector takes the value of 1 for the occurrence of a fact, and 0 otherwise. The loss function expressions of the entity embedding fitting and the relationship embedding fitting are as follows:

[0144]

[0145] where T represents the total number of time slices in the subgraph training set, that is, the length of the knowledge graph sequence divided. and represent the probability scores of the entity and the relationship respectively.

[0146] By setting the learnable parameters λ1 and λ2, the final total loss expression of the model is as follows:

[0147] L = λ1Le +λ2L r +L st

[0148] By optimizing the total loss of the model, the dynamic embedding fitting results for different entity and relationship elements in the time-varying knowledge graph under the decision-making mode constraints are obtained.

[0149] Then, the dynamic embeddings H t , R t of entities and relationships are input into the scoring module. Given a decision label entity pair (head entity s, tail entity o), let the head entity s traverse all possible valid relationships r. Among them, the valid relationships are divided into beneficial relationships and non-beneficial relationships. Using the default prediction triple set of these two types of relationships, model their conditional probability vectors. Finally, through the ConvTransE decoder including a one-dimensional convolutional layer and a fully connected layer, the conditional probability score of the tail entity o at time t under each relationship r i is simulated and obtained. The calculation formula is as follows:

[0150]

[0151] Use the scoring module to calculate the comprehensive score of the tendency change of the target label (the relationship of entities such as the person, place, and content describing the event) at a certain moment. This score can reflect the possibility of forming the target label using the current relationship information and reflect the development trend of the event. Using the change of this score, the judgment of whether the event occurs can be realized in combination with expert knowledge, improving the practicality of event prediction from the perspective of the actual work process.

[0152] Then, according to the score ranking, the scores under all relationships are weighted and summed to obtain the comprehensive score of the decision label at time t, reflecting the possibility of forming the target label in the current relationship environment (s, o). The expression is as follows:

[0153]

[0154] Among them, w i represents the weight of the entity pair (s, o) in the relationship r i , reflecting the importance of different relationships to this entity pair. The higher the ranking of the conditional probability score, the greater the weight value; att i represents the weighted coefficient of the relationship type of the relationship r i . It is defined as: +1 for beneficial relationships, -1 for non-beneficial relationships, and 0 for others. n r represents the number of possible relationships.

[0155] After calculating the scores of all decision label entity pairs, the scores of all labels are weighted and summed and normalized to obtain the event occurrence probability at time t. The expression is as follows:

[0156]

[0157] Finally, determine whether there is new data in the current time-varying sub-graph. If so, go back to step S3 until all time-varying sub-graphs of the event are all completed by model fitting. Output the decision label tendency graph and the event occurrence probability graph, complete the hot event prediction, and display the tendency change of the relationship between key entities and the possibility trend of the event for the user to refer to and analyze.

[0158] In this embodiment, it further includes step S5, which is specifically as follows:

[0159] Based on step S4, for the events that have occurred historically and are known, determine whether the predicted results output are correct, that is, compare the actual occurrence time of the event with the predicted results to verify and optimize the effectiveness and accuracy of the event prediction model.

[0160] In summary, the method of the present invention abstracts decision label entity pairs from hot events, effectively transforms complex event behaviors into structured decision frameworks, establishes associations between hot events and temporal knowledge and extracts event-associated sub-graphs, with high practicality and accuracy; by slicing the time-varying knowledge graph, using graph embedding and graph convolutional neural networks to vectorize entities and relationships, and dynamically fitting according to the time-varying knowledge within the slices, it can more accurately capture the potential laws and trend changes in the process of event development, can adapt to the rapid changes of knowledge triples in the time-varying knowledge graph, and improve the timeliness and accuracy of event prediction.

[0161] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A hot event prediction method based on time-varying knowledge reasoning, the specific steps are as follows: S1. Construct a hot event prediction model based on time-varying knowledge reasoning; The time-varying knowledge graph is regarded as a knowledge graph sequence, and the entire knowledge graph sequence is uniformly modeled. All historical facts are encoded as entity and relationship representations, and a hot event prediction model based on time-varying knowledge reasoning is constructed. The model adopts a cyclic evolutionary network based on the graph convolutional neural network GCN, and learns the evolutionary representation of entities and relationships at each timestamp through cyclic modeling of the knowledge graph sequence; in, The event prediction model includes: an event time-varying knowledge subgraph extraction module, two evolutionary reasoning units and an event prediction unit; the two evolutionary reasoning units are respectively a hot event model fitting and an initialization hot event model; S2. Based on step S1, event data collection and event time-varying knowledge subgraph extraction are performed through the event time-varying knowledge subgraph extraction module; S3, based on the time-varying knowledge subgraph obtained in step S2, input it into the evolutionary reasoning unit to perform temporal knowledge graph evolutionary reasoning; For different event time-varying knowledge subgraphs, different evolutionary reasoning units are used to perform fitting and initialization processing. The fitting mode is used to fit the hot event model, and new temporal knowledge information features are continued to be fitted based on the learned historical entity and relationship features. The initialization mode is used to initialize the hot event model, and the original model is trained for new events and the structural dependencies and temporal feature information of entities and relationships are learned. After the event-time-varying knowledge subgraph is sliced ​​and divided according to the set time interval, the sliced ​​data is subjected to model structure dependency learning and temporal feature learning in chronological order; S4, based on step S3, the event prediction unit performs decision label analysis and event probability prediction until all time-varying subgraphs of the event are fitted by the event prediction model, outputs a decision label tendency trend graph and an event probability graph, and completes the hot event prediction; After updating the entity and relationship embedding information and fitting the graph network structure, the evolution score of the decision label at each time slice is calculated. After the model is fitted for all slice data, the decision label trend graph and event occurrence probability graph are output.

2. According to claim 1, a hot event prediction method based on time-varying knowledge reasoning is characterized in that: The step S2 is specifically as follows: According to the constraints of various factors affecting character decision-making on entity and relationship types, a pattern of a time-varying knowledge graph is defined, wherein the pattern includes entity types and relationship types; Among them, entity types include: people, organizations, geographic locations and events, and relationship types include: positive emotional attitudes, negative emotional attitudes, general association relationships and event participation methods; Then crawl text data from social media platforms, news release platforms, and current affairs commentary platforms, extract quadruple data that conforms to the pattern, and store it in the main database to form an original data set; Then, the hot event names and the preset decision label set in the original data set are input, the target entity is retrieved from the total database according to the decision label, and a set of four-tuples centered on the entity and with associations extending to the two-degree neighborhood is extracted; the four-tuple set is denoised and cleaned, including: removing data with missing timestamps or incorrect formats, and filtering low-reliability relationships based on confidence thresholds to generate a time-varying knowledge subgraph, that is, sorting the cleaned four-tuples in timestamp order to construct a time-varying knowledge subgraph, whose nodes represent entities, edges represent relationship types, and edge attributes include timestamps and confidences; The decision label is a key political entity pair of the event.

3. The method for predicting hot events based on time-varying knowledge reasoning according to claim 1 is characterized in that: The step S3 is specifically as follows: Based on the event time-varying knowledge subgraph obtained in step S2, the evolutionary reasoning unit to be used is selected, that is, whether the event model already exists is determined. If so, the hot event model is selected for fitting. If not, the hot event model is initialized to learn structural dependencies and temporal features. First, the time-varying knowledge graph is sliced ​​in the time dimension and represented in the form of a knowledge graph sequence, that is, {G1,G2,…,G t ,…}; For each knowledge graph G in the sequence t =(V,R,E t ), expressed in the form of a four-tuple (s, r, o, t); Among them, V represents the entity set, R represents the relationship set, and E t represents the set of facts at discrete timestamp t; s, r, o represent the head entity, relationship, and tail entity respectively; Then, we use a recurrent neural network (RNN) to model event information and encode all historical facts into embedding representations of entities and relationships in the same vector space. We also use a stack of multi-layer relationship-aware graph convolutional neural networks to obtain the structural features in the knowledge graph and capture the dependencies between entities in the knowledge graph. Then, the knowledge graph sequence of a specific length is randomly initialized and mapped into entity sequence and relationship sequence. For the static knowledge graph slice at a specified timestamp, the tail entity of the lth layer obtains information from the head entity under the graph convolutional neural network message passing framework, sets the embedding relationship at the lth layer, and obtains the embedding representation based on structural dependency at the l+1th layer. The expression is as follows: in, Represents the entity o at time t, s at level l according to the related triple ε t The learned embedding representation, represents the evolution of the relation r at time t, represents the real number field, d represents the dimension of the vector; They represent the weight matrix used to aggregate node surrounding information and the weight matrix used to transmit node information; c o represents the normalization constant, which is equal to the in-degree of entity o, and f(·) represents the ReLU activation function; For entities that do not involve any facts, only the functions with the extra parameters are executed The relationship-aware graph convolutional neural network obtains the evolutionary representation of entity embedding according to the facts occurring at each timestamp, and realizes self-evolution with isolated entities through self-loop operations, aggregating all their static attribute information; Entity embedding is modeled by stacking the sequential patterns of the relation-aware graph convolutional neural network, and the evolutionary reasoning unit uses the gated unit to learn the temporal features. The specific expression is as follows: in, represents the Hadamard product, H t represents the cumulative representation of the entity embedding matrix by the model at time step t, Represents the model's initial response to the input at the current time step t; Candidate hidden states The calculation formula is as follows: Among them, W h represents the weight matrix, r t Represents the reset gate, X t represents the input of the current time step, b h represents the bias term, ⊙ represents element-by-element multiplication; Time Gate U t The expression for nonlinear transformation is as follows: U t =σ(W4H t-1 +b) Among them, σ(·) represents the sigmoid function, W4 represents the weight matrix of the time gate, and b represents the bias term; The sequential pattern of a relation captures information about the entities involved in the corresponding facts. The embedding at timestamp t is affected by the related entity The influence of embedding at timestamps t and t-1 is then modeled using a gated unit component to model the sequential pattern of relations. The relations are then averaged by pooling the embedding matrices of the relevant entities. The expression of the gate unit input at timestamp t is as follows: in, Represents the embedded representation of relation r in R, and is vector-connected with the output of the pooling layer; for relations that occur at timestamp t without corresponding facts The relation is embedded into the matrix R through the gated recurrent unit t-1 Update to R t , the expression is as follows: R t =GRU(R t-1 ,R t ′ ) Then, the static knowledge graph is a multi-relation graph, and each entity uses a self-looping single-layer graph convolutional neural network to aggregate all its static attribute information; the update rule definition expression of the static graph is as follows: in, and Indicates H s and H ′s The corresponding row in H s and H ′s denote the output and randomly initialized input embedding matrices respectively; Represents the relationship matrix in the graph convolutional neural network unit; Represents r in R-GCN s , Υ(·) represents the ReLu activation function, c i is a normalization constant, indicating the number of entities connected to entity i, ε s Represents the set of all relations; After obtaining the relationship matrix, the cosine similarity of the embedding is used to measure the embedding, and the angle between the evolving embedding and the static embedding of the same entity is limited to 90°. The specific definition expression is as follows: i x =min(γx,90°),x∈[0,1,…,v] Among them, the γ parameter represents the step size of the angle growth, which controls the speed of the angle increase between embeddings; θ x represents the angle between the evolving embedding and the static embedding of entity x, and v represents the total number of entities; then the static constraint component loss at timestamp t is defined as follows: Then, by optimizing the loss let and Gradually approaching in space, if the angle between them is less than θ x , the loss is 0; the final static constraint total loss expression is as follows: Finally, through multiple rounds of evolutionary reasoning units, until L st Obtain the minimum value so that the static constraints finally converge and obtain the final embedding vectors of entities and relationships at the current timestamp.

4. The method for predicting hot events based on time-varying knowledge reasoning according to claim 1 is characterized in that: The step S4 is specifically as follows: Decision label analysis is used to predict the probability of the tendency relationship between entity pairs corresponding to decision labels at several time stamps in the event time-varying knowledge subgraph. Event occurrence probability prediction is used to predict the probability of an event occurring at several time stamps by integrating the tendency relationship of several decision labels. The expressions are as follows: Among them, the entity pair corresponding to the i-th decision label of the event is defined as s i and i , the event time-varying knowledge graph is divided into k subgraphs {TG1, TG2, …, TG k }, each TG i Includes events that occurred during the time period and their related entities and relationships; About TG i The tendency relationship prediction of decision labels depends on the knowledge graph slice {G t-Δt+1 ,…,G t } information, and {G t-Δt+1 ,…,G t } is modeled as the entity embedding matrix at time t and the relation embedding matrix Then the event prediction unit uses a step-by-step iterative training strategy to perform decision label analysis and event probability prediction, as follows: First, TG1 is used as the initial training set of the model, and TG2 is used as the validation set. The model parameter learning is optimized through three types of features, namely structural dependency, temporal features, and static attributes. After the training is completed, the model is adjusted based on the performance in the validation set TG2, and the entity and relationship embedding of the TG1 subgraph is updated to predict the decision label tendency relationship and event probability at time t corresponding to the subgraph. After the training of TG1 is completed and verified, TG2 is used as a new training set to further train the model, and TG3 is used as a verification set. This process is repeated until the last subgraph TG is processed. k ; Consider entity embedding fitting and relation embedding fitting as a multi-label learning problem. and Respectively represent the label vectors of the entity and relationship fitting tasks at time t+1; each element in the vector takes the value of 1 for the occurrence of the fact, otherwise it takes the value of 0; the loss function expressions of entity embedding fitting and relationship embedding fitting are as follows: Where T represents the total number of time slices in the subgraph training set, which is the length of the divided knowledge graph sequence; and Represent the probability scores of entities and relations respectively; By setting the learnable parameters λ1 and λ2, the total loss expression of the final model is as follows: L=λ1L e +λ2L r +L st By optimizing the total loss of the model, the dynamic embedding fitting results of different entities and relationship elements in the time-varying knowledge graph under the constraints of the decision mode are obtained; Then, the dynamic embedding of entities and relations H t ,R t Input to the scoring module, given a decision label entity pair, let the head entity s traverse all possible valid relations r, where valid relations are divided into favorable relations and unfavorable relations; use these two types of relations to default predict triple sets, model their conditional probability vectors, and finally simulate the ConvTransE decoder containing a one-dimensional convolutional layer and a fully connected layer to obtain the tail entity o at time t in each relation r i The conditional probability score is calculated as follows: Then, the scores of all relations are weighted and summed according to the score ranking to obtain the comprehensive score of the decision label at time t, which reflects the possibility of forming the target label in the current relationship environment (s, o). The expression is as follows: Among them, w i In the relationship r i The weight of the entity pair (s,o); att i Represents the relationship i The weighted coefficient of the relationship type is defined as: +1 for favorable relationships, -1 for unfavorable relationships, and 0 for others; n r represents the number of possible relationships; After the scores of all decision label entity pairs are calculated, the scores of all labels are weighted and normalized to obtain the probability of the event at time t. The expression is as follows: Finally, determine whether the current time-varying subgraph has new data. If so, return to step S3 until all time-varying subgraphs of the event have been fitted by the model, output the decision label tendency trend graph and the event occurrence probability graph, and complete the hot event prediction.