Method, system, computing device, and readable medium for predicting future events
By obtaining three elements in the quadruple of future events and using pre-trained language model and graph neural network model for vector representation fusion, the problem of insufficient semantic expression ability of the existing model is solved, and effective prediction of new events is achieved.
Patent Information
- Application Number
- CN202111641675.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-12-29
AI Technical Summary
Existing machine models used to predict event development directions, such as CyGNet and CluSTeR models, fail to effectively consider the semantic information contained in entity nodes and event edges, resulting in poor semantic expression capabilities and inability to predict new events.
By obtaining the three elements (timestamp, subject element and type element) in the quadruple of future events, the vector representation of these elements is obtained using the trained pre-trained language model and the graph neural network model, and fuses based on these vector representations to predict the object elements of future events.
It enhances semantic expression ability, can predict new events that may occur in the future, and improves the prediction accuracy of the model.
Smart Images

Figure CN114462673B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more particularly, to a method, system, computing device, and readable medium for predicting future events. Background Art
[0002] With the development of science and technology, information on the Internet is changing rapidly. In the face of the complex and ever-changing development of events, public opinion analysts need to quickly judge the subsequent development direction of hot events with the support of real-time event information, so as to give early warnings of events that may cause serious social impacts. This puts higher requirements on the analysts' ability to recognize public opinion hot events and give the development direction of events. For a long time, predicting the development direction of events has mainly been completed by professional public opinion analysts through summarizing and analyzing information on Internet hot events and combining their own professional experience. Usually, a large amount of time is spent on manually sorting and analyzing real-time event information. With the advent of the self-media era, relying entirely on manual work to complete the analysis of the subsequent development of hot events cannot quickly give early warnings of events that may cause serious social impacts, which may bring significant losses.
[0003] There are mainly two existing machine models for predicting the development direction of events: the CyGNet model and the CluSTeR model.
[0004] Among them, the CyGNet model proposes an event prediction method based on time perception with a copy-generation mechanism. This method can predict future events from the entire intention space and can select future events from historical events by learning repeated historical events. However, the CyGNet model does not consider the semantic information contained in each entity node itself, nor does it model the association between different entity nodes, and does not construct a graph structure for the historical events that have occurred to learn the semantic information in the graph structure. Therefore, the semantic expression ability of the model is poor.
[0005] The CluSTeR model decomposes the prediction of future events into two stages: clue search and temporal reasoning according to the decision dual-system theory. First, it retrieves relevant historical clue information to obtain candidate answers, and then considers the temporal information of the clues to select the optimal result from the candidate answers. However, the way the CluSTeR model retrieves relevant historical data to form candidate answers limits the answer domain of the model, that is, the prediction result of the model can only appear in known historical events and cannot predict new events. In addition, the CluSTeR model also does not consider the semantic information contained in entity nodes and event edges, so the semantic expression ability of the model is poor.
[0006] Therefore, a new type of method, system, computing device, and readable medium for predicting future events is needed to solve the above problems. Summary of the Invention
[0007] A series of simplified concepts are introduced in the Summary of the Invention section, which will be further elaborated in the Detailed Description section. The Summary of the Invention section of the present invention is not intended to attempt to define the key features and essential technical features of the claimed technical solution, let alone attempt to determine the protection scope of the claimed technical solution.
[0008] According to one aspect of the present invention, there is provided a method for predicting future events, the method comprising: obtaining three elements in a quadruple of the future event to be predicted, the three elements including a timestamp, a subject element, and a type element, wherein the object element in the quadruple is unknown; encoding the timestamp to obtain a time vector representation of the timestamp; using a trained pre-trained language model and a trained graph neural network model to respectively obtain a subject vector representation of the subject element and a type vector representation of the type element; and based on the time vector representation, the subject vector representation, and the type vector representation, obtaining a prediction result of the object element of the future event.
[0009] In one embodiment, encoding the timestamp includes: encoding the year, month, and day in the timestamp respectively to obtain vector representations of the year, month, and day; and fusing the vector representations of the year, month, and day respectively to obtain a time vector representation of the timestamp.
[0010] In one embodiment, using a trained pre-trained language model and a trained graph neural network model to obtain the subject vector representation of the subject element includes: using a trained pre-trained language model to obtain a text semantic vector of the subject element; using a trained graph neural network model to obtain a graph structure semantic vector of the subject element; and fusing the text semantic vector and the graph structure semantic vector of the subject element to obtain a subject vector representation of the subject element.
[0011] In one embodiment, using a trained pre-trained language model and a trained graph neural network model to obtain the type vector representation of the type element includes: using a trained pre-trained language model to obtain a text semantic vector of the type element; using a trained graph neural network model to obtain a graph structure semantic vector of the type element; and fusing the text semantic vector and the graph structure semantic vector of the type element to obtain a type vector representation of the type element.
[0012] In one embodiment, obtaining a prediction result of the object element of the future event based on the time vector representation, the subject vector representation, and the type vector representation includes: fusing the time vector representation, the subject vector representation, and the type vector representation based on an attention mechanism to obtain an event vector representation of the future event; and obtaining a prediction result of the object element of the future event based on the event vector representation.
[0013] In one embodiment, the method further includes: training a pre-trained model with training data including external knowledge to obtain the trained pre-trained language model.
[0014] In one embodiment, the training data including external knowledge is obtained through the following steps: using an entity linking algorithm to obtain corresponding entity nodes of the subject elements of the quadruples of stored historical events in a domain knowledge graph; using the entity nodes and the relevant knowledge of the entity nodes to form knowledge triples as the external knowledge; retrieving relevant texts from a corpus based on the knowledge triples and the triples of the historical events, where the triples of the historical events are obtained by removing the time stamps from the quadruples of the historical events; and processing the relevant texts to obtain the training data including external knowledge.
[0015] In one embodiment, the relevant knowledge of the entity node includes the knowledge within one hop of the entity node.
[0016] In one embodiment, processing the relevant texts includes: using at least one of the subject elements and the object elements included in the relevant texts corresponding to the knowledge triples and the triples of the historical events as labels of the training data.
[0017] In one embodiment, training the pre-trained model includes: training the pre-trained model to predict whether two sentences are from the same paragraph.
[0018] In one embodiment, the method further includes: using an entity linking algorithm to obtain corresponding entity nodes of the subject elements of the quadruples of stored historical events in a domain knowledge graph; using the entity nodes and the relevant knowledge of the entity nodes to form knowledge triples; removing the time stamp from the quadruples of the historical events to obtain a historical event graph of the historical events; combining the knowledge triples with the historical event graph to obtain an updated historical event graph; and using the updated historical event graph as training data to train a graph neural network model to obtain the trained graph neural network model.
[0019] According to another aspect of the present invention, there is provided a system for predicting future events, the system comprising: a processor for using one or more neural networks: obtaining three elements in a quadruple of the future event to be predicted, the three elements including a timestamp, a subject element, and a type element, wherein the object element in the quadruple is unknown; encoding the timestamp to obtain a temporal vector representation of the timestamp; using a trained pre-trained language model and a trained graph neural network model to respectively obtain a subject vector representation of the subject element and a type vector representation of the type element; obtaining a prediction result of the object element of the future event based on the temporal vector representation, the subject vector representation, and the type vector representation, and a memory for storing network parameters of the one or more neural networks.
[0020] According to yet another embodiment of the present invention, there is provided a computing device, the computing device comprising a memory and a processor, and a computer program is stored on the memory, and when the computer program is run by the processor, the processor is caused to execute the method as described above.
[0021] According to still another embodiment of the present invention, there is provided a computer-readable medium, and a computer program is stored on the computer-readable medium, and when the computer program is run, it executes the method as described above.
[0022] The method, system, computing device, and readable medium for predicting future events according to the embodiments of the present invention can obtain semantic representations of entity nodes and event types, enhance semantic expression capabilities, and can predict new events that may occur in the future. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The following drawings of the present invention are hereby incorporated as part of the present invention for understanding the present invention. Embodiments of the present invention and their descriptions are shown in the drawings to explain the principles of the present invention.
[0024] In the drawings:
[0025] Figure 1 It is a schematic structural block diagram of an electronic device for implementing a method, system, computing device, and computer-readable medium for predicting future events according to an embodiment of the present invention.
[0026] Figure 2 It is an exemplary step flowchart of a method for predicting future events according to an embodiment of the present invention.
[0027] Figure 3 It shows a schematic diagram of an exemplary knowledge graph according to an embodiment of the present invention.
[0028] Figure 4Shows a schematic diagram of an exemplary historical event graph according to an embodiment of the present invention.
[0029] Figure 5 Shows a schematic diagram of an exemplary updated historical event graph according to an embodiment of the present invention.
[0030] Figure 6 Shows a schematic structural block diagram of a system for predicting future events according to an embodiment of the present invention.
[0031] Figure 7 Shows a schematic structural block diagram of a computing device according to an embodiment of the present invention. Detailed Description of the Invention
[0032] In order to make the objectives, technical solutions, and advantages of the present invention more apparent, exemplary embodiments according to the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein. Based on the embodiments of the present invention described herein, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0033] As described above, existing models for predicting the development direction of events do not consider the semantic information contained in entity nodes and event edges, resulting in poor semantic expression ability of the models.
[0034] Therefore, in order to improve the semantic expression ability of the model and enable the model to predict new events, the present invention provides a method for predicting future events, which includes: obtaining three elements in the quadruple of the future event to be predicted, where the three elements include a timestamp, a subject element, and a type element, and the object element in the quadruple is unknown; encoding the timestamp to obtain a time vector representation of the timestamp; using a trained pre-trained language model and a trained graph neural network model to respectively obtain a subject vector representation of the subject element and a type vector representation of the type element; and based on the time vector representation, subject vector representation, and type vector representation, obtaining a prediction result for the object element of the future event.
[0035] According to the method for predicting future events of the present invention, it is possible to obtain semantic representations of entity nodes and event types, enhance the semantic expression ability, and be able to predict new events that may occur in the future.
[0036] The following describes in detail the method, system, computing device, and computer-readable medium for predicting future events according to the present invention in combination with specific embodiments.
[0037] First, refer to Figure 1 to describe an electronic device 100 for implementing a method, a system, a computing device, and a computer-readable medium for predicting future events according to an embodiment of the present invention.
[0038] In one embodiment, the electronic device 100 can be, for example, a laptop computer, a desktop computer, a tablet computer, a learning machine, a mobile device (such as a smart phone, a phone watch, etc.), an embedded computer, a tower server, a rack server, a blade server, or any other suitable electronic device.
[0039] In one embodiment, the electronic device 100 can include at least one processor 102 and at least one memory 104.
[0040] Among them, the memory 104 can be a volatile memory, such as a random access memory (RAM), a cache memory, a dynamic random access memory (DRAM) (including stacked DRAM), or a high bandwidth memory (HBM), etc., or can also be a non-volatile memory, such as a read-only memory (ROM), a flash memory, 3D Xpoint, etc. In one embodiment, some parts of the memory 104 can be volatile memory, while another part can be non-volatile memory (for example, using a two-level memory hierarchy). The memory 104 is used to store a computer program, which, when run, can implement the client functions in the embodiments of the present invention described below (implemented by the processor) and / or other desired functions.
[0041] The processor 102 can be a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor, or other processing units with data processing capabilities and / or instruction execution capabilities. The processor 102 can be communicatively coupled to any suitable number or type of components, peripherals, modules, or devices via a communication bus. In one embodiment, the communication bus can be implemented using any suitable protocol, such as a peripheral component interconnect (PCI), a rapid peripheral component interconnect (PCIe), an accelerated graphics port (AGP), a hypertransport, or any other bus or one or more point-to-point communication protocols.
[0042] The electronic device 100 can also include an input device 106 and an output device 108. Among them, the input device 106 is a device for receiving user input, which can include a keyboard, a mouse, a touchpad, a microphone, etc. In addition, the input device 106 can also be any interface for receiving information. The output device 108 can output various information (such as images or sounds) to the outside (such as the user), which can include one or more of a display, a speaker, etc. In addition, the output device 108 can also be any other device with an output function, such as a printer, etc.
[0043] The following refers to Figure 2 the exemplary step flowchart of method 200 for predicting future events according to an embodiment of the present invention. As Figure 2 shown, the method 200 for predicting future events may include the following steps:
[0044] In step S210, three elements in the quadruple of the future event to be predicted are obtained, and the three elements include a timestamp, a subject element, and a type element, wherein the object element in the quadruple is unknown.
[0045] In step S220, the timestamp is encoded to obtain the time vector representation of the timestamp.
[0046] In step S230, the subject vector representation of the subject element and the type vector representation of the type element are obtained by using the trained pre-trained language model and the trained graph neural network model respectively.
[0047] In step S240, based on the time vector representation, the subject vector representation, and the type vector representation, a prediction result of the object element of the future event is obtained.
[0048] In one embodiment, the timestamp in the quadruple may be represented in the form of year, month, and day, such as 2020-10-15, July 12, 2019, etc. Therefore, examples of the quadruple may be, for example, (July 12, 2019, LeBron, called, Anthony Davis), (July 15, 2019, LeBron, teammate, Anthony Davis), (September 27, 2020, LeBron, won, NBA Western Conference Championship), (October 12, 2020, LeBron, won, NBA Championship), (October 15, 2020, LeBron, returned, Los Angeles), etc. Taking (October 12, 2020, LeBron, won, NBA Championship) as an example, where "October 12, 2020" is the timestamp of the quadruple, "LeBron" is the subject element (also called the subject entity) of the quadruple, "won" is the type element of the quadruple, and "NBA Championship" is the object element (also called the object entity) of the quadruple. Since the present invention aims to predict future events, the object element is unknown and the object to be predicted. At this time, the quadruple can be expressed as (October 12, 2020, LeBron, won,?). Among them, the subject entity and the object entity are collectively referred to as entities.
[0049] In one embodiment, encoding the timestamp may include: encoding the year, month, and day in the timestamp respectively to obtain their respective vector representations of the year, month, and day; and fusing the respective vector representations of the year, month, and day to obtain the time vector representation of the timestamp. Among them, the respective vector representations of the year, month, and day can be expressed as the year time vector, the month time vector, and the day time vector respectively.
[0050] In one embodiment, any suitable neural network model well-known in the art can be used to encode the timestamp, such as the One-Hot encoding model, the Word2Vec model, the FastText model, the BERT model, etc., and the present invention does not limit this.
[0051] Since the year changes consistently over time, while the month and day can be fixed at 12 and 31 respectively, the present invention uses the earliest year stored in the business system as the reference year, initializes the year time vector using the vector automatic generation formula, and randomly initializes the month time vector and the day time vector to generate 12 month time vectors and 31 day time vectors, which are updated during the model training process. The process of initializing the year time vector using the vector automatic generation formula can be expressed as follows:
[0052]
[0053] Among them, year is the difference between the year of the timestamp in the quadruple and the reference year, d m is a hyperparameter of the time vector dimension that can be set. The vector values at odd positions in the time vector are calculated using the cosine function, and the vector values at even positions are calculated using the sine function.
[0054] In one embodiment, fusing the respective vector representations of the year, month, and day may include: calculating the weighted average of the year time vector, the month time vector, and the day time vector as the time vector representation of the timestamp in the quadruple. In one embodiment, the weights of the year time vector, the month time vector, and the day time vector can be reasonably set as needed, such as 0.1, 0.3, and 0.6 respectively, and the present invention does not limit this. In one embodiment, fusing the respective vector representations of the year, month, and day may also include: concatenating, adding, subtracting, etc. the year time vector, the month time vector, and the day time vector.
[0055] In one embodiment, obtaining the subject vector representation of the subject element using a trained pre-trained language model and a trained graph neural network model may include: obtaining the text semantic vector of the subject element using the trained pre-trained language model; obtaining the graph structure semantic vector of the subject element using the trained graph neural network model; and fusing the text semantic vector and the graph structure semantic vector of the subject element to obtain the subject vector representation of the subject element.
[0056] In one embodiment, fusing the text semantic vector and the graph structure semantic vector of the subject element may include: concatenating, adding, subtracting, etc. the text semantic vector and the graph structure semantic vector of the subject element.
[0057] In one embodiment, obtaining the type vector representation of the type element using a trained pre-trained language model and a trained graph neural network model may include: obtaining the text semantic vector of the type element using the trained pre-trained language model; obtaining the graph structure semantic vector of the type element using the trained graph neural network model; and fusing the text semantic vector and the graph structure semantic vector of the type element to obtain the type vector representation of the type element.
[0058] In one embodiment, the pre-trained language model may be a BERT (Bidirectional Encoder Representations from Transformers) model, and may also be an XLNet model, a ROBERTa model, an ELECTRA model, etc. The present invention does not limit this.
[0059] In one embodiment, the graph neural network model may be a CompGCN model, and may also be a GCN (Graph Convolutional Network) model, a GGNN (Gated Graph Neural Network) model, etc. The present invention does not limit this.
[0060] In one embodiment, fusing the text semantic vector and the graph structure semantic vector of the type element may include: concatenating, adding, subtracting, etc. the text semantic vector and the graph structure semantic vector of the type element.
[0061] In one embodiment, training the pre-trained language model may include: training the pre-trained model using training data containing external knowledge to obtain a trained pre-trained language model. The trained pre-trained language model is trained using training data containing external knowledge, incorporates more semantic information of entities, and can significantly improve the model's semantic expression ability for event types and entities, thereby making the model's prediction more accurate.
[0062] When a business system stores historical events, it usually adopts a graph form with timestamps. Each historical event is in the form of a quadruple, which also indicates that the stored event information is very brief, only storing the most core element information of the event, and not mentioning the relevant information of the entities involved in the event. However, this graph does not consider the semantic information contained in each entity itself, and in fact, there may be a certain connection behind two independent entities. For example, when the entity nodes are "LeBron" and "Davis", these are two independent entities, but both of these entities are NBA players; the same is true for event types. For example, "access" and "visit" are two different event types, but the semantics of these two event types are highly similar. Therefore, if only the entity nodes and event edges are simply represented by vectors, these implicit semantics will be ignored. Therefore, when building a future fact prediction model, if these implicit information can be known, it can better help the model to make predictions.
[0063] The entities that appear in historical events and their related implicit information are usually stored in a knowledge graph. Therefore, the implicit information related to the entities that appear in historical events can be found through the knowledge graph, which is also called external knowledge in this article. See Figure 3 , Figure 3 shows a schematic diagram of an exemplary knowledge graph according to an embodiment of the present invention.
[0064] In one embodiment, the training data containing external knowledge is obtained through the following steps: using an entity linking algorithm to obtain the corresponding entity nodes of the main elements of the quadruples of the stored historical events in the domain knowledge graph; using the relevant knowledge between entity nodes to form knowledge triples as external knowledge; retrieving relevant texts from the corpus based on the knowledge triples and the triples of historical events, where the triples of historical events are obtained by removing the timestamp from the quadruples of historical events; and processing the relevant texts to obtain the training data containing external knowledge.
[0065] Specifically, when obtaining the implicit knowledge of each entity in a historical event, first find the corresponding entity node of the entity in the domain knowledge graph, obtain the relevant knowledge of the entity node, and then form a knowledge triple with the relevant knowledge between entity nodes as external knowledge. In one embodiment, the relevant knowledge of the entity node includes the knowledge within one hop of the entity node. For example, Figure 3 the triple formed by the entity node "LeBron Raymone James" (abbreviated as "James" in this article) and its knowledge within one hop in can include: (James, cooperation, Anthony Davis), (James, type, basketball player), (James, affiliated unit, Lakers), (James, alma mater, St. Vincent-St. Mary High School).
[0066] However, the expression forms of entities are diverse. For example, for the entity "James", there may be aliases such as "LeBron James", "LeBron James", and "The King of Basketball". Therefore, an entity linking algorithm is needed to accurately obtain the corresponding entity nodes of entities in historical events in the domain knowledge graph.
[0067] In one embodiment, an entity linking algorithm based on the literal features of entity names can be used to obtain the corresponding entity nodes of each entity in historical events in the domain knowledge graph to ensure the operation efficiency of the entity linking algorithm. The literal features of the entity linking algorithm are divided into three parts. The first part calculates the character co-occurrence of the entity in the historical event and the entity nodes in the domain knowledge graph at the character level, and normalizes it by dividing the length of the entity in the historical event to calculate the character similarity score. The second part calculates the word co-occurrence of the entity in the historical event and the entity nodes in the domain knowledge graph, and normalizes it by dividing the number of word segments of the entity in the historical event to calculate the word similarity score. Among them, the weight of the co-occurring words can be assigned by the inverse document frequency (IDF), so as to reduce the score of high-frequency words. The third part calculates the co-occurrence of the tri-gram sequential feature segments of the entity in the historical event and the entity nodes in the domain knowledge graph, and normalizes it by dividing the number of entity tri-grams to calculate the sequential similarity score. To prevent the denominator from being 0 due to the entity length in the historical event being less than 3, the present invention sets the minimum denominator to 1. After obtaining the above three scores, the three scores can be weighted and averaged to obtain the entity linking scores of the entity in the historical event and all entity nodes in the domain knowledge graph, and sorted according to the scores. The entity node with the highest score is used as the entity node corresponding to the entity in the historical event, so as to obtain the final entity linking result.
[0068] In order to introduce the structured knowledge in the domain knowledge graph when training the pre-trained model, in one embodiment, based on the above knowledge triples as external knowledge, combined with the triples obtained by removing the timestamp from the quadruple of the historical event, a remote supervision method can be used to search for relevant texts (such as sentences or paragraphs, etc.) containing this structured knowledge in a large corpus, and then process the relevant text to obtain training data containing external knowledge.
[0069] In one embodiment, processing the relevant text may include: cleaning, preprocessing, segmenting, clause splitting, word segmentation, etc. of the relevant text. In one embodiment, processing the relevant text may also include: processing the relevant text into the input format required by the pre-trained model, for example, [CLS] Yao Ming and [MASK][MASK] are going to get married [SEP] They are very happy [SEP].
[0070] In one embodiment, processing the relevant text includes: masking at least one of the subject element and the object element corresponding to the knowledge triple and the triple of historical events included in the relevant text as the label of the training data.
[0071] In one embodiment, training the pre-trained model includes: training the pre-trained model to predict whether two sentences are from the same paragraph.
[0072] The present invention makes the following improvements in the pre-trained model:
[0073] 1) Since the Next Sentence Prediction task is too simple, the Next Sentence Prediction task used in the original pre-trained model is removed and replaced with a task of predicting whether two sentences input into the pre-trained model are obtained from the same paragraph;
[0074] 2) The pre-training of the conventional pre-trained model adopts a character-level MASK (masking) strategy, allowing the model to predict the masked characters. This method is not friendly to the pre-trained Chinese BERT model because the vast majority of Chinese words with practical meanings are words. Therefore, the present invention adopts a word-level MASK to allow the model to predict the masked words with practical semantics, so that the model incorporates more semantic knowledge;
[0075] 3) When performing word-level MASK on the pre-trained model, since the training data of the training model is obtained through distant supervision based on structured knowledge, the training data will contain entities in the structured knowledge. When performing word masking, it is ensured that at least one of the entities appearing in the structured knowledge will be masked, so that the model learns the semantics in the structured knowledge;
[0076] 4) After the pre-trained model performs multiple rounds of pre-training tasks, the pre-trained model BERT is used to continue fine-tuning training on downstream tasks such as entity recognition and event type recognition, allowing the model to predict which words are entities in the structured knowledge and which words are related. Through fine-tuning training, the pre-trained model incorporates more information in the structured knowledge.
[0077] In one embodiment, training the graph neural network model may include: removing the timestamp from the quadruple of historical events to obtain the triple of historical events, and using the triple of historical events to construct a historical event graph; combining the knowledge triple as the external knowledge with the historical event graph to obtain an updated historical event graph; and using the updated historical event graph as training data to train the graph neural network model to obtain a trained graph neural network model. Refer to Figure 4 andFigure 5 , Figure 4 shows a schematic diagram of an exemplary historical event map according to an embodiment of the present invention, Figure 5 shows a schematic diagram of an exemplary updated historical event map according to an embodiment of the present invention.
[0078] After obtaining the updated historical event map, a graph neural network model (e.g., compGCN) is used to train on the updated historical event map to obtain the graph structure semantic representations of each node of the updated historical event map. Among them, the graph neural network model is well-known in the art and will not be elaborated here. Examples of the graph structure semantics output by the graph neural network model are as follows:
[0079] James = [0.992734, -0.476647,..., 0.217249]
[0080] Call = [-0.135216, 0.156160,..., 0.001139]
[0081] Anthony Davis = [0.088582, 0.240145,..., -0.006931]
[0082] Since the graph neural network model performs semantic representation training according to the graph structure composed of the nodes and edges of the map, if two different nodes often point to the same other nodes, for example, both the node "James" and the node "Anthony Davis" point to the node "Lakers", then the semantic representations of these two nodes trained by the model will have semantic similarity, which also increases the internal connection between the node "James" and the node "Anthony Davis" from the perspective of the map structure.
[0083] In one embodiment, based on the time vector representation, the subject vector representation, and the type vector representation, obtaining the prediction result of the object element of the future event may include: fusing the time vector representation, the subject vector representation, and the type vector representation based on the attention mechanism to obtain the event vector representation of the future event; and obtaining the prediction result of the object element of the future event based on the event vector representation.
[0084] In one embodiment, fusing the time vector representation, the subject vector representation, and the type vector representation based on the attention mechanism may include: performing attention weighting and summing on the time vector representation, the subject vector representation, and the type vector to obtain the event vector representation of the future event.
[0085] In one embodiment, obtaining a prediction result of an object element of a future event based on an event vector representation may include: inputting the event vector representation into a fully connected layer neural network and a classification layer (e.g., a softmax layer) to obtain a prediction probability of the object element of the future event, and taking the result with the maximum prediction probability as the prediction result of the object element of the future event.
[0086] In another embodiment, the present invention provides a system for predicting future events. Referring to Figure 6 , Figure 6 FIG. shows a schematic structural block diagram of a system 600 for predicting future events according to an embodiment of the present invention. As Figure 6 shown, the system 600 for predicting future events may include a processor 610 and a memory 620. Among them, the processor 610 is used to implement the following steps using one or more neural networks: obtaining three elements in the quadruple of the future event to be predicted, the three elements including a timestamp, a subject element, and a type element, where the object element in the quadruple is unknown; encoding the timestamp to obtain a time vector representation of the timestamp; using a trained pre-trained language model and a trained graph neural network model to respectively obtain a subject vector representation of the subject element and a type vector representation of the type element; and obtaining a prediction result of the object element of the future event based on the time vector representation, the subject vector representation, and the type vector representation.
[0087] Exemplarily, the processor 610 may be any processing device well known in the art, such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor, a microcontroller, a field programmable gate array (FPGA), etc., and the present invention is not limited thereto.
[0088] Among them, the memory 620 is used to store network parameters of one or more neural networks. Exemplarily, the memory 620 may be RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage devices, magnetic tape cartridges, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by the processor 610.
[0089] The system 600 for predicting future events according to an embodiment of the present invention may execute the method 200 for predicting future events according to an embodiment of the present invention described above. Those skilled in the art can understand the specific implementation method of the system 600 for predicting future events according to an embodiment of the present invention in combination with the content described above. For the sake of brevity, the specific details are not described herein again.
[0090] In yet another embodiment, the present invention provides a computing device. Referring to Figure 7 , Figure 7 FIG. Figure 7 shows a schematic structural block diagram of a computing device 700 according to an embodiment of the present invention. As Figure 7 shown, the computing device 700 may include a memory 710 and a processor 720, where a computer program is stored on the memory 710, and when the computer program is run by the processor 720, the processor 720 is caused to execute the method 200 for predicting future events as described above.
[0091] Those skilled in the art can understand the specific operations of the computing device 700 according to the embodiments of the present invention in combination with the content described above. For the sake of brevity, the specific details are not described herein again, and only some main operations of the processor 720 are described as follows:
[0092] Obtain three elements in the quadruple of the future event to be predicted, where the three elements include a timestamp, a subject element, and a type element, and the object element in the quadruple is unknown;
[0093] Encode the timestamp to obtain a time vector representation of the timestamp;
[0094] Use a trained pre-trained language model and a trained graph neural network model to respectively obtain a subject vector representation of the subject element and a type vector representation of the type element; and
[0095] Based on the time vector representation, the subject vector representation, and the type vector representation, obtain a prediction result of the object element of the future event.
[0096] The computing device 700 according to the embodiment of the present invention can execute the method 200 for predicting future events according to the embodiment of the present invention described above. Those skilled in the art can understand the specific implementation method of the computing device 700 according to the embodiment of the present invention in combination with the content described above. For the sake of brevity, the specific details are not described herein again.
[0097] In another embodiment, the present invention provides a computer-readable medium having a computer program stored thereon, and the computer program, when running, executes the method 200 for predicting future events as described in the above embodiments. Any tangible, non-transitory computer-readable medium can be used, including magnetic storage devices (hard disks, floppy disks, etc.), optical storage devices (CD-ROMs, DVDs, Blu-ray discs, etc.), flash memories, and / or the like. These computer program instructions can be loaded onto a general-purpose computer, a special-purpose computer, or other programmable data processing devices to form a machine, such that the instructions executed on the computer or other programmable data processing devices can generate a device for implementing the specified functions. These computer program instructions can also be stored in a computer-readable memory, which can direct the computer or other programmable data processing devices to operate in a specific manner, so that the instructions stored in the computer-readable memory can form a manufactured article, including an implementation device for implementing the specified functions. The computer program instructions can also be loaded onto a computer or other programmable data processing devices, thereby performing a series of operational steps on the computer or other programmable devices to generate a computer-implemented process, such that the instructions executed on the computer or other programmable devices can provide steps for implementing the specified functions.
[0098] The beneficial effects of the present invention are as follows:
[0099] (1) The present invention obtains the semantic representations of entity nodes and event types, enhances the semantic expression ability, and can predict new events that may occur in the future.
[0100] (2) The present invention uses an entity linking algorithm to introduce external knowledge into the historical event graph, and models the updated graph structure through a graph neural network model to obtain the graph structure semantic vectors of each node with semantic information. It also integrates the structured graph knowledge into the pre-trained model through the methods of distant supervision and model pre-training, improving the model's representation ability for domain data.
[0101] (2) The present invention proposes a strategy of obtaining semantic vector representations for years, months, and days respectively and then performing weighting, which makes the expression of timestamps more accurate.
[0102] (4) The present invention adopts a future fact prediction method that combines time vectors, text semantic vectors, and graph structure semantic vectors, forming a complete set of data processing and modeling solutions to achieve auxiliary decision-making.
[0103] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present invention thereto. Those of ordinary skill in the art can make various changes and modifications therein without departing from the scope and spirit of the present invention. All such changes and modifications are intended to be included within the scope of the present invention as claimed in the appended claims.
[0104] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0105] Similarly, it should be understood that, in order to streamline the present invention and assist in understanding one or more of the various inventive aspects, in the description of the exemplary embodiments of the present invention, various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present invention should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected by the corresponding claims, the inventive point lies in that the corresponding technical problem can be solved with features less than all the features of a single disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim itself serves as a separate embodiment of the present invention.
[0106] Those skilled in the art will appreciate that, except where features are mutually exclusive, any combination can be used of all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or apparatus so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature providing the same, equivalent, or similar purpose.
[0107] In addition, those skilled in the art will be able to understand that, although some of the embodiments described herein include certain features included in other embodiments but not others, the combination of features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.
[0108] It should be noted that the above embodiments are illustrative of the present invention rather than restrictive of the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names.
[0109] As described above, this is only a specific embodiment of the present invention or an illustration of the specific embodiment. The protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for predicting future events, characterized in that, The method includes: Obtaining three elements in the quadruple of the future event to be predicted, where the three elements include a timestamp, a subject element, and a type element, and the object element in the quadruple is unknown; Encoding the timestamp to obtain a temporal vector representation of the timestamp; Using a trained pre-trained language model and a trained graph neural network model to respectively obtain a subject vector representation of the subject element and a type vector representation of the type element; and Based on the temporal vector representation, the subject vector representation, and the type vector representation, obtaining a prediction result for the object element of the future event; where obtaining a prediction result for the object element of the future event based on the temporal vector representation, the subject vector representation, and the type vector representation includes: Fusing the temporal vector representation, the subject vector representation, and the type vector representation based on an attention mechanism to obtain an event vector representation of the future event; and Based on the event vector representation, obtaining a prediction result for the object element of the future event.
2. The method according to claim 1, wherein Where encoding the timestamp includes: Encoding the year, month, and day in the timestamp respectively to obtain vector representations of the year, month, and day; and Fusing the vector representations of the year, month, and day respectively to obtain a temporal vector representation of the timestamp.
3. The method according to claim 1, characterized in that, Where using a trained pre-trained language model and a trained graph neural network model to obtain a subject vector representation of the subject element includes: Using a trained pre-trained language model to obtain a text semantic vector of the subject element; Using a trained graph neural network model to obtain a graph structure semantic vector of the subject element; and Fusing the text semantic vector and the graph structure semantic vector of the subject element to obtain a subject vector representation of the subject element.
4. The method according to claim 1, wherein Where using a trained pre-trained language model and a trained graph neural network model to obtain a type vector representation of the type element includes: Using a trained pre-trained language model to obtain a text semantic vector of the type element; Using a trained graph neural network model to obtain a graph structure semantic vector of the type element; and Fusing the text semantic vector and the graph structure semantic vector of the type element to obtain a type vector representation of the type element.
5. The method according to claim 1, characterized in that, The method further includes: training a pre-trained model using training data containing external knowledge to obtain the trained pre-trained language model.
6. The method according to claim 5, wherein The training data containing external knowledge is obtained through the following steps: Using an entity linking algorithm to obtain corresponding entity nodes of the subject elements in the quadruples of stored historical events in the domain knowledge graph; Using the entity nodes and the relevant knowledge of the entity nodes to form knowledge triples as the external knowledge; Retrieving relevant texts from a corpus based on the knowledge triples and the triples of the historical events, where the triples of the historical events are obtained by removing the timestamp from the quadruples of the historical events; And Processing the relevant texts to obtain the training data containing external knowledge.
7. The method according to claim 6, characterized in that, Where the relevant knowledge of the entity nodes includes the knowledge within one hop of the entity nodes.
8. The method according to claim 6, wherein Among them, processing the relevant text includes: Taking at least one of the subject elements and object elements corresponding to the knowledge triple and the triple of the historical event included in the relevant text as the label of the training data.
9. The method according to claim 5, wherein Among them, training the pre-trained model includes: training the pre-trained model to predict whether two sentences come from the same paragraph.
10. The method according to claim 1, wherein The method further includes: Using an entity linking algorithm to obtain the corresponding entity node of the subject element of the quadruple of the stored historical event in the domain knowledge graph; Using the relevant knowledge of the entity node and the entity node to form a knowledge triple; Removing the time stamp from the quadruple of the historical event to obtain the historical event graph of the historical event; Combining the knowledge triple with the historical event graph to obtain an updated historical event graph; and Using the updated historical event graph as training data to train a graph neural network model to obtain the trained graph neural network model.
11. A system for predicting future events, characterized in that, The system includes: A processor for using one or more neural networks: Obtaining three elements in the quadruple of the future event to be predicted, the three elements including a time stamp, a subject element, and a type element, where the object element in the quadruple is unknown; Encoding the time stamp to obtain a time vector representation of the time stamp; Using the trained pre-trained language model and the trained graph neural network model to respectively obtain a subject vector representation of the subject element and a type vector representation of the type element; Based on the time vector representation, the subject vector representation, and the type vector representation, obtaining a prediction result of the object element of the future event, where, based on the time vector representation, the subject vector representation, and the type vector representation, obtaining a prediction result of the object element of the future event includes: Fusing the time vector representation, the subject vector representation, and the type vector representation based on an attention mechanism to obtain an event vector representation of the future event; and Obtaining a prediction result of the object element of the future event based on the event vector representation; A memory for storing network parameters of the one or more neural networks.
12. A computing device, characterized in that, The computing device includes a memory and a processor, and a computer program is stored on the memory. When the computer program is run by the processor, the processor executes the method according to any one of claims 1-10.
13. A computer-readable medium, characterized in that, A computer program is stored on the computer-readable medium. When the computer program is run, it executes the method according to any one of claims 1-10.
Citation Information
Patent Citations
Intelligent workshop production optimization method based on dynamic bottleneck prediction
CN110163436A
Video system using dual stage attention based recurrent neural network for future event prediction
US20180060666A1