Identification method, device and storage medium for event causal relationship identification model

By building a relational graph model in event causal relationship recognition and using graph convolution network for training, the problem of large error in causal relationship recognition in the existing technology is solved, and a higher accuracy event causal relationship recognition is achieved.

CN115292551BActive Publication Date: 2025-05-16HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210793375.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2025-05-16
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

The existing causal relationship recognition methods are mainly judged through the causal relationship between text instances, resulting in large errors in the identification results and cannot be applied to actual life.

Method used

By establishing instance nodes, event nodes and document nodes characterized by vectors based on the text content of the document, a relational graph model is constructed, and the relational graph model is trained using graph convolution network to update the representation vectors of each node, and then identify the causal relationship between the two event nodes.

Benefits of technology

Causal relationship recognition is realized based on events, reducing the error of causal relationship recognition, and making the causal relationship recognized with higher accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115292551B_ABST
    Figure CN115292551B_ABST
Patent Text Reader

Abstract

The present application discloses an identification method, device and storage medium for an event causal relationship identification model. The method includes: establishing instance nodes, event nodes and document nodes represented by vectors based on the text content of the document; conditionally connecting each instance node, event node and document node to construct a relationship graph model; training the relationship graph model through a graph convolutional network to update each node corresponding to each instance node, each event node and document node using other nodes connected to each instance node, each event node and document node; based on each updated node, obtaining a connection path between two event nodes through the established connection relationship between the nodes; based on the connection path between the two event nodes, calculating the causal probability between the two event nodes to identify the causal relationship between the text events corresponding to the two event nodes. In the above manner, the present application realizes causal relationship identification based on events, reducing the error of causal relationship identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data mining technology, and in particular to an identification method, device and storage medium for an event causal relationship identification model. Background Art

[0002] Causation is the relationship between cause and effect, and is an important type of relationship. If the occurrence of event A increases the probability of event B, it can be considered that there is a causal relationship between the two. Identifying the causal relationship between events in a document is conducive to logical sorting and understanding the context, which in turn facilitates appropriate predictions and decisions.

[0003] However, in the current event causal relationship recognition method, the causal relationship between different text instances is mainly used for judgment, which leads to large errors in the recognition results and cannot be applied to real life. Summary of the invention

[0004] The main technical problem solved by the present application is to provide an identification method, device and storage medium for an event causal relationship identification model, which can perform event-level causal relationship identification.

[0005] In order to solve the above technical problems, a technical solution adopted by the present application is: to provide an identification method for an event causal relationship identification model, the method comprising: establishing instance nodes, event nodes and document nodes represented by vectors based on the text content of the document; wherein the text event corresponding to the event node is a collection of text instances corresponding to each instance node pointing to the same text event, and the text content includes each text event and each text instance; connecting all instance nodes pointing to the same text event, connecting each instance node to the corresponding event node, connecting different event nodes in which instance nodes appear simultaneously in a preset text segment, and connecting all event nodes to document nodes. The relationship graph model is constructed by connecting them to each other; the relationship graph model is trained through a graph convolutional network to update each instance node, event node and document node correspondingly with other nodes connected to each instance node, each event node and document node; based on the updated instance nodes, event nodes and document nodes, a connection path between two event nodes is obtained through the established connection relationships between instance nodes, between instance nodes and corresponding event nodes, between event nodes and between event nodes and document nodes; based on the connection path between two event nodes, the causal probability between the two event nodes is calculated to identify the causal relationship between the text events corresponding to the two event nodes.

[0006] In order to solve the above technical problems, the second technical solution adopted in the present application is: to provide a computer device, which includes a processor and a memory, the memory stores a computer program, and the processor is used to execute the computer program to implement the method provided by the first technical solution of the present application as mentioned above.

[0007] In order to solve the above technical problems, the third technical solution adopted in the present application is: to provide a computer-readable storage medium, on which a computer program is stored to implement the method provided by the first technical solution of the present application.

[0008] The beneficial effects of the present application are as follows: different from the prior art, by establishing instance nodes, event nodes and document nodes represented by vectors based on the text content in the document, the instance nodes, event nodes and document nodes are conditionally connected to construct a relationship graph model, and then the relationship graph model is trained through a graph convolutional network, so that each node is updated, so that the representation vector of each node is more accurate and can better reflect the information of the node, and based on the updated nodes, the connection path between the two event nodes is obtained through the established connection relationship between the nodes, and then the causal probability between the two event nodes is calculated to identify the causal relationship between the text events corresponding to the two event nodes, so as to realize causal relationship recognition based on events, improve the traditional recognition method of using the causal relationship between text instances to represent the causal relationship between text events, reduce the error of causal relationship recognition, and make the recognized causal relationship have higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 It is a flow chart of an embodiment of an identification method of an event causal relationship identification model of the present application;

[0010] Figure 2 It is a scenario diagram of the identification method of the event causal relationship identification model of the present application;

[0011] Figure 3 is a schematic block diagram of the structure of an embodiment of a computer device of the present application;

[0012] Figure 4 It is a schematic block diagram of the structure of a computer-readable storage medium embodiment of the present application. DETAILED DESCRIPTION

[0013] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0014] With the development of Internet technology, the phenomenon of "information explosion" has emerged. A large amount of document information is generated every day, and the events described in these documents may have a certain causal relationship, such as "economic crisis" and "wage reduction", "earthquake" and "tsunami", etc. The causal relationship between these events may be obvious, or may not be judged by simple reading. In order to facilitate logical sorting and understanding of the context, many models and methods for causal relationship identification of events in documents have been spawned.

[0015] After long-term research, the inventor of the present application found that there are text events and text instances in the text content of a document composed of words. A text instance is a certain language expression of a text event. A text event can have multiple text instances, and a text event is represented by a set consisting of all its text instances. For example, the event "dog" can have text instances such as "dog", "puppy", "pooch", and "canine" in a document. These text instances point to the same event, and the set consisting of these text instances represents the text event. The current causal relationship recognition method usually uses the causal relationship judgment results between different text instances to represent the causal relationship between text events, but this is one-sided. In some cases, there is a certain difference between the causal relationship of the text instance and the causal relationship of the text event, resulting in errors in the recognition result. In order to improve or solve the above technical problems, the present application proposes at least the following embodiments.

[0016] like Figure 1 As shown, the identification method of the event causal relationship identification model of the present application may include:

[0017] S100: Establishing instance nodes, event nodes and document nodes represented by vectors based on the text content of the document.

[0018] S200: Connect all instance nodes pointing to the same text event, connect each instance node with the corresponding event node, connect different event nodes in which instance nodes appear simultaneously in a preset text segment, and connect all event nodes with document nodes to construct a relationship graph model.

[0019] S300: Train the relationship graph model through a graph convolutional network to update each instance node, event node, and document node accordingly using other nodes connected to each instance node, each event node, and document node.

[0020] S400: Based on the updated instance nodes, event nodes and document nodes, the connection path between two event nodes is obtained through the established connection relationships between instance nodes, between instance nodes and corresponding event nodes, between event nodes, and between event nodes and document nodes.

[0021] S500: Based on the connection path between the two event nodes, the causal probability between the two event nodes is calculated to identify the causal relationship between the text events corresponding to the two event nodes.

[0022] By establishing instance nodes, event nodes and document nodes represented by vectors based on the text content in the document, the instance nodes, event nodes and document nodes are conditionally connected to build a relationship graph model, and then the relationship graph model is trained through a graph convolutional network to update each node, so that the representation vector of each node is more accurate and can better reflect the information of the node, and based on the updated nodes, the connection path between the two event nodes is obtained through the established connection relationship between the nodes, and then the causal probability between the two event nodes is calculated to identify the causal relationship between the text events corresponding to the two event nodes, so as to realize causal relationship recognition based on events, improve the traditional recognition method of using the causal relationship between text instances to represent the causal relationship between text events, reduce the error of causal relationship recognition, and make the recognized causal relationship have higher accuracy.

[0023] The following is a detailed description of an embodiment of the method for identifying the event causal relationship identification model of the present application.

[0024] S100: Establishing instance nodes, event nodes and document nodes represented by vectors based on the text content of the document.

[0025] The text event corresponding to the event node is a collection of text instances corresponding to each instance node pointing to the same text event, and the text content includes each text event and each text instance.

[0026] A document is composed of many words, which are arranged in a certain sequence to form the text content of the document. The text content includes various text events and various text instances. A text instance refers to the specific text description of an event in a document, the text mention of an event in a document, and a part of the field in a document. A text event is a collection of text instances that describe the same event. For example, the field "seismic sea wave" from the 5th to the 7th word in document D is a text instance, and the field "tidal wave" from the 13th to the 14th word is a text instance. These two text instances refer to the same event "tsunami". The collection of all text instances in document D that point to the event "tsunami" is the text event corresponding to the event "tsunami".

[0027] Optionally, the text content corresponding to the text instance, the text event and the entire text content is vectorized to obtain a first initial vector representing the instance node, a second initial vector representing the event node and a third initial vector representing the document node. For details, refer to the following steps included in S100:

[0028] S110: Use a bidirectional threshold recurrent network to serialize and encode the text content to output an initial representation vector of each word that connects the forward hidden state and the backward hidden state.

[0029] The sequence encoder can serialize and encode the text content in the input document to capture the sequence characteristics between words. Specifically, it generates a corresponding representation vector based on each word in the document, that is, each word is represented by a vector.

[0030] This application uses a gated recurrent neural network (GRU) as the most basic word-level encoder. In specific applications, a vector is randomly initialized for each word to represent it. This representation vector is called a hidden state. By updating, the hidden state is made more accurate. The number of updates can be set, and each update is a time step. Assume that the tth word in the document is w t , and its corresponding hidden state is h t , at each time step, the hidden state h t Specifically, it is updated through the following formula:

[0031] Zt=σ(W z x t +U z h t-1 +b z )

[0032] r t =σ(W r x t +Ur h t-1 +b r )

[0033]

[0034]

[0035] Among them, h t-1 is the hidden state of the (t-1)th word in the document, x t It is the word w t The representation vector, z t 、r t and is the intermediate state vector in the calculation process, σ and tanh are activation functions, ⊙ is the dot product operation, and W z , W r , W h , U z , U r , U h is the parameter matrix to be learned, b z , b r , b h is the parameter vector to be learned. The parameter matrix and parameter vector will be randomly initialized and updated through training.

[0036] In a specific document, since the words are arranged in a certain sequence order, for each word, the information of the pre-order and post-order is very important for its representation vector. For example, if the document length is k, the pre-order is arranged from the first word to the kth word, and the post-order is arranged from the kth word to the first word, and the tth word w t The positions in these two arrangements will be different. Therefore, this application uses a bidirectional gated recurrent network (GRU) as a word-level encoder and updates the hidden state using the following formula:

[0037]

[0038]

[0039] in, is the forward hidden state, is the backward hidden state, k is the length of the document. Finally, the forward hidden state and the backward hidden state Connect as word w t The initial representation vector is as follows:

[0040]

[0041] Among them, the forward hidden state and the backward hidden state The connection between is the connection between vectors. For example, the connection between vectors is to splice two vectors end to end to form a vector.

[0042] S120: Perform mean processing on the initial representation vectors of all words in each text instance to obtain a first initial vector.

[0043] A text instance is a field in the text content that is composed of words and points to an event. It is a textual reference to an event in a document. Therefore, the first initial vector representing the instance node is obtained by averaging the initial representation vectors of all words that make up the text instance.

[0044] Specifically, the field consisting of the sth word to the tth word in document D is the text instance m i , then it represents the text instance m i The first initial vector of the instance node is the average value of the initial representation vectors from the sth word to the tth word, as follows:

[0045]

[0046] where x j is the initial representation vector of the jth word in the field consisting of the sth word to the tth word.

[0047] S130: Performing mean processing on the first initial vectors of all text instances in each text event to obtain a second initial vector.

[0048] A text event is a collection of text instances describing the same event, and is a collection of related fields in the text content. Therefore, by averaging the first initial vectors of all text instances in the text event, a second initial vector representing the event node is obtained.

[0049] Specifically, a text event e in document D i Is a collection of text instances pointing to the text event. The text instance set is Represents the text event e i The second initial vector of the event node Text instance set The average value of the first initial vector of the instance nodes in is as follows:

[0050]

[0051] in Text instance set Text instance m in jThe first initial vector of .

[0052] S140: Perform mean processing on the initial representation vectors of all words in the document to obtain a third initial vector.

[0053] The text content of a document is composed of words arranged in a certain sequence, and all the words in the document constitute the text content of the document. Therefore, the third initial vector representing the document node is obtained by averaging the initial representation vectors of all the words in the document.

[0054] Specifically, the document node g representing the document D contains the global information of the entire document. The third initial vector x representing the document node g is g is the average value of the initial representation vectors of all words in document D, as follows:

[0055]

[0056] Where k is the length of document D, x j is a word in document D.

[0057] S200: Connect all instance nodes pointing to the same text event, connect each instance node with the corresponding event node, connect different event nodes in which instance nodes appear simultaneously in a preset text segment, and connect all event nodes with document nodes to construct a relationship graph model.

[0058] This application constructs a relationship graph model by conditionally connecting each instance node, each event node, and document node. Two nodes are connected by edges. Due to different node types, four different types of edges are constructed: between instances, between instances and events, between events and events, and between events and documents.

[0059] To construct edges between instances, it is necessary to connect the instance nodes corresponding to all text instances pointing to the same text event. i The corresponding text instance set is but The instance nodes corresponding to all text instances in will be connected in pairs. At this time, all text instances pointing to the same text event are in a co-referential relationship.

[0060] To construct the edge between instances and events, each instance node needs to be connected to its corresponding event node. i The corresponding text event is e i , then the text instance m i The corresponding instance node and text event ei The corresponding event nodes are connected.

[0061] For the construction of edges between events, when two text instances appear in a preset text segment at the same time, and the two text instances do not point to the same text event, the event nodes corresponding to the text events pointed to by the two text instances are connected. At this time, the two text events are called co-occurrence relations. The preset text segment is any text segment in the document, and the text segments are divided by punctuation marks, which include but are not limited to periods, commas, semicolons, etc. For example, the text event e i Text instance m i , and points to the text event e j Text instance m j If both of them appear in the preset text segment P, the text event e i The corresponding event node and text event e j The corresponding event nodes are connected.

[0062] To construct the edge between an event and a document, it is necessary to connect the event nodes corresponding to all text events in the document with the document node of the document. i The corresponding event node will be connected to the document node g corresponding to document D.

[0063] Optionally, for event nodes corresponding to text events that do not have text instances in the same preset text segment, an indirect connection can be made through a document node. i With text event j No text instance appears in the same preset text segment, and the text event e i With text event j The corresponding event nodes are all connected to the document node g, then the text event e i The corresponding event node is connected to the document node g, which is connected to the text event e j The corresponding event nodes are connected to realize the text event e i Corresponding event node and text event e j The corresponding event nodes are connected indirectly.

[0064] S300: Train the relationship graph model through a graph convolutional network to update each instance node, event node, and document node accordingly using other nodes connected to each instance node, each event node, and document node.

[0065] In order to achieve richer information exchange between nodes, this application uses a graph convolutional neural network (GCN) training graph model for training. The graph convolutional neural network (GCN) training graph model can take into account different edge types and update the representation vector of the node to better model the various relationships in the graph.

[0066] Optionally, the first initial vector, the second initial vector and the third initial vector are updated layer by layer through a relational graph convolutional network using other nodes connected to each instance node, each event node and document node to obtain a first non-local feature vector representing the updated instance node, a second non-local feature vector representing the updated event node and a third non-local feature vector representing the updated document node.

[0067] Specifically, the graph convolutional neural network (GCN) training graph model is assumed to have a total of L layers. The first initial vector, the second initial vector, and the third initial vector of each instance node, each event node, and each document node will start from the first layer, and use other nodes connected to it to perform layer-by-layer forward updates to higher layers. Each layer is updated once until it is updated to the Lth layer. Finally, the vector obtained after the node is updated in the Lth layer is defined as the non-local feature vector of the node. Optionally, if there is a node n in the lth layer, the representation vector of the node n in the lth layer is Then its forward pass update at the (l+1)th layer is:

[0068]

[0069] in, is the representation vector of node n at the (l+1)th layer, R is the set of edge types, W l , W l and are the weight matrix and bias vector of the lth layer, respectively, which are randomly initialized and updated during training. is the set of nodes directly connected to node n through edges, and σ is the activation function.

[0070] When the update reaches the Lth layer, the update ends and the representation vector of each node at the Lth layer is used as the non-local feature vector of the node, that is, the representation vector of the text instance m j The first non-local eigenvector of the corresponding instance node Representing text events i The second non-local eigenvector of the corresponding event node The third non-local eigenvector of the document node g corresponding to the text content of document D

[0071] S400: Based on the updated instance nodes, event nodes and document nodes, the connection path between two event nodes is obtained through the established connection relationships between instance nodes, between instance nodes and corresponding event nodes, between event nodes, and between event nodes and document nodes.

[0072] Event nodes have established certain connections through edges. When an event node has edges connected to other event nodes, this event node can be indirectly connected to any other event node through at least one edge. This directional connection channel between event nodes is called a connection path. Any two event nodes can be connected through an intermediate node and an edge connected to the intermediate node, that is, an event node can be connected to another event node through at least one connection path.

[0073] Based on the updated instance nodes, event nodes and document nodes, a connection path between two event nodes is obtained through the established connection relationship between the nodes. For details, refer to the following steps included in S400:

[0074] S410: Calculate a combined feature vector representing the event node using the second non-local feature vector and the corresponding first non-local feature vectors.

[0075] A text event i The corresponding event node can be represented by the second non-local eigenvector representation, while textual events i Also by its corresponding text instance set , so we can represent the instance set The first non-local feature vectors of the instance nodes corresponding to the text instances in the text event are processed to represent the text event e i The corresponding event node. For details, please refer to the following steps included in S410:

[0076] S411: Calculate the weight of each text instance in the corresponding text event using the second non-local feature vector and the corresponding first non-local feature vectors based on the attention mechanism.

[0077] Text Event i By its corresponding text instance set All text instances in the text event e are composed of different text instances. i The importance of the instance set is different, so the attention mechanism can be used to Different weights are assigned to the text instances in . Specifically, the following formula is used to assign different weights to the instance set Assign weights to text instances within:

[0078]

[0079] Among them, α i,j Represents a text event i The corresponding text instance set Text instance m j The weight held, is the text instance m j The transpose of the first non-local eigenvector corresponding to m k Text instance set Any instance of text within .

[0080] S412: Perform weighted summation on each first non-local feature vector corresponding to the text event and the corresponding weight to obtain a central representation vector representing the text event centered on the text instance.

[0081] Will form the text event e i Text instance set The first non-local feature vectors corresponding to all text instances in the text instance are weighted summed with the weights assigned by the attention mechanism, and the text event e centered on the text instance is obtained. i The center representation vector of can focus more on more important text instances. Specifically, the text event e is calculated using the following formula: i The center representation vector

[0082]

[0083] S413: Connect the second non-local feature vector, the center representation vector, and the indication vector for indicating whether the text event is a main event to obtain a combined feature vector.

[0084] The main event describes the most important content in the document, which usually appears repeatedly throughout the document, and the most important event relationship in the document usually exists between the main events. This application counts the number of occurrences of the text events corresponding to each text instance based on the coreference relationship, sorts the number of occurrences, and selects the two text events with the largest number of occurrences as the main events. Optionally, two indicator vectors are randomly initialized. One of them indicates that the text event is the main event, and the other indicates that the text event is not the main event.

[0085] will represent the text event e i The second non-local eigenvector of the corresponding event node Center Representation Vector and an indication vector indicating whether the text event is a main event Connect them to get the combined feature vector to further characterize the text event e i The corresponding event nodes are specifically expressed as follows:

[0086]

[0087] By representing the text event i The vectors of the corresponding event nodes are connected into a combined feature vector, so that both the local features and non-local features of the text event are considered and reflected, and by emphasizing whether the event is the main event, the text event e i The description is more comprehensive and accurate.

[0088] S420: Calculate the connection vector of the edge between the two event nodes using the combined feature vectors corresponding to the two event nodes having a direct connection relationship, and then obtain the connection vectors of all edges corresponding to the direct connection relationship.

[0089] This application extracts an event-level graph (EG) from a graph convolutional neural network (GCN) training graph model, that is, removes the remaining nodes and leaves only the event nodes and the edges connecting the event nodes. The connection vector of the edge between the two event nodes is calculated using the combined feature vectors corresponding to the two event nodes with a direct connection relationship, and the text event e is defined. i The corresponding event node and text event e j The connection vector of the edges between the corresponding event nodes is:

[0090] e ij =σ(W e |e i -e j |+b e )

[0091] Among them, W e and b e are the weight matrix and bias vector respectively, where the weight matrix and bias vector will be randomly initialized and updated through training.

[0092] S430: Using all connection vectors to establish a connection path between two preset nodes in the event node.

[0093] In the event level graph (EG), any two event nodes are selected. These two event nodes are the preset nodes, that is, the two preset nodes are any two event nodes among all event nodes. The text events corresponding to these two preset nodes are called head events e h and tail event e t . The connection path is from the beginning event e hStarting from the corresponding event node, through the intermediate nodes and the edges connected to the intermediate nodes, until the end event e t The corresponding event node ends. For example, the head event e h The corresponding event node is connected to the text event e through the edge d The corresponding event nodes are connected, and the text event e d The corresponding event node then passes through the edge and the tail event e t The corresponding event nodes are connected, then the text event e d The corresponding event node is called the intermediate node. h The corresponding event node to the last event e t The edges connecting the corresponding event nodes form a connection path. Due to different intermediate nodes and different numbers of intermediate nodes, there can be multiple different connection paths between two preset nodes.

[0094] Optionally, a connection path between two preset nodes in the event node is established by using all connection vectors, and details may refer to the following steps included in S430:

[0095] S431: Connect the combined feature vectors of two preset nodes to obtain the node combination vectors of the two preset nodes.

[0096] Header event h With the tail event e t The node combination vectors of the two preset nodes can be obtained by connecting the combined feature vectors of the event nodes corresponding to each other. h ;e t ], this combined vector can be used to represent the relevant information of the text events specifically corresponding to the two preset nodes.

[0097] S432: Connect the connection vectors of the edges connecting the two preset nodes to the intermediate event node respectively, and calculate the edge combination vector of the corresponding connection path.

[0098] This application will be sent via text event d The head event e connected to the corresponding event node h With the tail event e t The i-th path between the corresponding event nodes The definition is as follows:

[0099]

[0100] where e hd Indicates the header event h The corresponding event node and text event e d The connection vector of the edges connecting the corresponding intermediate nodes, edt Represents a text event d The corresponding intermediate node and tail event e t The connection vector of the edges connecting the corresponding event nodes.

[0101] S433: Calculate the weight of each connection path in all connection paths between two preset nodes using the edge combination vector and the corresponding node combination vector based on the attention mechanism.

[0102] There are multiple connection paths between two preset nodes, and the importance of each path to the two preset nodes is different. Therefore, this application assigns different weights to the multiple connection paths connecting the two preset nodes through the attention mechanism, which are specifically defined as follows:

[0103]

[0104] Among them, q i Indicates the connection path The weight of the corresponding edge combination vector in all two preset nodes, For connecting paths The transpose of the corresponding edge combination vector.

[0105] S430: performing weighted summation on all connection paths and weights of two preset nodes to obtain a sum value, and performing operation on the sum value through a linear mapping function to obtain a connection path vector of the two preset nodes.

[0106] The edge combination vectors of all connection paths between two preset nodes are weighted summed with their corresponding weights, and then operated by the linear mapping function f(·) to obtain the connection path vector p of the two preset nodes. ht , the specific formula is as follows:

[0107]

[0108] S500: Based on the connection path between the two event nodes, the causal probability between the two event nodes is calculated to identify the causal relationship between the text events corresponding to the two event nodes.

[0109] The present application synthesizes the connection path vector between two event nodes and other related data to express the relationship between the two event nodes, and obtains the causal probability between the two event nodes by calculation. For details, refer to the following steps included in S500:

[0110] S510: Calculate a distance vector representing the distance between two preset nodes.

[0111] Since the relationship between two events usually exists in a shorter context in the document, the head event e h and tail event e t The relative distance δ between the corresponding event nodes ht Relationship identification is performed to strengthen the relationship expression between the two. Optionally, the present application randomly initializes a vector Δ(δ ht ) to represent the head event e h and tail event e t The relative distance between the corresponding event nodes is calculated and updated during training.

[0112] S520: Connect the combined feature vector of two preset nodes, the connection path vector, the third non-local feature vector and the distance vector to obtain a final feature vector.

[0113] The two preset nodes, namely the head event e h and tail event e t The combined feature vector of the corresponding event node [e h ;e t ], connection path vector p ht , the third non-local eigenvector h g and the distance vector Δ(δ ht ) to obtain the final feature vector f that can represent the characteristics and relationships of the two preset nodes in many aspects. ht , specifically expressed as follows:

[0114] f ht =[e h ;e t ;p ht ;h g ; Δ(δ ht )]

[0115] S530: Use the SoftMax function to calculate the causal probability of the final feature vector to obtain the causal probability between two preset nodes.

[0116] For the final eigenvector f ht , so as to obtain the causal probability between two preset nodes. Optionally, the SoftMax function is used to process the final feature vector f ht The causal probability calculation is performed, which is specifically expressed as follows:

[0117] p(c ht |e h , e t )=softmax(W f f hf +b f )

[0118] Among them, c ht is the causal probability between two preset nodes, i.e., the head event e h and tail event e t The causal probability between f and b f are the weight matrix and bias vector respectively, where the weight matrix and bias vector will be randomly initialized and updated through training.

[0119] After obtaining the causal probability between two preset nodes, it is necessary to identify the causal probability to determine whether there is a causal relationship between the text events corresponding to the two preset nodes. Optionally, by calculating the loss value between the causal probability and the true causal probability, the causal relationship identification model of this event can be further adjusted and improved. For details, refer to the following steps included in S500:

[0120] S540: Determine whether the causal probability between two preset nodes is greater than or equal to the preset probability.

[0121] If yes, there is a causal relationship between the two preset nodes. If no, there is no causal relationship between the two preset nodes.

[0122] The present application sets a preset probability as a criterion for determining whether there is a causal relationship between two preset nodes, and compares the causal probability between the two preset nodes with the preset probability. Optionally, if the causal probability between the two preset nodes is greater than or equal to the preset probability, it is considered that there is a causal relationship between the two preset nodes. If the causal probability between the two preset nodes is less than the preset probability, it is considered that there is no causal relationship between the two preset nodes. For example, if the preset probability is set to 0.5, then when the causal probability C between the two preset nodes is ht When it is greater than or equal to 0.5, it is considered that there is a causal relationship between the two preset nodes; when the causal probability c between the two preset nodes ht When it is less than 0.5, it is considered that there is no causal relationship between the two preset nodes.

[0123] S550: Calculate the loss value between the causal probability and the true causal probability to adjust the event causal relationship identification model through the loss value.

[0124] This application can be set by setting the real label To indicate the known causal relationship between the text events corresponding to the two preset nodes. Optionally, when the real label When it is "1", it indicates that there is a causal relationship between the two text events; when the true label When it is "0", it indicates that there is no causal relationship between the two text events.

[0125] By calculating the causal probability c of two preset nodes ht The true labels of the two preset nodes The loss value between and is used to judge the accuracy of this model. The smaller the loss value, the higher the accuracy, and then the causal relationship identification model of this event is adjusted. The calculation is as follows:

[0126]

[0127] in, is the overall data set, that is, all the data in this model.

[0128] The following combination Figure 2 The following is an illustrative description of the usage scenarios of the event causal relationship identification model of the present application.

[0129] Module (a) in the figure is the encoding layer, and the document will be input into the sequence encoder for processing. This application uses a gated recurrent network (GRU) to serialize and encode the words in the document and generate an initial representation vector corresponding to each word. For example, if document D contains the words "vigil", "marred", "shooting" and "shot", it will be serialized and encoded to generate the initial representation vector x corresponding to these four words. a 、x b 、x c and x d .

[0130] Module (b) in the figure is the graph network construction layer. In this module, the first initial vector representing the instance node, the second initial vector representing the event node, and the third initial vector representing the document node are calculated based on the initial representation vector of each word, and then the nodes are conditionally connected to build a relationship graph model. For example, the word "vigil" and the word "marred" are text instances m representing the events "vigil" and "destruction" respectively. a and m b ; The words "shooting" and "shot" are two text instances of the event "shooting" m c and m d , which constitutes the set of text instances corresponding to the event “shooting” First, generate a representation text instance m a 、m b 、m c and m d The first initial vector of the corresponding instance node and The events "vigil", "destruction" and "shooting" exist in document D, and the text instance m a 、m b and mc 、m d Corresponds to text event e i 、e f and e f , then the representation text event e is generated i 、e j and e f The second initial vector of the corresponding event node as well as All the words in the document constitute the text content of document D in a certain order, and the third initial vector x corresponding to the document node g is generated g Since the text instance m c and m d Corresponding to the same text event e f , then the text instance m c and m d Then connect the corresponding instance nodes. a 、m b and m c 、m d The corresponding instance nodes are respectively associated with their corresponding text events e i 、e j and e f The corresponding event nodes are connected. Finally, each event node is connected to the document node g. In this way, the relationship graph model is constructed.

[0131] Module (c) in the figure is the graph network coding layer, in which each node is updated and a combined feature vector representing the event node is generated. This application trains the graph model through a graph convolutional neural network (GCN), updates each node, and obtains the corresponding first non-local feature vector representing the instance node, the second non-local feature vector representing the event node, and the third non-local feature vector representing the document node. For example, for the first initial vector and Update and get the first non-local eigenvector and For the second initial vector as well as Update to get the second non-local eigenvector and For the third initial vector x g Update and get the third non-local eigenvector h g . Text Event f Two text instances of m c and m d Composition instance set By using the attention mechanism to Different weights are assigned to the text instances in the text, and then the representation of the text event e is obtained through calculation. f The center representation vector of the corresponding event node Then the second non-local eigenvector Center Representation Vector and the indicator vector Connect to get the combined feature vector e f .

[0132] Module (d) in the figure is the event path reasoning layer, in which the event level graph (EG) is extracted and then the path of each event node is inferred to achieve multi-hop reasoning of event relations. For example, first extract the event level graph (EG) and select the text event e i and e j For the head event h and tail event e t , head event h and tail event e t There are i paths between them. The attention mechanism is used to assign weights to these i paths, and then the connection path vector p is obtained after processing. ht , thereby generating a comprehensive representation of the head event e h and tail event e t The final eigenvector f ht , and finally the final eigenvector f ht Calculate the head event e h and tail event e t The causal probability c ht , to identify text events i and e j The causal relationship between them.

[0133] This application processes text instances, text events, and the text content of documents into nodes represented by vectors, and uses co-occurrence relationships and co-reference relationships to create connection relationships between nodes. It then uses a graph convolutional neural network (GCN) to train a graph model to aggregate information from different text instances corresponding to a text event into the text event, represents it through vectors, and then processes and calculates it to achieve causal relationship recognition and judgment based on text events, thereby improving the accuracy of event causal relationship recognition.

[0134] like Figure 3 As shown, the electronic device 10 described in the electronic device embodiment of the present application may specifically include a processor 110 and a memory 120. The memory 120 is coupled to the processor 110.

[0135] The processor 110 is used to control the operation of the electronic device 10. The processor 110 may also be referred to as a CPU (Central Processing Unit). The processor 110 may be an integrated circuit chip having signal processing capabilities. The processor 110 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0136] The memory 120 is used to store computer programs, and may be a RAM, a ROM, or other types of storage devices. Specifically, the memory 120 may include one or more computer-readable storage media, which may be non-transitory. The memory 120 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 120 is used to store at least one program code.

[0137] The processor 110 is used to execute the computer program stored in the memory 120 to implement the event causal relationship recognition model recognition method described in the embodiment of the present application.

[0138] like Figure 4 As shown, if the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium 200. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions / computer programs to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical disks, and electronic devices such as computers, mobile phones, laptops, tablet computers, cameras, etc. having the above-mentioned storage media.

[0139] The description of the execution process of the program data in the computer-readable storage medium can refer to the description in the above-mentioned embodiment of the identification method of the event causal relationship identification model of the present application, which will not be repeated here.

[0140] The above descriptions are merely embodiments of the present application and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for identifying an event causal relationship identification model, characterized in that: include: Establishing instance nodes, event nodes and document nodes represented by vectors based on the text content of the document; wherein the text event corresponding to the event node is a collection of text instances corresponding to each of the instance nodes pointing to the same text event, and the text content includes each of the text events and each of the text instances; Connecting all the instance nodes pointing to the same text event, connecting each instance node with the corresponding event node, connecting different event nodes in which the instance nodes appear simultaneously in a preset text segment, and connecting all the event nodes with the document node, so as to construct a relationship graph model; The relationship graph model is trained by a graph convolutional network to update each instance node, each event node, and each document node correspondingly using other nodes connected to each instance node, each event node, and each document node; Based on the updated instance nodes, event nodes and document nodes, a connection path between two event nodes is obtained through established connection relationships between the instance nodes, between the instance nodes and the corresponding event nodes, between the event nodes, and between the event nodes and the document nodes; Based on the connection path between the two event nodes, the causal probability between the two event nodes is calculated to identify the causal relationship between the text events corresponding to the two event nodes; The step of establishing instance nodes, event nodes, and document nodes represented by vectors based on the text content of the document includes: Vectorizing the text content corresponding to the text instance, the text event, and the entire text content to obtain a first initial vector representing the instance node, a second initial vector representing the event node, and a third initial vector representing the document node; The training of the relationship graph model by a graph convolutional network to update each instance node, each event node, and each document node correspondingly using other nodes connected to each instance node, each event node, and each document node includes: The first initial vector, the second initial vector and the third initial vector are updated layer by layer by using other nodes connected to each of the instance nodes, each of the event nodes and the document node through a relational graph convolutional network to obtain a first non-local feature vector representing the updated instance node, a second non-local feature vector representing the updated event node and a third non-local feature vector representing the updated document node; The obtaining, based on the updated instance nodes, the event nodes, and the document nodes, of a connection path between two event nodes through established connection relationships between the instance nodes, between the instance nodes and the corresponding event nodes, between the event nodes, and between the event nodes and the document nodes, comprises: Calculating a combined feature vector representing the event node using the second non-local feature vector and the corresponding first non-local feature vectors; Calculating the connection vector of the edge between the two event nodes using the combined feature vector corresponding to the two event nodes having a direct connection relationship, and then obtaining the connection vectors of all edges corresponding to the direct connection relationship; A connection path between two preset nodes in the event node is established by using all the connection vectors; wherein the two preset nodes are any two event nodes in all the event nodes.

2. The method according to claim 1, characterized in that The step of vectorizing the text content corresponding to the text instance, the text event, and the entire text content to obtain a first initial vector representing the instance node, a second initial vector representing the event node, and a third initial vector representing the document node includes: The text content is serialized and encoded using a bidirectional threshold recurrent network to output an initial representation vector of each word that connects a forward hidden state and a backward hidden state; Performing mean processing on the initial representation vectors of all the words of each of the text instances to obtain the first initial vector; Performing mean processing on the first initial vectors of all the text instances in each of the text events to obtain the second initial vector; The initial representation vectors of all the words in the document are averaged to obtain the third initial vector.

3. The method according to claim 1, characterized in that: The step of calculating a combined feature vector representing the event node by using the second non-local feature vector and the corresponding first non-local feature vectors includes: Calculate the weight of each of the text instances in the corresponding text event using the second non-local feature vector and the corresponding first non-local feature vectors based on the attention mechanism; Performing a weighted summation on each of the first non-local feature vectors corresponding to the text event and the corresponding weight to obtain a center representation vector representing the text event centered on the text instance; The second non-local feature vector, the center representation vector, and an indication vector for indicating whether the text event is a main event are connected to obtain the combined feature vector.

4. The method according to claim 1, characterized in that: The step of using all the connection vectors to establish a connection path between two preset nodes in the event node includes: Connecting the combined feature vectors of the two preset nodes to obtain a node combination vector of the two preset nodes; Connecting the connection vectors of the edges connecting the two preset nodes and the middle event node respectively, and calculating the edge combination vector of the corresponding connection path; Calculate the weight of each connection path in all the connection paths of the two preset nodes using each edge combination vector and the corresponding node combination vector based on the attention mechanism; A weighted sum is performed on all the connection paths of the two preset nodes and their weights to obtain a sum value, and the sum value is operated through a linear mapping function to obtain a connection path vector of the two preset nodes.

5. The method according to claim 4, characterized in that The calculating the causal probability between the two event nodes based on the connection path between the two event nodes comprises: Calculating a distance vector representing the distance between the two preset nodes; Connecting the combined feature vector of the two preset nodes, the connection path vector, the third non-local feature vector and the distance vector to obtain a final feature vector; The causal probability is calculated on the final feature vector using the SoftMax function to obtain the causal probability between the two preset nodes.

6. The method according to claim 5, characterized in that The identifying the causal relationship between the text events corresponding to the two event nodes includes: Determine whether the causal probability between the two preset nodes is greater than or equal to a preset probability; If so, there is a causal relationship between the two preset nodes; If not, there is no causal relationship between the two preset nodes.

7. The method according to claim 5, characterized in that After calculating the causal probability between the two event nodes based on the connection path between the two event nodes, the method further comprises: A loss value between the causal probability and the true causal probability is calculated to adjust the event causal relationship identification model according to the loss value.

8. A computer device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and the processor is used to execute the computer program to implement the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: A computer program is stored to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Network coupling time sequence information flow prediction method based on causal logic and graph convolution feature extraction

    CN112348222A