Natural disaster joint entity and relation extraction method based on multilayer deep learning framework

Through the natural disaster entity and relationship extraction method of a multi-layer deep learning framework, the Chinese version of the RoBERTa model and context dependency enhancement module are used to solve the accuracy and robustness of entity and relationship extraction in the natural disaster field, and efficient entity and relationship identification and complex relationship capture are achieved.

CN120373443APending Publication Date: 2025-07-25HUNAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510454246.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the field of natural disasters, the entity and relationship extraction methods have problems such as mispropagation, inaccurate recognition of professional nouns, and difficulty in capturing complex contexts and long-distance relationships, resulting in poor accuracy and robustness of extraction.

Method used

Using a multi-layer deep learning framework method, the Chinese version of RoBERTa model is used to perform word embedding processing, and the context dependency enhancement module is constructed. Local and long-distance dependencies are captured through BiLSTM and Self-Attention layers. Combined with the binary classifier and Biaffine model, iterative extraction of head entities and tail entities is realized, and the semantic interaction of entity pairs is enhanced by using Coordinate Attention.

Benefits of technology

It significantly improves the semantic representation ability and recognition accuracy of natural disaster entities, alleviates the problem of error propagation, improves the accuracy and robustness of entity extraction, and enhances the accuracy of relation extraction, especially in complex and long-distance relationship scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373443A_ABST
    Figure CN120373443A_ABST
Patent Text Reader

Abstract

The invention discloses a natural disaster joint entity and relationship extraction method based on a multilayer deep learning framework, and belongs to the technical field of text recognition, and the method comprises the following steps: S1, constructing a natural disaster entity and relationship extraction corpus; s2, carrying out word embedding processing by utilizing a Chinese version RoBERTa model; s3, constructing a context dependency enhancement module, and capturing local context features and long-distance dependency relationships of word vectors; s4, respectively extracting head entities and tail entities in the words and sentences by an entity extraction module; s5, an entity pair extraction module identifies a tail entity related to the head entity based on the iteration mark in the step S4 and identifies the head entity related to the tail entity to form an entity pair; s6, the relation extraction module enhances the characteristics of the entity pair through Coordinate Attention, and calculates the relation type of the entity pair through a Biaffine model; according to the method, the entity and relation extraction accuracy is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of text recognition, and particularly relates to a natural disaster joint entity and relationship extraction method based on a multi-layer deep learning framework. Background Art

[0002] Natural disasters, such as earthquakes, landslides, typhoons, etc., seriously threaten the safety of human life and property and the stable development of society. With the rapid development of the Internet, a vast amount of natural disaster information is usually published in the forms of blogs, tweets, news reports, and disaster situation announcements. These natural disaster-related text data contain a large amount of key information, including entity information such as the time, location, type, and affected objects of the disaster, as well as the relationship information between them. How to efficiently and accurately extract natural disaster-related entities and relationships from numerous unstructured texts is the key foundation for constructing a natural disaster knowledge graph and is of great significance for realizing efficient disaster relief and disaster management.

[0003] In the field of natural disasters, the entity and relationship extraction task faces higher requirements for accuracy and efficiency. The relationships between disaster entities are complex and diverse, and many relationships are long-distance relationships, that is, there are associations between entities that are far apart in the text. Modeling such long-distance relationships usually requires information extraction and association across sentences or paragraphs, further increasing the difficulty of the task.

[0004] Existing entity and relationship extraction methods are mainly oriented to the general field and have the following technical limitations in the field of natural disasters: adopting a pipeline architecture, separating entity recognition and relationship extraction for processing, unable to fully utilize the correlation between the two, easily leading to error propagation, and reducing the overall extraction performance; the field of natural disasters contains a large number of professional terms and has strong domain characteristics, making it difficult for general models to accurately identify and represent; natural disaster-related texts usually have complex contexts, and semantic dependencies may span sentences or paragraphs. Existing methods have insufficient modeling capabilities for long-distance relationships and complex contexts; existing models mostly rely on shallow features or simple attention mechanisms, making it difficult to mine the deep semantic information of texts, resulting in poor accuracy and robustness for the extraction of complex relationships.

[0005] In view of this, a natural disaster joint entity and relationship extraction method based on a multi-layer deep learning framework is designed to solve the above problems. Summary of the Invention

[0006] To solve the problems raised in the above background art, the present invention provides a natural disaster joint entity and relationship extraction method based on a multi-layer deep learning framework, which has the characteristics of alleviating the problem of error propagation and improving the accuracy of entity and relationship extraction.

[0007] To achieve the above object, the present invention provides the following technical solutions: A method for jointly extracting natural disaster entities and relationships based on a multi-layer deep learning framework, comprising the following steps:

[0008] S1: Obtain text data related to natural disasters, define target entity types and relationship types, and annotate the text data to form a natural disaster entity and relationship extraction corpus;

[0009] S2: Use the Chinese version of the RoBERTa model to perform word embedding processing on the sentence sequences of the text data in the natural disaster entity and relationship extraction corpus to obtain word vector representations;

[0010] S3: Construct a context dependence enhancement module to further capture the local context features and long-distance dependence relationships of the word vectors, fuse the captured feature information into the word vectors, and then generate three context feature representations through three different linear functions respectively;

[0011] S4: An entity extraction module, through two binary classifiers, predicts the start and end positions of entities in the input context feature vectors, and extracts the head entity and the tail entity in the sentence respectively;

[0012] S5: The entity pair extraction module identifies all tail entities related to the previously extracted head entity based on the iterative marking in step S4, and identifies all head entities related to the previously extracted tail entity to form entity pairs;

[0013] S6: The relationship extraction module retains the position information of the entity pair in the sentence through Coordinate Attention to ensure that the model can perceive the relative position of the entity pair, and then captures the semantic interaction between the entity pair and the relationship through the Biaffine model to obtain the entity pair relationship type.

[0014] Further, in the step S1, the entity types in the natural disaster entity and relationship extraction corpus include disasters, locations, times, levels, economic losses, names, depths, weather, and the number of people, and the relationship types include place of occurrence, time of occurrence, level, name, economic loss, number of affected people, number of deaths, weather, and focal depth.

[0015] Further, the specific steps of the step S2 include:

[0016] Use the Chinese version of the RoBERTa model to segment the sentences of the text data in the natural disaster entity and relationship extraction corpus to generate corresponding encodings, that is, sentence sequences, and then calculate the output vector h of each word in the sentence sequence i , to obtain the word vector representation, and the expression is:

[0017] H RoBERTa = RoBERTa(Xinput )

[0018] Where: H RoBERTa =(h1, h2, …, h t ) represents the context embedding vector output after RoBERTa encoding, RoBERTa represents the pre-trained model, X input =(x1, x2, …, x t ) represents the input text sequence, and t represents the word length.

[0019] Furthermore, the specific steps of step S3 include:

[0020] Construct a context-dependence enhancement module including a first BiLSTM layer, a Self-Attention layer, and a second BiLSTM layer, where:

[0021] The first BiLSTM layer takes the word vector representation output by the Chinese version of the RoBERTa model as input, captures the local context features of the sentence, and at the same time uses a bidirectional structure to integrate the context information, and the expression is:

[0022] H1 = BiLSTM(H RoBERTa )

[0023] Where: H1 represents the result output by BiLSTM, BiLSTM represents the bidirectional long short-term memory network, H RoBERTa =(h1, h2, …, h t ) represents the context embedding vector output after RoBERTa encoding;

[0024] The Self-Attention layer calculates the attention weights between words, assigns different weights to each word, and highlights the words related to natural disaster information, and the expression is:

[0025]

[0026] Where: Att represents the self-attention calculation result obtained through softmax, Q, K, and V represent the 3 matrices obtained by linearly transforming the result H1 output by BiLSTM respectively, and d k represents the dimension of matrix K;

[0027] The second BiLSTM layer performs secondary processing on the features after fusing self-attention, further extracts local semantic information, and enhances the sensitivity of the model to sequence patterns, and the expression is:

[0028] H2 = BiLSTM(Att + H RoBERTa )

[0029] Where: H2 represents the result output by BiLSTM, BiLSTM represents a bidirectional long short-term memory network, Att represents the self-attention calculation result obtained through softmax, and H RoBERTa =(h1, h2,..., h t ) represents the context embedding vector output after RoBERTa encoding;

[0030] The head entity, tail entity, and relationship have their own features, and three context feature representations are respectively generated through three different linear functions, which are respectively represented as H S 、H o and H r , and their expressions are:

[0031] H S =w s H2 + b s

[0032] H o =w o H2 + b o

[0033] H r =w r H2 + b r

[0034] Where: the subscripts s, o, and r respectively represent the head entity, tail entity, and relationship, w represents the weight matrix, H2 represents the feature vector, and b represents the bias vector;

[0035] Fuse the feature vectors of the head entity and tail entity with the CLS vectors of each other, and the expression is:

[0036]

[0037] Where: H' S represents the feature vector of the head entity after fusing the CLS vector of the tail entity, H S represents the feature vector of the head entity, represents the global semantic representation of the CLS vector in the input sequence for the tail entity, H' o represents the feature vector of the tail entity after fusing the CLS vector of the head entity, H o represents the feature vector of the tail entity, represents the global semantic representation of the CLS vector in the input sequence for the head entity.

[0038] Furthermore, the specific steps of the said steps S4 and S5 include:

[0039] First, extract the head entity, then extract the tail entity conditional on the head entity, and also extract the head entity conditional on the tail entity by first extracting the tail entity. The features obtained from the encoder part are shared in both directions. Since the structures in both directions are similar, the head entity is used as the base entity and the tail entity as the paired entity for description:

[0040] The head entity extraction module is based on a binary pointer network. By predicting two probability values for each position in the input sentence and using a preset threshold of 0.5 to dichotomize the result into 1 or 0, if the value at a certain position is 1, it indicates that the position is the start or end position of a head entity. The expression is:

[0041]

[0042] In the formula: and respectively represent the probabilities that the i-th position is the start and end positions of the head entity. σ(·) represents the sigmoid activation function. and respectively represent the weight matrices for the extraction of the start and end positions of the head entity. represents the input sequence of the i-th word encoding. and respectively represent the bias vectors for the extraction of the start and end positions of the head entity;

[0043] For each selected head entity, each token in the input sentence is evaluated and two probability values are assigned to determine whether the token represents the start or end position of a tail entity related to the head entity. The calculation method of the probability follows the shared dependency architecture of the model, and its probability calculation formula is as follows:

[0044]

[0045] In the formula: represents the representation of the k-th head entity, and maxpool(·) represents the max pooling operation. represents the vector representation of the token in the k-th head entity. and respectively represent the probabilities of the start token and end token of the i-th tail entity related to the k-th head entity. σ(·) represents the sigmoid activation function. and represent the weight matrices. represents the input sequence of the i-th word encoding. ○ represents the hadamard product operation. and represent the bias vectors;

[0046] Since all extraction modules in both directions work in a multi-task learning manner, each of the two extraction modules in each direction has its own loss function, and the losses are respectively denoted as L s1 and L o1 , both of which are defined using binary cross-entropy loss, and the expression is:

[0047] ce(p,t) = -[tlogp + (1 - t)log(1 - p)]

[0048]

[0049] In the formula: ce(p,t) represents the binary cross-entropy loss, p ∈ (0,1) represents the predicted probability, t represents the true label, and l represents the number of tokens in the input sentence;

[0050] When extracting the head entity conditional on the tail entity, the two losses are respectively L s2 and L o2 , and the calculation method is the same as that of the above L s1 and L o1 .

[0051] Furthermore, the specific steps of the step S6 include:

[0052] For the entity pair (s k , o j ), obtain the representation vectors of the two entities and , and the expression is:

[0053]

[0054] In the formula: maxpool(·) represents the max pooling operation, represents the vector representation of the token in the k-th topic;

[0055] Given the input Use pooling kernels of size (H,1) and (1,W) to encode the features of each channel along the horizontal coordinate direction and the vertical coordinate direction, so that each feature can capture information in different spatial dimensions, where:

[0056] The expression for horizontal dimension encoding is:

[0057]

[0058] In the formula: represents encoding the given input in the horizontal dimension;

[0059] The expression for vertical dimension encoding is:

[0060]

[0061] Wherein: is expressed as the given input encoded in the vertical dimension;

[0062] The information of different spatial dimensions captured by each feature is concatenated together, and through the shared 1×1 convolutional transformation function F1, the expression is:

[0063]

[0064] Wherein: f ∈ R C / r×(H+W) is expressed as the intermediate feature encoding the spatial information in the horizontal and vertical directions, σ is expressed as the sigmoid activation function, [z h , z w represents the concatenation operation along the spatial dimension;

[0065] Split f along the spatial dimension into two independent tensors f h ∈ R C / r×H and f w ∈ R C / r×W , and then through two 1×1 convolutions F h and F w respectively transform f h and f w into tensors with the same number of channels as the input , and the expression is:

[0066] g h = σ(F h (f h ))

[0067] g w = σ(F w (f w ))

[0068] Wherein: σ(·) is expressed as the sigmoid activation function for generating attention weights;

[0069] The final output of the coordinate attention mechanism is calculated as the product of the input feature and the attention weights along the two dimensions, and the expression is:

[0070]

[0071] Calculate the probability that the entity pair (s k , o j ) composed of two entities has the i-th relationship through the Biaffine model The expression is:

[0072]

[0073] where: σ(·) represents the sigmoid activation function, represents the feature vector of the k-th head entity, represents the feature vector of the j-th tail entity, represents the feature of the k-th head entity enhanced by the coordinate attention mechanism, represents the feature of the j-th tail entity enhanced by the coordinate attention mechanism, represents the parameter matrix of the i-th relationship;

[0074] The defined cross-entropy based loss function expression is:

[0075]

[0076] where: R represents the predefined relationship set, |R| represents the total number of relationships, ce(p,t) represents the binary cross-entropy loss, p∈(0,1) represents the prediction probability, and t represents the true label.

[0077] Compared with the prior art, the beneficial effects of the present invention are:

[0078] The present invention uses the Chinese version of the RoBERTa model to encode and represent natural disaster related texts, solves the problem of inaccurate recognition of professional terms in the field of natural disasters, significantly improves the semantic representation ability and recognition accuracy of natural disaster entities, introduces the constructed context-dependent enhancement module, which can capture the dependency information of words at different distances, further improves the representation ability of complex dependencies, introduces the entity pair extraction module based on the parallel bidirectional framework, extracts entity pairs from the head entity and the tail entity respectively, effectively alleviates the error propagation problem in traditional models, improves the accuracy and robustness of entity extraction, and proposes to use Coordinate Attention to enhance the feature representation of the head entity and the tail entity in the relationship extraction stage, extracts richer semantic information, and significantly improves the accuracy of relationship extraction BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 is the flowchart of the method of the present invention;

[0080] Figure 2 is the example diagram of natural disaster corpus annotation of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0081] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0082] As shown in the attached Figure 1 figures, the present invention provides the following technical solutions: A method for joint entity and relationship extraction of natural disasters based on a multi-layer deep learning framework, including the following steps:

[0083] S1: Obtain text data related to natural disasters, define target entity types and relationship types, and annotate the text data to form a corpus for entity and relationship extraction of natural disasters;

[0084] S2: Use the Chinese version of the RoBERTa model to perform word embedding processing on the sentence sequences of the text data in the corpus for entity and relationship extraction of natural disasters to obtain word vector representations;

[0085] S3: Construct a context-dependent enhancement module to further capture the local context features and long-distance dependence relationships of the word vectors, fuse the captured feature information into the word vectors, and then generate three context feature representations through three different linear functions respectively;

[0086] S4: An entity extraction module, through two binary classifiers, predicts the start and end positions of entities in the input context feature vectors, and extracts the head entity and the tail entity in the sentence respectively;

[0087] S5: The entity pair extraction module identifies all tail entities related to the previously extracted head entity based on the iterative labeling in step S4, and identifies all head entities related to the previously extracted tail entity to form entity pairs;

[0088] S6: The relationship extraction module retains the position information of the entity pair in the sentence through Coordinate Attention to ensure that the model can perceive the relative position of the entity pair, and then captures the semantic interaction between the entity pair and the relationship through the Biaffine model to obtain the relationship type of the entity pair.

[0089] Specifically, in step S1, the entity types in the corpus for entity and relationship extraction of natural disasters include disasters, locations, times, levels, economic losses, names, depths, weather, and the number of people, and the relationship types include place of occurrence, time of occurrence, level, name, economic loss, number of affected people, number of deaths, weather, and focal depth;

[0090] Specifically, the specific steps of step S2 include:

[0091] The sentences in the text data of the natural disaster entity and relationship extraction corpus are segmented using the Chinese version of the RoBERTa model to generate corresponding encodings, that is, word and sentence sequences, and then the output vector h of each word in the word and sentence sequence is calculated. i , obtaining the word vector representation, and the expression is:

[0092] H RoBERTa = RoBERTa(X input )

[0093] In the formula: H RoBERTa =(h1, h2,..., h t ) represents the context embedding vector output after RoBERTa encoding, RoBERTa represents the pre-trained model, and X input =(x1, x2,..., x t ) represents the input text sequence, and t represents the word length;

[0094] Specifically, the specific steps of step S3 include:

[0095] Construct a context dependence enhancement module including a first BiLSTM layer, a Self-Attention layer, and a second BiLSTM layer, where:

[0096] The first BiLSTM layer takes the word vector representation output by the Chinese version of the RoBERTa model as input, captures the local context features of the words and sentences, and at the same time uses a bidirectional structure to integrate the context information before and after, and the expression is:

[0097] H1 = BiLSTM(H RoBERTa )

[0098] In the formula: H1 represents the result output by BiLSTM, BiLSTM represents the bidirectional long short-term memory network, and H RoBERTa =(h1, h2,..., h t ) represents the context embedding vector output after RoBERTa encoding;

[0099] The Self-Attention layer calculates the attention weights between words, assigns different weights to each word, and highlights the words related to natural disaster information, and the expression is:

[0100]

[0101] In the formula: Att represents the self-attention calculation result obtained through softmax, and Q, K, and V represent the 3 matrices obtained by respectively processing the result H1 output by BiLSTM through linear transformation, and d k represents the dimension of matrix K;

[0102] The second BiLSTM layer further processes the features fused with self-attention to extract local semantic information and enhance the model's sensitivity to sequence patterns. The expression is as follows:

[0103] H2 = BiLSTM(Att + H RoBERTa )

[0104] Where: H2 represents the result output by BiLSTM, BiLSTM represents the bidirectional long short-term memory network, Att represents the self-attention calculation result obtained through softmax, and H RoBERTa =(h1, h2,..., h t ) represents the context embedding vector output after RoBERTa encoding;

[0105] The head entity, tail entity, and relation have their own features, and three context feature representations are generated through three different linear functions, which are respectively represented as H S , H o and H r , and their expressions are as follows:

[0106] H S = w s H2 + b s

[0107] H o = w o H2 + b o

[0108] H r = w r H2 + b r

[0109] Where: the subscripts s, o, and r represent the head entity, tail entity, and relation respectively, w represents the weight matrix, H2 represents the feature vector, and b represents the bias vector;

[0110] At the same time, the head entity and the tail entity are highly correlated, and the features of one entity will help improve the extraction effect of the other entity. Therefore, the CLS vector of the tail entity representation sequence is added to the head entity features to enhance the representation ability of the head entity features, and the same operation is also performed on the tail entity part, that is, the expression is changed to:

[0111]

[0112]

[0113] Where: H' S represents the head entity feature vector after fusing the tail entity CLS vector, and H S represents the feature vector of the head entity, Denoted as the global semantic representation of the tail entity by the CLS vector in the input sequence, H' o Denoted as the feature vector of the tail entity after fusing the CLS vector of the head entity, H o Denoted as the feature vector of the tail entity Denoted as the global semantic representation of the head entity by the CLS vector in the input sequence.

[0114] Specifically, the specific steps of steps S4 and S5 include:

[0115] First extract the head entity, then extract the tail entity conditional on the head entity, and first extract the tail entity, then extract the head entity conditional on the tail entity. The features obtained by the encoder part are shared in both directions;

[0116] Described with the head entity as the base entity and the tail entity as the paired entity:

[0117] The head entity extraction module is based on a binary pointer network. By predicting two probability values for each position in the input sentence and using a preset threshold of 0.5 to dichotomize the result into 1 or 0, if the value of a certain position is 1, it means that this position is the start or end position of a head entity. The expression is:

[0118]

[0119] In the formula: and respectively represent the probabilities that the i-th position is the start and end positions of the head entity, and σ(·) represents the sigmoid activation function. and respectively represent the weight matrices for the extraction of the start and end positions of the head entity. represents the input sequence encoding of the i-th word of and respectively represent the bias vectors for the extraction of the start and end positions of the head entity;

[0120] For each selected head entity, each token in the input sentence is evaluated and two probability values are assigned to determine whether the token represents the start or end position of the tail entity related to the head entity. The calculation method of the probability follows the shared dependency architecture of the model. The probability calculation formula is as follows:

[0121]

[0122]

[0123] In the formula: represents the representation of the k-th head entity, and maxpool(·) represents the max pooling operation. Denoted as the vector representation of the token in the $k$-th head entity, and respectively denote the probabilities of the start and end tokens of the $i$-th tail entity related to the $k$-th head entity. $\sigma(\cdot)$ denotes the sigmoid activation function, and denote the weight matrix, denotes the input sequence encoded for the $i$-th word of denotes the hadamard product operation, and denote the bias vector;

[0124] Described with the tail entity as the base entity and the head entity as the paired entity:

[0125] The tail entity extraction module is based on a binary pointer network. By predicting two probability values for each position in the input sentence and using a preset threshold of 0.5 to dichotomize the result into 1 or 0, if the value at a certain position is 1, it indicates that the position is the start or end position of a tail entity. The expression is:

[0126]

[0127] Where: and respectively denote the probabilities that the $i$-th position is the start and end positions of the tail entity. $\sigma(\cdot)$ denotes the sigmoid activation function, and respectively denote the weight matrices for the extraction of the start and end positions of the head entity, denotes the input sequence encoded for the $i$-th word of and respectively denote the bias vectors for the extraction of the start and end positions of the tail entity;

[0128] For each selected tail entity, each token in the input sentence is evaluated and assigned two probability values to determine whether the token represents the start or end position of the head entity related to the tail entity. The calculation method of the probability follows the shared dependency architecture of the model, and its probability calculation formula is as follows:

[0129]

[0130]

[0131] Where: denotes the representation of the $k$-th tail entity, and $\maxpool(\cdot)$ denotes the max pooling operation, Denoted as the vector representation of the tokens in the k-th tail entity, and Denoted as the probabilities of the start and end tokens of the i-th head entity related to the k-th tail entity respectively, σ(·) is denoted as the sigmoid activation function, and Denoted as the weight matrix, Denoted as the input sequence The encoding of the i-th word of the input sequence, o is denoted as the hadamard product operation, and Denoted as the bias vector;

[0132] Since all extraction modules in both directions work in a multi-task learning manner, each of the two extraction modules in each direction has its loss function, and the losses are denoted as L s1 and L o1 respectively, and both are defined using the loss based on binary cross-entropy, and the expression is:

[0133] ce(p,t) = -[tlogp + (1 - t)log(1 - p)]

[0134]

[0135] where: ce(p,t) is denoted as the binary cross-entropy loss, p ∈ (0,1) is denoted as the predicted probability, t is denoted as the true label, and l is denoted as the number of tokens in the input sentence;

[0136] In extracting the head entity conditional on the tail entity, the two losses are L s2 and L o2 respectively, and the calculation method is the same as that of the above L s1 and L o1 ;

[0137] Specifically, the specific steps of step S6 include:

[0138] For the entity pair (s k , o j ), obtain the representation vectors of the two entities and The expression is:

[0139]

[0140] where: maxpool(·) is denoted as the max pooling operation, Denoted as the vector representation of the tokens in the k-th topic;

[0141] Given the input Encode the features of each channel along the horizontal and vertical coordinate directions using pooling kernels of size (H,1) and (1,W), so that each feature can capture information in different spatial dimensions, where:

[0142] The horizontal dimension encoding expression is:

[0143]

[0144] In the formula: denotes the given input encoded in the horizontal dimension;

[0145] The vertical dimension encoding expression is:

[0146]

[0147] In the formula: denotes the given input encoded in the vertical dimension;

[0148] Concatenate the information in different spatial dimensions captured by each feature, and pass it through the shared 1×1 convolution transformation function F1, the expression is:

[0149]

[0150] In the formula: f∈R C / r×(H+W) denotes the intermediate feature encoding the spatial information in the horizontal and vertical directions, σ denotes the sigmoid activation function, [z h ,z w denotes the concatenation operation along the spatial dimension;

[0151] Split f into two independent tensors f h ∈R C / r×H and f w ∈R C / r×W along the spatial dimension, and then transform f h and f w into tensors with the same number of channels as the input h and f w respectively through two 1×1 convolutions F , the expression is:

[0152] g h =σ(F h (f h ))

[0153] g w =σ(F w (f w ))

[0154] where: σ(·) represents the sigmoid activation function used to generate attention weights;

[0155] The final output of the coordinate attention mechanism is calculated as the product of the input features and the attention weights along two dimensions, and the expression is:

[0156]

[0157] The probability of the entity pair (s k , o j ) composed of two entities having the i-th relationship is calculated through the Biaffine model The expression is:

[0158]

[0159] where: σ(·) represents the sigmoid activation function, represents the feature vector of the k-th head entity, represents the feature vector of the j-th tail entity, represents the feature of the k-th head entity enhanced by the coordinate attention mechanism, represents the feature of the j-th tail entity enhanced by the coordinate attention mechanism, represents the parameter matrix of the i-th relationship;

[0160] The defined cross-entropy-based loss function expression is:

[0161]

[0162] where: R represents the predefined relationship set, |R| represents the total number of relationships, ce(p, t) represents the binary cross-entropy loss, p ∈ (0, 1) represents the predicted probability, and t represents the true label.

[0163] Experimental demonstration

[0164] Experimental dataset:

[0165] The datasets used in this experiment include DuIE and NDERECorpus. Among them, NDERECorpus represents a self-built corpus for natural disaster entity and relationship extraction;

[0166] DuIE is an open-source dataset of the 2019 Baidu Information Extraction Competition, which is widely used to evaluate the performance of entity and relationship extraction tasks. It contains 49 relationship types, 239,076 entities, and 194,642 Chinese sentences. The dataset is split in a ratio of 7:2:1 to generate 136,249 training sentences, 38,928 validation sentences, and 19,465 test sentences;

[0167] The NDERECorpus, namely the Natural Disaster Entity and Relationship Extraction Corpus, collects natural disaster-related texts from news media and social media. After screening the obtained original texts, the content closely related to natural disasters is selected, and the entity types and relationship types to be extracted are defined. According to these definitions, the screened texts are manually annotated, and the annotation results are stored in JSON format, finally constructing the Natural Disaster Entity and Relationship Extraction Corpus;

[0168] The above annotation process takes natural language sentences and phrases as units, and annotates the word units in the sentences and phrases. The annotation examples are as shown in the appendix Figure 2 as follows;

[0169] The NDERECorpus contains 3,570 sentences and phrases, 14,800 relationship triples. The entity types are defined as 9 categories, namely disasters, locations, times, levels, economic losses, names, depths, weather, and the number of people. The relationship types are divided into 9 categories, including the place of occurrence, time of occurrence, level, name, economic loss, number of affected people, number of deaths, weather, and focal depth;

[0170] The NDERECorpus is divided into 2,499 training sentences and phrases, 714 validation sentences and phrases, and 357 test sentences and phrases according to the ratio of 7:2:1;

[0171] Experimental settings:

[0172] Three commonly used evaluation metrics in the field of natural language processing are used to evaluate the model, including accuracy (P), recall (R), and F1 value. The definitions of each evaluation metric are as follows:

[0173]

[0174] In the formula: T P represents the number of positive class samples judged as positive class, F P represents the number of negative class samples judged as positive class, F N represents the number of positive class samples judged as negative class;

[0175] Whether the prediction result is correct is evaluated for triples. Only when the head entity, relationship, and tail entity in the triple are all correctly predicted, the prediction result of the triple is regarded as correct;

[0176] The model runs on a server configured with an RTX 3060 GPU (12GB video memory) and a Windows 10 operating system. The system environment includes Python 3.11 and PyTorch 2.4.0;

[0177] During the training process of the model, combined with the experiments on the NDERECorpus, the optimal parameter configuration was finally determined. The final parameter settings of the model are shown in the following table:

[0178] Table 1: Model Parameter Settings

[0179]

[0180] Comparative Experiment:

[0181] The model of this application was experimentally compared with the CasRel (Wei, et al. ACL 2020), ETL-Span (Yu, et al. ECAI 2020), SpERT (Eberts and Ulges. ECAI 2020) and BiRTE (Ren, et al. WSDM 2022) models on the NDERECorpus and DuIE datasets. The experimental results are shown in the following table:

[0182] Table 2: Experimental Results of NDERECorpus and DuIE Datasets (%)

[0183]

[0184] The experimental results show that on the NDERECorpus dataset, the F1 value of the model of this application is increased by 2.88% compared with the CasRel model, 2.49% compared with the ETL-Span model, 2.16% compared with the SpERT model, and 2.40% compared with the BiRTE model, showing good results; on the DuIE dataset, the F1 value is increased by 1.22% compared with the CasRel model, 1.07% compared with the ETL-Span model, 2.72% compared with the SpERT model, and 1.06% compared with the BiRTE model, still maintaining the best level; the model of this application achieves a good balance between precision and recall, can significantly improve the recall rate while maintaining a high accuracy, and shows good robustness, proving its good effect in the entity and relationship extraction tasks of natural disaster texts and can effectively solve the problem of complex relationship extraction;

[0185] Comparative Experiment on Complex Relationship Extraction:

[0186] The sentences and phrases in the NDERECorpus dataset were divided into four categories according to the number of triples they contain, corresponding to the cases where the sentences and phrases contain 1, 2, 3, and 4 or more triples respectively. The ability of the model of this application to extract multiple triples from a single sentence and phrase was verified. The experimental results are shown in the following table:

[0187] Table 3: F1 Values (%) of Extracting Relational Triples from Sentences and Phrases with Different Numbers of Triples in NDERECorpus

[0188]

[0189] The results show that, compared with the baseline, the model of the present application has the highest F1 score in all categories, including more complex scenarios (such as N = 2, N = 3 or N ≥ 4), which fully demonstrates that the model of the present application has excellent ability to process triple extraction in complex scenarios.

[0190] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for jointly extracting natural disaster entities and relationships based on a multi-layer deep learning framework, characterized in that It includes the following steps: S1: Obtain text data related to natural disasters, define target entity types and relationship types, annotate the text data to form a corpus for natural disaster entity and relationship extraction; S2: Use the Chinese version of the RoBERTa model to perform word embedding processing on the sentence sequences of the text data in the natural disaster entity and relationship extraction corpus to obtain word vector representations; S3: Construct a context dependence enhancement module to further capture the local context features and long-distance dependence relationships of the word vectors, fuse the captured feature information into the word vectors, and then generate three context feature representations through three different linear functions respectively; S4: An entity extraction module, through two binary classifiers, predicts the start and end positions of entities in the input context feature vector, and extracts the head entity and tail entity in the sentence respectively; S5: The entity pair extraction module identifies all tail entities related to the previously extracted head entity and all head entities related to the previously extracted tail entity based on the iterative marking in step S4 to form entity pairs; S6: The relationship extraction module retains the position information of the entity pair in the sentence through Coordinate Attention to ensure that the model can perceive the relative position of the entity pair, and then captures the semantic interaction between the entity pair and the relationship through the Biaffine model to obtain the entity pair relationship type.

2. A method for jointly extracting natural disaster entities and relationships based on a multi-layer deep learning framework according to claim 1, characterized in that: In step S1, the entity types in the natural disaster entity and relationship extraction corpus include disasters, locations, times, levels, economic losses, names, depths, weather, and the number of people, and the relationship types include place of occurrence, time of occurrence, level, name, economic loss, number of affected people, number of deaths, weather, and focal depth.

3. The method for jointly extracting natural disaster entities and relationships based on a multi-layer deep learning framework according to claim 2, characterized in that: The specific steps of step S2 include: Segment the sentences of the text data in the natural disaster entity and relationship extraction corpus using the Chinese version of the RoBERTa model to generate corresponding encodings, that is, word and sentence sequences, and then calculate the output vector h of each word in the word and sentence sequences i , and obtain the word vector representation, with the expression: H RoBERTa = RoBERTa(X input ) Where: H RoBERTa =(h1, h2, …, h t ) represents the context embedding vector output after RoBERTa encoding, RoBERTa represents the pre-trained model, X input =(x1, x2, …, x t ) represents the input text sequence, and t represents the length of the sentence or phrase.

4. A method for jointly extracting natural disaster entities and relationships based on a multi-layer deep learning framework according to claim 3, characterized in that: The specific steps of step S3 include: Construct a context dependence enhancement module including a first BiLSTM layer, a Self-Attention layer, and a second BiLSTM layer, where: The first BiLSTM layer takes the word vector representation output by the Chinese version of the RoBERTa model as input, captures the local context features of the sentence, and at the same time uses a bidirectional structure to integrate the context information, and the expression is: H1 = BiLSTM(H RoBERTa ) Where: H1 represents the result output by BiLSTM, and BiLSTM represents a bidirectional long short-term memory network, H RoBERTa = (h1, h2, …, h t ) represents the context embedding vector output after being encoded by RoBERTa; The Self-Attention layer calculates the attention weights between words, assigns different weights to each word, and highlights the words related to natural disaster information, and the expression is: Where: Att represents the self-attention calculation result obtained through softmax, and Q, K, and V represent three matrices obtained by linearly transforming the result H1 output by BiLSTM respectively, and d k represents the dimension of matrix K; The second BiLSTM layer performs secondary processing on the features after fusing self-attention to further extract local semantic information, and the expression is: H2 = BiLSTM(Att + H RoBERTa ) Where: H2 represents the result output by the BiLSTM, BiLSTM represents the bidirectional long short-term memory network, Att represents the self-attention calculation result obtained through softmax, and H RoBERTa =(h1, h2, …, h t ) represents the context embedding vector output after RoBERTa encoding; The head entity, tail entity, and relation have their respective features, and three context feature representations are generated through three different linear functions, denoted as H s , H o , and H r , and their expressions are as follows: H s = w s H2 + b s H o = w o H2 + b o H r = w r H2 + b r In the formula: the subscripts s, o, and r represent the head entity, the tail entity, and the relationship respectively, w represents the weight matrix, H2 represents the feature vector, and b represents the bias vector; Fuse the feature vectors of the head entity and the tail entity with the CLS vector of each other, and the expression is: Where: H' S represents the head entity feature vector after fusing the tail entity CLS vector, H S represents the feature vector of the head entity, represents the global semantic representation of the CLS vector in the input sequence for the tail entity, H' o represents the tail entity feature vector after fusing the head entity CLS vector, H o represents the feature vector of the tail entity, represents the global semantic representation of the CLS vector in the input sequence for the head entity.

5. A method for jointly extracting natural disaster entities and relationships based on a multi-layer deep learning framework according to claim 4, characterized in that: The specific steps of steps S4 and S5 include: First extract the head entity, then extract the tail entity conditional on the head entity, and first extract the tail entity, then extract the head entity conditional on the tail entity. The features obtained by the encoder part are shared in both directions. Since the structures in both directions are similar, the head entity is used as the basic entity and the tail entity is used as the paired entity for description: The head entity extraction module is based on a binary pointer network. By predicting two probability values for each position in the input sentence and using a preset threshold of 0.5, the result is binary classified into 1 or 0. If the value of a certain position is 1, it means that this position is the start or end position of a head entity. The expression is: Wherein: and respectively represent the probabilities that the i-th position is the start and end positions of the head entity, and σ(·) represents the sigmoid activation function, and respectively represent the weight matrices extracted for the start and end positions of the head entity, represents the input sequence encoded for the i-th word of and respectively represent the bias vectors extracted for the start and end positions of the head entity; For each selected head entity, each token in the input sentence is evaluated and assigned two probability values to determine whether this token represents the start or end position of a tail entity related to this head entity. The calculation method of the probability follows the shared dependency architecture of the model, and its probability calculation formula is as follows: Wherein: represents the representation of the k-th head entity, and maxpool(·) represents the max pooling operation, represents the vector representation of the token in the k-th head entity, and respectively represent the probabilities of the start token and the end token of the i-th tail entity related to the k-th head entity, and σ(·) represents the sigmoid activation function, and represent the weight matrices, represents the input sequence encoding of the i-th word, represents the hadamard product operation, and represent the bias vectors; Since all extraction modules in both directions work in a multi-task learning manner, each of the two extraction modules in each direction has its own loss function, and the losses are respectively denoted as L s1 and L o1 , both of which are defined using a loss based on binary cross-entropy, and the expression is: ce(p,t)=-[tlogp+(1-t)log(1-p)] In the formula: ce(p,t) represents the binary cross-entropy loss, p∈(0,1) represents the predicted probability, t represents the true label, and l represents the number of tokens in the input sentence; For extracting head entities conditional on tail entities, the two losses are \(L\) s2 and \(L\) o2 , and their calculation methods are the same as those of the above \(L\) s1 and \(L\) o1 .

6. The method for jointly extracting natural disaster entities and relationships based on a multi-layer deep learning framework according to claim 5, wherein: The specific steps of step S6 include: For entity pair (s k , o j ), obtain the representation vectors of the two entities and The expression is: where: maxpool(·) represents the max pooling operation, represents the vector representation of the tokens in the k-th topic; Given input Encode features for each channel along the horizontal and vertical coordinate directions using pooling kernels of size (H,1) and (1,W) so that each feature can capture information in different spatial dimensions, where: The expression for horizontal dimension encoding is: Wherein: is expressed as a given input is encoded in the horizontal dimension; The expression for vertical dimension encoding is: In the formula: Denoted as the given input Encoded in the vertical dimension; Connect the information of different spatial dimensions captured by each feature together through the shared 1×1 convolutional transformation function F1. The expression is: f = σ(F1([z h , z w )) where: f ∈ R C / r×(H+W) represents the intermediate feature for encoding the spatial information in the horizontal and vertical directions, σ(·) represents the sigmoid activation function, [z h , z w represents the concatenation operation along the spatial dimension; Split f into two independent tensors f along the spatial dimension h ∈R C / r×H and f w ∈R C / r×W , and then transform f h and f w into tensors with the same number of channels as the input h and f w respectively through two 1×1 convolutions F The expression is: g h = σ(F h (f h )) g w = σ(F w (f w )) In the formula: σ represents the sigmoid activation function used to generate attention weights; The final output of the coordinate attention mechanism is calculated as the product of the input feature and the attention weights along two dimensions. The expression is: Calculate the probability that the entity pair (s k , o j ) composed of two entities has the i-th relationship The expression is as follows: where: σ(·) represents the sigmoid activation function, represents the feature vector of the k-th head entity, represents the feature vector of the j-th tail entity, represents the feature of the k-th head entity enhanced by the coordinate attention mechanism, represents the feature of the j-th tail entity enhanced by the coordinate attention mechanism, represents the parameter matrix of the i-th relationship; The defined loss function based on cross-entropy is expressed as: In the formula: R represents the predefined relation set, |R| represents the total number of relations, ce(p,t) represents the binary cross-entropy loss, p∈(0,1) represents the predicted probability, and t represents the true label.