Entity Relationship Extraction Method and Its Application Integrating BERT Network and Location Feature Information
By integrating BERT network and location feature information, combined with Bi-LSTM and Attention layers, the intrinsic relationship ignorance problem of entity recognition and relationship classification in the traditional entity relationship extraction method is solved, and high-precision entity extraction and relationship recognition are achieved.
Patent Information
- Application Number
- CN202210791774.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-07
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-07-07
AI Technical Summary
The traditional entity relationship extraction method ignores the inherent connection between entity recognition and relationship classification, has problems of error propagation and information redundancy, and it is difficult to effectively solve the problem of overlapping entity relationships.
The entity relationship extraction method that integrates the BERT network and location feature information is adopted, context semantic information is captured through the BERT model, combined with Bi-LSTM, Attention layer and fully connected layer, the feature representation of the entity is extracted, and the attention mechanism is enhanced through the location feature information to achieve accurate identification and classification of entity relationships.
In complex fields and small samples, high-precision entity extraction is achieved, effectively solving the problem of overlapping entity relationships and improving the accuracy of entity classification.
Smart Images

Figure CN115203434B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for entity relation extraction integrating BERT network and position feature information and its application, belonging to the field of knowledge extraction. Background Art
[0002] With the multi-field development of knowledge graphs, through the analysis and mining of a large amount of heterogeneous data, the comprehensive processing and utilization capabilities of data in most vertical fields can be greatly improved, and entity relation extraction is an important step in constructing large-scale domain knowledge graphs. The big data era has swept in, and knowledge graphs must also integrate complex data. Traditional entity relation extraction methods mainly ignore the internal connection between entity recognition and relation classification, and have problems such as error propagation and information redundancy, and cannot effectively solve the problem of entity relation overlap. To solve related problems, deep learning methods are introduced. Deep learning methods can alleviate the disadvantages of relation extraction models based on traditional features to a certain extent, with less cumulative error. However, in many fields, due to the different scales of data volume, the direct application of deep learning methods to entity relation extraction has many limitations, and it is urgent to improve the relevant network models to solve the entity relation extraction problem in specific application fields. Summary of the Invention
[0003] To solve the above existing problems, the present invention provides a method for entity relation extraction integrating BERT network and position feature information and its application.
[0004] The purpose of the present invention is achieved by the following technical solutions:
[0005] A method for entity relation extraction integrating BERT network and position feature information, the steps of which are:
[0006] 1) Entity data collection: Obtain a publicly available dataset in a professional vertical field, and collect samples according to BIO annotation;
[0007] 2) Data processing: Divide the obtained entity samples into a training set, a validation set, and a test set, and perform maximum-minimum normalization processing;
[0008] 3) Propose a model structure: The proposed entity relation extraction model integrating BERT network and position feature information is composed of a BERT model, a Bi-LSTM, a linear layer, an Attention layer, and a fully connected layer; First, the BERT layer captures stronger context semantic information. Secondly, the complete time-step hidden state sequence is concatenated through Bi-LSTM and mapped to the linear layer to obtain the scores of each category label. Then, through the Attention layer, weights are assigned to select specific information and concatenate it with the position embedding matrix P in the fully connected layer to distinguish the feature representations of the same entity in different relations. Finally, entity classification prediction is completed through the entity classification matrix C;
[0009] 4) Offline training: Use the training set and regularization strategy to train the model and save the optimal parameters of the entity classification matrix C;
[0010] 5) Online testing: Apply the test set to verify the model performance or load the pre-trained parameters to fine-tune the entire model, and use parameter sharing and transfer learning to achieve timely training of the model.
[0011] In step 3) mentioned above, the specific method is as follows:
[0012] 3.1) The data passes through the BERT model and Bi-LSTM to obtain the scores of each category label corresponding to the word level:
[0013] Take BERT as the encoder of the input text sequence, and use BERT to obtain the hidden layer state vector X t , as shown in formula (1) specifically:
[0014] X t = Bert base (w t ) (1)
[0015] Record it as sequence X = (x 1 , x 2 , …, x n );
[0016] Take the sequence X obtained by the BERT network as the input of each time step of the bidirectional long short-term memory network Bi-LSTM, and obtain the forward hidden state sequence and the backward hidden state sequence where h t Introduce corresponding memory units in the hidden layer, as shown in formula (2);
[0017]
[0018] Among them, h t-1 is the output result of the hidden layer of the previous long short-term memory network unit; C t-1 is the state result of the previous long short-term memory network unit; x t is the word vector input result of this article; f t is the output result of the forgetting gate, where σ is the sigmoid activation function; i t and are the output results of the input gate; tanh is the tanh activation function; Ο t is the output result of the output gate; C t is the state value of the current unit; h t is the output of the hidden layer of the current unit;
[0019] Concatenate the forward and backward hidden state sequences along the time steps to obtain the complete hidden state sequence, denoted as H = (h 1 , h 2 , …, h n );
[0020] Finally, map the hidden state sequence to s dimensions through a linear layer to obtain the mapped sequence denoted as L = (l 1 , l 2 , …, l n ); where L represents the score of y i corresponding to each category label of word X j ;
[0021] 3.2) Propose the relational position feature attention mechanism: Input the scores of each category label into the relational position feature attention mechanism to obtain the entity classification matrix C;
[0022] Use the attention mechanism QKV model to calculate the weight values of the entity relationship classification results; Obtain the Query matrix through a vector matrix query k*1 sampled randomly by uniform distribution, where k is the output vector dimension of the hidden layer of the bidirectional long short-term memory network; Obtain the Key matrix through the feature matrix generated by the word vectors of the Chinese word segmentation in the sentence; Obtain the Value matrix through the matrix composed of the output vectors of the hidden layer of the bidirectional long short-term memory network;
[0023] The weight values of the attention mechanism in entity relationship extraction are calculated according to formula (3):
[0024] Attention_w n×1 = softmax(key_w n×k * query_w k×1 ) (3)
[0025] where the softmax function is used for vector normalization; key_w n×k is the Key vector matrix of the attention mechanism; query_w k×1 is the Query vector matrix in the attention mechanism; Attention_w n×1 is the weight value of the attention mechanism;
[0026] The output matrix of the attention mechanism in entity relationship extraction is calculated according to formula (4):
[0027] Attention_r k×1 = (Attention_w T * value_w n×k ) T (4)
[0028] Among them, value_w n×k is the output vector matrix of the hidden layer of the bidirectional long short-term memory network; Attentiont_r k×1 is the output vector matrix of the attention mechanism;
[0029] In addition, establish the relationship between the attention output sequence and the output sequence passing through BERT, distinguish the feature representations of the same entity in different relationships, and make the final relationship classification prediction;
[0030] Calculate the distance between each word in the output sequence of BERT and the trigger word of the current attention output sequence, then randomly initialize the position embedding matrix P according to the maximum sentence length m and the position feature size n, and obtain the relationship position feature Pf of each word by querying the position embedding matrix P t , and the relationship position feature is calculated according to formula (5):
[0031] Pf t = P Pr-Pw (5)
[0032] Among them, Pr represents the position of the relationship trigger word; Pw represents the position of the word in the BERT output sequence;
[0033] Finally, splice the position embedding matrix and the attention mechanism output matrix in the fully connected layer to obtain the entity classification matrix (6):
[0034] C = f(Attention_r k×1 ; P) (6)
[0035] Among them, f(·) represents the fully connected layer; Attentiont_r k×1 is the output vector matrix of the attention mechanism; P is the position embedding matrix, and C is the current entity classification matrix.
[0036] Application of the entity relationship extraction method integrating BERT network and position feature information in medicine: Input CNMER data, and the specific application method is: Use the entity relationship extraction method integrating BERT network and position feature information to perform entity extraction on CNMER data, complete entity classification such as disease signs and disease treatments, and improve the accuracy of medical entity classification.
[0037] Application of the entity relation extraction method integrating BERT network and location feature information in the military: Input AIR FORCE MIL-HDBK-310-1997 data, and the specific application method is as follows: Use the entity relation extraction method integrating BERT network and location feature information to perform entity extraction on the AIR FORCE MIL-HDBK-310-1997 data, complete the classification of different climate entities for the development of military products, solve the problem of difficult relation extraction caused by overlapping entity relations in the climate field of military product development, and predict the climate suitable for the development of military products.
[0038] Application of the entity relation extraction method integrating BERT network and location feature information in the finance field: Input LendingClub data, and the specific application method is as follows: Use the entity relation extraction method integrating BERT network and location feature information to perform entity extraction on the LendingClub data, complete the entity classification of loan customers, loan services, and loan default factors, comprehensively understand the development trend of loan financial events, and predict the development law of the loan financial market.
[0039] Application of the entity relation extraction method integrating BERT network and location feature information in the legal field: Use the entity relation extraction method integrating BERT network and location feature information to perform entity extraction on the CALL2018 data, complete the entity classification of crime names, legal articles, and prison terms, improve the entity classification of criminal law, and improve the accuracy of entity classification of criminal law for crime name prediction, legal article recommendation, and prison term prediction.
[0040] The beneficial effects of the present invention are as follows:
[0041] The present invention adopts the above scheme, generates multi-data samples conforming to the field through BIO annotation sampling, and then divides and normalizes the data set. Use the BERT network for unsupervised pre-training on the text, and then add Bi-LSTM for fine-tuning on specific downstream tasks to obtain a stronger ability to capture context semantic information. Secondly, add location feature information on the basis of the attention mechanism to selectively focus on certain information and better extract information features. Finally, realize entity relation recognition and classification. The proposed entity relation extraction method integrating BERT network and location feature information can achieve high-precision entity extraction in complex fields with few samples considering the entity overlap of domain data. The present invention performs entity extraction on the data in the FB15K data set. Brief Description of the Drawings
[0042] Figure 1 Improved attention mechanism model structure diagram.
[0043] Figure 2Entity relation extraction model diagram integrating BERT network and position feature information.
[0044] Figure 3 Basic BERT structure diagram for entity extraction.
[0045] Figure 4 ACC value diagram under different models. Specific implementation manner
[0046] The entity relation extraction method integrating BERT network and position feature information comprises the following steps:
[0047] 1) Entity data collection: Obtain a publicly available dataset in a professional vertical domain and collect samples according to BIO annotation;
[0048] 2) Data processing: Divide the obtained entity samples into a training set, a validation set and a test set, and perform maximum-minimum normalization processing;
[0049] 3) Propose a model structure: The proposed entity relation extraction model integrating BERT network and position feature information is as Figure 2 shown, and is composed of a BERT model, a Bi-LSTM, a linear layer, an Attention layer, and a fully connected layer; First, the BERT layer captures stronger context semantic information. Second, the complete time-step hidden state sequence is concatenated through the Bi-LSTM and mapped to the linear layer to obtain the scores of each category label. Then, weights are assigned through the Attention layer, and specific information is selected and concatenated with the position embedding matrix P in the fully connected layer to distinguish the feature representations of the same entity in different relations. Finally, entity classification prediction is completed through the entity classification matrix C;
[0050] 3.1) The data passes through the BERT model and the Bi-LSTM to obtain the scores of each category label corresponding to the word level:
[0051] Take BERT as the encoder of the input text sequence, and use BERT to obtain the hidden layer state vector X t , specifically as shown in formula (1):
[0052] X t = Bert base (w t ) (1)
[0053] Record it as the sequence X=(x 1 , x 2 , …, x n ).
[0054] Take the sequence X obtained by the BERT network as the input of each time step of the bidirectional long short-term memory network Bi-LSTM to obtain the forward hidden state sequence of the Bi-LSTM layer and the backward hidden state sequence where h t Introduce corresponding memory units in the hidden layer, as shown in formula (2).
[0055]
[0056] where, h t-1 is the output result of the hidden layer of the previous long short-term memory network unit; C t-1 is the state result of the previous long short-term memory network unit; x t is the word vector input result of this article; f t is the output result of the forgetting gate, where σ is the sigmoid activation function; i t and is the output result of the input gate; tanh is the tanh activation function; Ο t is the output result of the output gate; C t is the state value of the current unit; h t is the output of the hidden layer of the current unit.
[0057] Concatenate the forward and backward hidden state sequences according to the time step to obtain the complete hidden state sequence, denoted as H=(h 1 ,h 2 ,…,h n ).
[0058] Finally, map the hidden state sequence to s dimensions through a linear layer, and the obtained mapped sequence is denoted as L=(l 1 ,l 2 ,…,l n ), where L represents the score of y i corresponding to each category label of word X j .
[0059] 3.2) Propose a relational position feature attention mechanism: Input the scores of each category label into the relational position feature attention mechanism to obtain an entity classification matrix.
[0060] Use the attention mechanism QKV model to calculate the weight values of the entity relationship classification results. The vector matrix query k*1 obtained by random sampling through a uniform distribution yields the Query matrix, where k is the output vector dimension of the hidden layer of the bidirectional long short-term memory network. The Key matrix is obtained from the feature matrix generated by the word vectors of the Chinese word segmentation in the sentence. The Value matrix is obtained from the matrix composed of the output vectors of the hidden layer of the bidirectional long short-term memory network.
[0061] The weight values in the attention mechanism for entity relationship extraction are calculated according to formula (3):
[0062] Attention_w n×1 = softmax(key_w n×k * query_w k×1 ) (3)
[0063] Among them, the softmax function is used for vector normalization operation; key_w n×k is the Key vector matrix of the attention mechanism; query_w k×1 is the Query vector matrix in the attention mechanism; Attention_w n×1 is the weight value of the attention mechanism.
[0064] In entity relation extraction, the output matrix of the attention mechanism is calculated according to formula (4):
[0065] Attention_r k×1 = (Attention_w T * value_w n×k ) T (4)
[0066] Among them, value_w n×k is the output vector matrix of the hidden layer of the bidirectional long short-term memory network; Attentiont_r k×1 is the output vector matrix of the attention mechanism.
[0067] In addition, establish the relationship between the attention output sequence and the output sequence passing through BERT, distinguish the feature representations of the same entity in different relations, and make the final relation classification prediction.
[0068] Calculate the distance between each word in the output sequence of BERT and the trigger word of the current attention output sequence, then randomly initialize the position embedding matrix P according to the maximum sentence length m and the position feature size n, and obtain the relation position feature Pf of each word by querying the position embedding matrix P t , and the relation position feature is calculated according to formula (5):
[0069] Pf t = P Pr-Pw (5)
[0070] Among them, Pr represents the position of the relation trigger word; Pw represents the position of the word in the BERT output sequence.
[0071] Finally, splice the position embedding matrix and the attention mechanism output matrix in the fully connected layer to obtain the entity classification matrix (6):
[0072] C = f(Attention_r k×1 ; P) (6)
[0073] where f(·) represents the fully connected layer; Attentiont_r k×1 is the output vector matrix of the attention mechanism; P is the position embedding matrix, and C is the current entity classification matrix.
[0074] The improved attention mechanism model structure is as Figure 1 shown. In this model, the Query matrix is a vector matrix query randomly sampled from a uniform distribution k*1 , where k is the output vector dimension of the hidden layer of the bidirectional long short-term memory network, and the Key matrix is a feature matrix generated from the word vectors of the Chinese word segmentation in the sentence. The Value matrix is a matrix composed of the output vectors of the hidden layer of the bidirectional long short-term memory network.
[0075] 4) Offline training: Use the training set and regularization strategy to train the model and the entity classification matrix C, and save the optimal parameters;
[0076] 5) Online testing: Apply the test set to verify the model performance or load the pre-trained parameters to fine-tune the entire model. Use parameter sharing and transfer learning to achieve timely training of the model.
[0077] Example 1:
[0078] I. Theoretical basis of the solution of the present invention:
[0079] 1. BERT network
[0080] BERT consists of three modules: the Embedding module on the left, the Transformer module in the middle, and the pre-fine-tuning module on the right. The general BERT in entity extraction is as Figure 3 shown.
[0081] In entity extraction, Embedding includes three parts: the token embedding tensor Token Embedding, the segment embedding tensor Segment Embedding, and the position encoding tensor Position Embeddings. The output tensor of the entire Embedding module is the direct sum result of these 3 tensors.
[0082] In entity extraction, only the Encoder part of the classical Transformer architecture is used in BERT, and the Decoder part is completely discarded. The two major pre-training tasks are also concentrated in the training of the Transformer module. After being processed by the middle-layer Transformer, the last layer of BERT makes different adjustments according to different requirements of entity extraction
[0083] II. Implementation process of the technical solution of the present invention:
[0084] 1. Entity data collection: Obtain publicly available datasets in professional vertical fields and collect samples according to BIO annotation;
[0085] 2. Data processing: Divide the obtained entity samples into training set, validation set and test set, and perform maximum-minimum normalization processing;
[0086] 3. Propose a model structure: An entity relation extraction model that combines BERT network and positional feature information, which consists of a BERT model, Bi-LSTM, linear layer, Attention layer, and fully connected layer; First, the BERT layer captures stronger context semantic information. Secondly, the complete time-step hidden state sequence is concatenated through Bi-LSTM and mapped to the linear layer to obtain the scores of each category label. Then, weights are assigned through the Attention layer, and specific information is selected to be concatenated with the positional embedding matrix P in the fully connected layer to distinguish the feature representations of the same entity in different relations. Finally, entity classification prediction is completed through the entity classification matrix C;
[0087] 3.1 The data passes through the BERT model and Bi-LSTM to obtain the scores of each category label corresponding to the word level:
[0088] The BERT model internally uses multiple layers of Transformer as its encoding structure. Compared with the recurrent neural network based on time series, BERT has a stronger ability to capture context semantic information and contains richer syntactic, semantic and context information. Taking BERT as the encoder of the input text sequence, the output is used as the input of each time step of the Bi-LSTM bidirectional long short-term memory network, and the forward hidden state sequence of the Bi-LSTM layer is obtained and the backward hidden state sequence The forward and backward hidden state sequences are concatenated according to the time step to obtain a complete hidden state sequence, and then the hidden state sequence is mapped to s dimensions through the linear layer, that is, the number of label categories in the annotation set.
[0089] 3.2 Propose a relationship position feature attention mechanism: Input the scores of each category label into the relationship position feature attention mechanism to obtain the entity classification matrix;
[0090] In order to further focus on specific information and thus better extract information features, a relationship position feature attention mechanism is proposed. The "QKV" model of the attention mechanism is used to calculate the weight value of the attention mechanism and the attention mechanism output matrix in entity relation extraction. Calculate the distance between each word in the output sequence of BERT and the trigger word of the current attention output sequence, and then randomly initialize the position embedding matrix P according to the maximum sentence length m and the position feature size n. Obtain the relationship position feature Pf of each word by querying the position embedding matrix P t, finally, entity classification prediction is completed through the entity classification matrix C.
[0091] 4. Offline training: Use the training set and regularization strategy to train the model and save the optimal parameters of the entity classification matrix C;
[0092] 5. Online testing: Apply the test set to verify the model performance or load the pre-trained parameters to fine-tune the entire model, and use parameter sharing transfer learning to achieve timely training of the model.
[0093] Evaluation metrics: In the field of entity relation extraction, compare the precision and recall of different models. When it is not easy to directly judge the performance advantages and disadvantages when the two metrics are high and low respectively, compare the F1-Score value. The calculation formulas of precision, recall, and F1-Score are shown in equations (7) - (9).
[0094]
[0095]
[0096]
[0097] Among them, TP represents the number of true predictions when it is actually true; FP represents the number of false predictions when it is actually false, that is, the error rate; FN represents the number of false predictions when it is actually true, that is, the false negative rate.
[0098] 5.1 FB15K dataset
[0099] FB15K is a subset of the knowledge graph Freebase, containing a large amount of general human knowledge. FB15k contains 14,951 entities and 592,213 triples. The approximate ratio of the training set, validation set, and test set is 9:1:1. The experimental results are shown in Table 1.
[0100] Table 1 Accuracy, recall, and F1-Score values under different models
[0101]
[0102] It can be seen from Table 1 that BERT_BAP is the algorithm of the present invention. The algorithm of the present invention has achieved the best performance in terms of accuracy, recall, and F1-Score on the dataset. Compared with the word embedding model, the BERT network pre-training model can better extract the feature information between corpora.
[0103] Figure 4It is the comparison result of the ACC of different models on the FB15K dataset for experiments. BERT_BAP is the algorithm of the present invention. It can be seen that the change of the ACC value along with the training process of the dataset, and finally it converges to a stable value. It can be concluded from the figure that the ACC value of BERT_BAP is better. Generally speaking, the BERT network model has better performance.
[0104] The algorithm proposed by the present invention can be applied in the military field, the medical field, etc. Through the unsupervised pre-training of the BERT network on the text, and then adding Bi-LSTM for fine-tuning on specific downstream tasks, a stronger ability to capture context semantic information is obtained. Based on the attention mechanism, position feature information is added to selectively focus on certain information and better extract information features. Furthermore, entity relationship recognition and classification are carried out to effectively extract entity relationships.
Claims
1. An entity relation extraction method integrating BERT network and position feature information, characterized in that, its steps are as follows: 1) Entity data collection: Obtain the publicly available dataset in the professional vertical field and collect samples according to BIO annotation; 2) Data processing: Divide the obtained entity samples into training set, validation set and test set, and perform maximum-minimum normalization processing; 3) Propose the model structure: The proposed entity relation extraction model integrating BERT network and position feature information is composed of BERT model, Bi-LSTM, linear layer, Attention layer, and fully connected layer; First, the BERT layer captures stronger context semantic information. Secondly, the complete time-step hidden state sequence is concatenated through Bi-LSTM and mapped to the linear layer to obtain the scores of each category label. Then, weights are assigned through the Attention layer, and specific information is selected to be concatenated with the position embedding matrix P in the fully connected layer to distinguish the feature representations of the same entity in different relations. Finally, entity classification prediction is completed through the entity classification matrix C; Calculate the distance between each word in the output sequence of BERT and the trigger word of the current attention output sequence, and then randomly initialize the position embedding matrix P according to the maximum sentence length m and the position feature size n. Obtain the relative position feature Pf of each word by querying the position embedding matrix P t , and the relative position feature is calculated according to formula (5): Pf t = P Pr-Pw (5) where Pr represents the position of the relation trigger word; Pw represents the position of the word in the BERT output sequence; The position embedding matrix and the output matrix of the attention mechanism are concatenated in the fully connected layer to obtain the entity classification matrix (6): C = f(Attention_r k×1 ; P)(6) where f(·) represents the fully connected layer; Attentiont_r k×1 is the output vector matrix of the attention mechanism; P is the position embedding matrix, and C is the current entity classification matrix; 4) Offline training: Use the training set and regularization strategy to train the model and save the optimal parameters of the entity classification matrix C; 5) Online testing: Apply the test set to verify the model performance or load the pre-trained parameters to fine-tune the entire model, and use parameter sharing transfer learning to realize the timely training of the model.
2. The entity relation extraction method integrating BERT network and position feature information according to claim 1, characterized in that, in the step 3), the specific method is: 3.1) The data passes through the BERT model and Bi-LSTM to obtain the scores of each category label corresponding to the word level: Taking BERT as the encoder of the input text sequence, BERT is used to obtain the hidden layer state vector X t , as specifically shown in formula (1): X t = Bert base (w t )(1) Denote it as sequence X = (x 1 , x 2 , …, x n ); Take the sequence X obtained by the BERT network as the input of each time step of the bidirectional long short-term memory network Bi-LSTM, and obtain the forward hidden state sequence of the Bi-LSTM layer and the backward hidden state sequence where h t Introduce corresponding memory units in the hidden layer, as shown in formula (2); Among them, h t-1 is the output result of the hidden layer of the previous long short-term memory network unit; C t-1 is the state result of the previous long short-term memory network unit; x t is the word vector input result of this article; f t is the output result of the forgetting gate, where σ is the sigmoid activation function; i t and are the output results of the input gate; tanh is the tanh activation function; Ο t is the output result of the output gate; C t is the state value of the current unit; h t is the output of the hidden layer of the current unit; Concatenate the forward and backward hidden state sequences along the time steps to obtain the complete hidden state sequence, denoted as H = (h 1 , h 2 , …, h n ); Finally, the hidden state sequence is mapped to s dimensions through a linear layer, and the resulting mapped sequence is denoted as L = (l 1 , l 2 , …, l n ); where L represents the score of y i corresponding to each class label of word X j . 3.2) Propose the relation position feature attention mechanism: Input the scores of each category label into the relation position feature attention mechanism to obtain the entity classification matrix C; Calculate the weight value of the entity relationship classification result using the attention mechanism QKV model; the vector matrix query is randomly sampled through a uniform distribution k*1 Obtain the Query matrix, where k is the output vector dimension of the hidden layer of the bidirectional long short-term memory network; obtain the Key matrix through the feature matrix generated by the word vectors of the Chinese word segmentation in the sentence; obtain the Value matrix through the matrix composed of the output vectors of the hidden layer of the bidirectional long short-term memory network; The weight value of the attention mechanism in entity relation extraction is calculated according to formula (3): Attention_w n×1 = softmax(key_w n×k * query_w k×1 )(3) Among them, the softmax function is used for vector normalization operation; key_w n×k is the Key vector matrix of the attention mechanism; query_w k×1 is the Query vector matrix in the attention mechanism; Attention_w n×1 is the weight value of the attention mechanism; The output matrix of the attention mechanism in entity relation extraction is calculated according to formula (4): Attention_r k×1 = (Attention_w T * value_w n×k ) T (4) Among them, value_w n×k is the output vector matrix of the hidden layer of the bidirectional long short-term memory network; Attentiont_r k×1 is the output vector matrix of the attention mechanism; Establish the relationship between the attention output sequence and the output sequence passing through BERT, distinguish the feature representations of the same entity in different relations, and make the final relation classification prediction.
3. The entity relation extraction method integrating BERT network and position feature information according to claim 2, characterized in that: Use the entity relation extraction method integrating BERT network and position feature information to perform entity extraction on CNMER data, and complete entity classification of disease signs and disease treatments.
4. The entity relation extraction method integrating BERT network and position feature information according to claim 2, characterized in that: Use the entity relation extraction method integrating BERT network and position feature information to perform entity extraction on AIR FORCE MIL-HDBK-310-1997 data, and complete entity classification of different climates for developing military products.
5. The entity relationship extraction method integrating the BERT network and location feature information according to claim 2, characterized in that: Entity extraction is performed on the LendingClub data using the entity relationship extraction method integrating the BERT network and location feature information to complete entity classification of loan customers, loan services, and loan default factors.
6. The entity relationship extraction method integrating the BERT network and location feature information according to claim 2, characterized in that: Entity extraction is performed on the CALL2018 data using the entity relationship extraction method integrating the BERT network and location feature information to complete entity classification of crime names, legal articles, and prison terms, and improve criminal law entity classification.