A Biological Nested Named Entity Recognition Method Based on an Attention State Transition Model

By adopting an attention state transfer model based on the biological named entity recognition, combining the semantic masking model and attention mechanism, the problem of nested named entities and long entity recognition in the prior art is solved, and more accurate and efficient biological entity extraction is achieved.

CN115017909BActive Publication Date: 2025-05-27ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210653161.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-09
Publication Date
2025-05-27
Estimated Expiration
2042-06-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify nested named and long entities in biological domain texts, especially in the absence of global context information and attention mechanisms for feature representations.

Method used

A biological nested named entity recognition method based on the attention state transfer model is adopted, candidate entities are extracted through the attention state transfer model, and entities are screened based on context information using the semantic mask model to increase the attention mechanism to strengthen the correlation between words.

Benefits of technology

Effectively extract five types of entities, including DNA, RNA, protein, cell lines and cells in biological texts, improve the recognition ability of nested entities and long entities, reduce error transmission, and enhance generalization of complex entities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115017909B_ABST
    Figure CN115017909B_ABST
Patent Text Reader

Abstract

A method for biological nested named entity recognition based on an attention state transition model, comprising: 1. Dividing biological domain texts containing entity tags of five types, namely DNA, RNA, protein, cell line, and cell, into training data and test data; 2. Adjusting the training data to meet the input form of the model according to the input forms of the attention state transition model and the semantic masking model; 3. Training the attention state transition model to learn the correlation between words, and extracting candidate entities from the text and judging their types through the states output by the model; 4. Training the semantic masking model to judge whether the candidate entities and their types conform to the context semantics; 5. Inputting the test data into the attention state transition model to extract candidate entities, then masking the extracted entities and sending them into the semantic masking model for screening, and finally confirming the real entities that conform to the context.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for recognizing biological nested named entities based on an attention state transition model, particularly for the recognition of five entity types: DNA, RNA, protein, cell line, and cell. For biological text, sentences are fed into the model. The model traverses each word in the sentence to adjust its state. The model output, i.e., information about the model state transition, identifies candidate entities and their types within the sentence. Finally, a semantic masking model is used to determine whether the candidate entity satisfies the contextual information, resulting in the final entity and its type. This entity may contain nested or long entities. Background Art

[0002] With the advancement of biological research, more and more biology-related texts have accumulated, and text mining technology for automated text processing has become increasingly important. One of the most basic and important tasks is biological named entity recognition. Named entity recognition technology can effectively extract entity names related to the biological field from biological texts, such as the five types of entities: DNA, RNA, protein, cell line, and cell. Due to the presence of nested entities, manual extraction is inefficient and costly. The extracted entities are also affected by human factors, leading to extraction errors. At the same time, the identification of nested entities also increases the difficulty of extraction. Therefore, using deep learning models to solve the problem of named entity recognition is a current trend, providing a foundation for subsequent tasks such as entity linking, relationship extraction, building biological knowledge graphs, and biological knowledge question answering.

[0003] The traditional approach is to treat named entity recognition as a sequence labeling task, assigning a label to each word in a sentence. However, this approach is not suitable for the recognition of nested named entities. In recent years, various methods for nested named entity recognition have been proposed:

[0004] (1) Modify the sequence labeling model, changing the labeling of each word from one label to multiple labels, where the number of labels for each word depends on the number of nested levels of the nested entity; or modify the meaning of the label to make the meaning of each label more specific to support the recognition of nested entities.

[0005] (2) The hypergraph method actually uses a decoder to label each word in the sentence with all possible labels. The combination of these labels can identify all possible nested entities in the text.

[0006] (3) Span-based method: The idea of ​​this method is to regard entity recognition as a span classification task, that is, to find all possible spans in the text and then predict their categories.

[0007] (4) The method based on reading comprehension is to find the boundaries of entities in the text through question-answering.

[0008] (5) The method based on the state transition model. The idea of ​​this method is to store each word in the sentence in a data structure, change the state of the data structure by parsing each word, and finally determine the nested entity through the output of the state.

[0009] For simple modification of sequence labeling models, assigning multiple labels to each word depends on the number of nested levels of the text, which lacks generalization. Modifying the meaning of the label may cause ambiguity in the label information, resulting in the inability to correctly identify nested entities. Span-based methods have problems such as high computational cost, neglect of boundary information, and difficulty in identifying long entities. For reading comprehension methods, different problem sets need to be constructed for different fields. Methods based on state transition models lack global contextual features in feature representation, require certain constraints on changes in specific states, and lack an attention mechanism during state output, ignoring the correlation between structures. Summary of the Invention

[0010] The present invention aims to overcome the above-mentioned shortcomings of the prior art and proposes a biological nested named entity recognition method based on an attention state transfer model.

[0011] In order to achieve the above object, the present invention provides the following technical solutions:

[0012] A method for biological nested named entity recognition based on an attention state transfer model includes the following steps:

[0013] Step 1: Divide the biological domain text containing five types of entity labels: DNA, RNA, protein, cell line, and cell into training data and test data;

[0014] Step 2: According to the input form of the attention state transfer model and the semantic mask model, adjust the training data to meet the model input form;

[0015] The input of the attention state transfer model is the state of the model, which is defined as a tuple (B1, S1, S2, B2), where B1 and B2 represent two queues, which are used as buffers to store context information, and S1 and S2 represent stacks, which are used to store words that the model needs to judge in the current state; the S2 structure only stores one word to judge whether a single word constitutes an entity, and the S1 structure stores words that may constitute an entity with S2; the words in the sentence are stored in the dictionary {'buffer1':[],'stack1':[],'stack2':[],'buffer2':[]} to represent the current state of the model; positive examples of the attention state transfer model dataset are generated based on the entities and their types in the sentence. When a single word in S1 constitutes an entity or S1 and words in S2 constitute an entity, the output label of the model is the type of its entity. When a word in S1 is associated with a word in S2 but does not constitute a complete entity, 'correlation' is used as the output label of the model. Randomly extract non-entity words to generate negative examples of the attention state transfer model dataset, and use 'not' to represent the label of the negative sample;

[0016] The input of the semantic masking model is a sentence with a special identifier separating the original sentence from the masked sentence. The masked sentence is based on the original sentence, with the type entity in the original sentence replaced with its type identifier. The positive examples of the semantic masking model dataset are generated based on the entities and their types in the sentence, and the negative examples of the attention state transfer model dataset are generated by randomly extracting non-entity words.

[0017] Step 3: Train the attention state transfer model to learn the association between words. The state output by the model can be used to extract candidate entities from the text and determine their types.

[0018] By splicing context representation Non-contextual representation and character-level representations As the word vector of the current word

[0019]

[0020] in, Obtained through pre-training model; non-contextual representation Obtained through pre-trained Wordvecs; Each character in the word is generated by the BiLSTM model; [;] represents the concatenation operation of the vector;

[0021] The state representations β1 and β2 of B1 and B2 are obtained by extracting features from the word vectors in the structure through a unidirectional LSTM model;

[0022]

[0023]

[0024] in, represents the d-dimensional vector representation of the i-th word in B1, represents the d-dimensional vector representation of the i-th word in B2;

[0025] For the type judgment of a single word, the state representation S1 is also obtained by extracting features from the word vectors in the structure through a unidirectional LSTM model;

[0026]

[0027] in, represents the d-dimensional vector representation of the i-th word in S1;

[0028] The state of S2 indicates that S2 is Represents the d-dimensional vector representation of the words in S2;

[0029] For the type judgment of multiple words, the attention mechanism is introduced because the model needs to pay attention to the correlation between the words in the two structures S1 and S2;

[0030]

[0031] in, Represents a scaling factor, which is used to optimize the defects of dot product attention, scaling the value to the area where the softmax function changes the most, and amplifying the gap.

[0032] At this time, the state of S1 indicates that S1 is the result of paying attention to the word vectors in S1 and S2 and passing them through LSTM. The state of S2 indicates that S2 is the result of paying attention to the word vectors in S2 and S1.

[0033] S1=LSTM(Attention(S′1,S′2)) (6)

[0034] S2=Attention(S′2,S′1) (7)

[0035] in, Represents the matrix composed of the word vectors of h words in S1, Represents the matrix composed of word vectors of the words in S2;

[0036] The state of the entire model is represented as It is composed of the state representations of 4 structures;

[0037] Pk =[β1; S1; S2; β2] (8)

[0038] Get the state representation P of the model k After that, it will be classified by the multi-layer perceptron MLP, and the words classified as entity types will be regarded as candidate entities;

[0039] Step 4: Train a semantic mask model to determine whether the candidate entity and its type conform to the context semantics;

[0040] Use the BERT model to obtain the vector representation v of the global features of the masked sentence and the original sentence [cls] 、v [sep] , splicing vector v [cls] and v [sep] , and classify it through the multi-layer perceptron MLP to determine whether the masked entity boundary and entity type are contextually semantic;

[0041] Step 5: Input the test data into the attention state transfer model to extract candidate entities. Then, mask the extracted entities and send them to the semantic mask model for screening, and finally identify the entities that meet the context.

[0042] The input of the model is a word sequence x0,x1,...,x n The overall operation process is as follows:

[0043] 1. All words are stored in B2 as the initial state, which can be expressed as {'buffer1':[],'stack1':[],'stack2':[],'buffer2':[x0,x1,...,x n ]};

[0044] 2. Put the first word of B2 into S2. The current state can be expressed as {'buffer1':[],'stack1':[],'stack2':[x0],'buffer2':[x1,...,x n ]};

[0045] 3. Determine the type of each word in S2; if the entity type of a single word is obtained, the word is used as a candidate entity, and is masked and sent to the semantic mask model for judgment;

[0046] 4. Store the words in S2 into S1. The current state can be expressed as {'buffer1':[],'stack1':[x0],'stack2':[],'buffer2':[x1,...,x n ]};

[0047] 5. Put the first word in B2 into S2. The current state can be expressed as {'buffer1':[],'stack1':[x0],'stack2':[x1],'buffer2':[x2,...,x n ]};

[0048] 6. Perform multiple word type judgments on the words in S1 and S2; if the judgment result is an entity type, the words in S1 and S2 are used as candidate entities, and are masked and sent to the semantic mask model for judgment; if the type is 'correlation', the words in S2 are placed in S1, and the first word of B2 is placed in S2, that is, {'buffer1':[],'stack1':[x0,x1],'stack2':[x2],'buffer2':[x3,...,x n ]}, then perform type judgment on multiple words, and repeat this process until the output type is 'not' or B2 is empty;

[0049] 7. Adjust the state to {'buffer1':[],'stack1':[],'stack2':[x0],'buffer2':[x1,...,x n ]}, store the elements of S2 into B1, that is, {'buffer1':[x0],'stack1':[],'stack2':[],'buffer2':[x2,...,x n ]};

[0050] 8. Repeat the process from 2 to 7 until each word in the sentence is traversed by S2 to obtain a set of candidate entities. Mask the entities in the set, and send the original sentence and the masked sentence to the semantic mask model for judgment. If it meets the context semantics, it is the final recognized entity. If it does not meet the context semantics, it is discarded.

[0051] Performing the above operations on all sentences in the test set will obtain all recognized entities, including nested entities and long entities.

[0052] Furthermore, the pre-trained model package in step 3 includes BERT, BioBERT, etc.; the pre-trained Wordvecs in step 3 include BioWord2Vec and GloVe.

[0053] Furthermore, the multilayer perceptron MLP described in steps 3 and 4 includes a linear layer and a GLUE activation function.

[0054] The present invention can effectively extract five types of entities, namely DNA, RNA, protein, cell line and cell, from biological texts. In view of the lack of global contextual information in feature extraction based on the state transfer model, a semantic mask model is added to screen the extracted entities according to contextual information. The structure of the original state transfer model is changed, and the model will automatically change the state according to the rules, eliminating the constraints of specific state changes, reducing the influence of the model on the results of the previous transformation, and reducing error transmission; in the process of judging nested entities and long entities, the attention mechanism is added to strengthen the association between words, which is helpful for the recognition of nested entities and long entities. Therefore, this method has two models, one is the attention state transfer model, and the other is the semantic mask model; the purpose of the attention state transfer model is to extract candidate entities in the text, and the semantic mask model is to screen candidate entities based on global context features.

[0055] The advantages of the present invention are: it can effectively extract five types of entities, namely DNA, RNA, protein, cell line and cell, from biological texts; in view of the lack of global context information in feature extraction based on the state transition model, a semantic mask model is added to screen the extracted candidate entities according to context information; the structure of the original state transition model is changed, and the model will automatically change the state according to the rules, eliminating the constraints of specific state changes, reducing the influence of the model on the results of the previous transformation, and reducing error transmission; in the process of judging nested entities and long entities, an attention mechanism is added to strengthen the association between words, which is helpful for the recognition of nested entities and long entities. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a schematic diagram of the overall process of the present invention.

[0057] Figure 2 Schematic diagram of the attention state transfer model.

[0058] Figure 3 Schematic diagram of the model workflow. DETAILED DESCRIPTION

[0059] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described below with reference to the accompanying drawings.

[0060] Step 1: Divide the biological domain text containing five entity labels of DNA, RNA, protein, cell line, and cell into training data and test data according to a certain ratio. The data sample is: {"tokens":["There","is","a","single","methionine","codon-initiated","open","reading","frame","of","1,458","nt","in","frame","with","a","homeobox","and","a","CAX","repeat",",","and","the","open","reading","frame","is","predicted","to","encode","a","protein","of","51,659","daltons."],"entities":[{"start":16,"end" :17,"type":"DNA"},{"start":4,"end":9,"type":"DNA"},{"start":24,"end":27,"type":"DNA"},{"start":19,"end":21,"type":"DNA"}]}, in the formed data, each data is a dictionary data composed of keys "tokens" and "entities", the value of the key "tokens" is a list, which means a sentence in the biological field, each word in the sentence is an element in the list, the value of the key "entities" is a list, composed of several dictionaries, the number of dictionaries in the list corresponds to the number of entities in the sentence, and each dictionary contains three information, "start", "end", and "type" respectively represent the starting position of the entity, the ending position of the entity, and the type of the entity.

[0061] Step 2: According to the input form of the attention state transfer model and the semantic mask model, adjust the training data to meet the model input form;

[0062] The attention state transfer model uses two queues and two stacks to store data. The queues store contextual information, namely words that have been judged and those that have not been judged, and the stacks store words that need to be judged. Each time a word in the stack is judged, the four structures undergo a corresponding transformation. The model state is defined as a tuple (B1, S1, S2, B2), where B1 and B2 represent two queues, which serve as buffers to store contextual information, and S1 and S2 represent stacks, which store words that the model needs to judge in the current state. The dictionary {'buffer1':[],'stack1':[],'stack2':[],'buffer2':[]} is used to store words in the sentence to represent the current model state. The elements in each state are the stored words.

[0063] Based on the characteristics of the language and the need for nested entity recognition, we judge entities consisting of a single word and entities consisting of multiple words (two or more) separately. For the judgment of a single word, the word currently to be judged will first be stored in S2. According to the type of the single word, the model outputs six states, namely 'not', 'DNA', 'RNA', 'cell_type', 'protein', and 'cell_line'. Among them, 'not' indicates that the current single word does not constitute an entity. The other states indicate that the single word constitutes an entity of the corresponding type and is used as a candidate entity. For the judgment of multiple words, the word that needs to be judged will be stored in S2 first, and the previous words associated with the current word will be stored in S1. According to the types of multiple words, the model outputs 7 states, namely 'not', 'DNA', 'RNA', 'cell_type', 'protein', 'cell_line', and 'correlation'. Among them, 'not' means that the word in S1 is not associated with the word in S2, 'correlation' means that the word in S1 is associated with the word in S2 but does not constitute a complete entity. The other states indicate that the words in S1 and S2 constitute entities of the corresponding type and are used as candidate entities.

[0064] Given a word sequence x0,x1,...,x n First, a dictionary {'buffer1':[],'stack1':[],'stack2':[],'buffer2':[]} is generated. According to the "entities" information in each data, the data is divided into single-word type judgment and multi-word type judgment.

[0065] If end-start is 1, the word is x i, the word is stored in stack2, and the words before it are stored in buffer1 and stack1. Specifically, the current dictionary is represented by {'buffer1':[x1,...,x i-2 ],'stack1':[x i-1 ],'stack2':[x i ],'buffer2':[x i+1 ...,x n ]}, the label is the information in "type", that is, x i Entity type; if a single word entity is nested in multiple word entities, such as x j ,...,x i Is a long entity, x i For a single word entity, x j ,...,x i-1 If stored in stack1, the current dictionary is {'buffer1':[x1,...,x j-1 ],'stack1':[x j ,...,x i-1 ],'stack2':[x i ],'buffer2':[x i+1 ...,x n ]}, the label is word x i The entity type represented by the word is randomly extracted as a negative example. The words that are not single-word entities are stored in the same way as the above operation in 4 structures with the label 'not'.

[0066] If end-start is greater than 1, x j ,...,x i Is a long entity, traverse x j to x i-1 , the current word x m Store (j<m≤i-1) in stack2, x j ,...,x m-1 Stored in stack1, the current dictionary is {'buffer1':[x1,...,x j-1 ],'stack1':[x j ,...,x m-1 ],'stack2':[x m ],'buffer2':[x m+1 ...,x n ]}, the label is 'correlation', when xi is stored in stack2, the current dictionary is {'buffer1':[x1,...,x j-1],'stack1':[x j ,...,x i-1 ],'stack2':[x i ],'buffer2':[x i+1 ...,x n ]}, the tag becomes the type tag corresponding to the long entity. If the left boundaries of the two entities are the same but the boundaries are different, the right boundary x of the nested long entity h When stored in stack2, the label becomes the type label corresponding to the nested long entity, and the rest remains unchanged, such as x j ,...,x h Long entity nested in x j ,...,x i In the long entity, traverse to x h When the dictionary is {'buffer1':[x1,...,x j-1 ],'stack1':[x j ,...,x h-1 ],'stack2':[x h ],'buffer2':[x h+1 ...,x n Randomly extract words that are not long entities as negative examples, and store them in 4 structures with the label 'not' as the above operation.

[0067] The semantic masking model begins operating when the attention state transfer model outputs five types: 'DNA', 'RNA', 'cell_type', 'protein', and 'cell_line'. The model masks the entity with the type label and then performs a binary classification. An output of '0' indicates that the candidate entity does not conform to the context semantics and is considered an incorrect entity. An output of '1' indicates that the candidate entity conforms to the context semantics and is confirmed to be a true entity.

[0068] Assume that the entity is x j ,...,x i , the original sequence is x0,x1,...,x n , mask the original sentence and convert the entity sequence into its type label [type] (such as '[DNA]'). The masked sequence is: x0,...,x j-1 ,[type],x i+1 ,...,x n . Use the special identifiers '[cls]' and '[sep]' to concatenate the original sequence with the masked sequence, and the sequence becomes: [cls],x0,x1,...,x n ,[sep],x0,...,x j-1 ,[type],xi+1 ,...,x n The label is set to "1", and words that are not entities are randomly selected and masked with random types as negative examples for model training, and the label is set to "0".

[0069] Step 3: Train the attention state transfer model to learn the association between words. The state output by the model can be used to extract candidate entities from the text and determine their types.

[0070] By splicing context representation Non-contextual representation and character-level representations As the word vector of the current word

[0071]

[0072] in, Obtained through pre-training model; non-contextual representation Obtained through pre-trained Wordvecs; Each character in the word is generated by the BiLSTM model; [;] represents the concatenation operation of the vector.

[0073] The state representations β1 and β2 of B1 and B2 are obtained by extracting features from the word vectors in the structure using a unidirectional LSTM model:

[0074]

[0075]

[0076] in, represents the d-dimensional vector representation of the i-th word in B1, Represents the d-dimensional vector representation of the i-th word in B2.

[0077] For the type judgment of a single word, the state representation S1 is also obtained by extracting features from the word vectors in the structure through a unidirectional LSTM model:

[0078]

[0079] in, Represents the d-dimensional vector representation of the i-th word in S1.

[0080] The state of S2 indicates that S2 is Represents the d-dimensional vector representation of the words in S2.

[0081] For the type judgment of multiple words, since the model needs to pay attention to the correlation between the words in the two structures S1 and S2, an attention mechanism is introduced:

[0082]

[0083] in, Represents a scaling factor, which is used to optimize the defects of dot product attention, scaling the value to the area where the softmax function changes the most, and amplifying the gap.

[0084] At this time, the state of S1 indicates that S1 pays attention to the word vectors in S1 and S2 and passes them through LSTM. The state of S2 indicates that S2 pays attention to the word vectors in S2 and S1.

[0085] S1=LSTM(Attention(S′1,S′2)) (6)

[0086] S2=Attention(S′2,S′1) (7)

[0087] in, Represents the matrix composed of the word vectors of h words in S1, Represents the matrix composed of word vectors of the words in S2.

[0088] The state of the entire model is represented as It is composed of the state representation of 4 structures:

[0089] P k =[β1; S1; S2; β2] (8)

[0090] Get the state representation P of the model k After that, it will be classified by the multi-layer perceptron MLP, and the words classified as entity types will be regarded as candidate entities.

[0091] The attention state transfer model transforms the named entity recognition process into a process of state changes between structures. This approach can identify nested and long entities in text and has a certain degree of generalization when dealing with complex entities. The introduction of the attention mechanism enhances feature extraction of contextual information, helps determine entity types, and further strengthens the correlation between words in determining whether multiple words constitute an entity.

[0092] The diagram of the attention state transfer model is shown in Figure 2 .

[0093] Step 4: Train a semantic mask model to determine whether the candidate entity and its type conform to the context semantics;

[0094] Use the BERT model to obtain the vector representation v of the global features of the masked sentence and the original sentence [cls] 、v [sep] , splicing vector v [cls] and v [sep] , and multi-layer perceptron MLP is used for classification to determine whether the masked entity boundary and entity type are contextually semantic.

[0095] Step 5: Connect the two models in series, input the test data into the attention state transfer model, extract candidate entities, then mask the extracted entities and send them to the semantic masking model for screening, and finally confirm the real entities that match the context;

[0096] The input of the model is a word sequence x0,x1,...,x n The overall operation process is as follows:

[0097] 1. All words are stored in B2 as the initial state, which can be expressed as {'buffer1':[],'stack1':[],'stack2':[],'buffer2':[x0,x1,...,x n ]}.

[0098] 2. Put the first word of B2 into S2. The current state can be expressed as {'buffer1':[],'stack1':[],'stack2':[x0],'buffer2':[x1,...,x n ].

[0099] 3. Perform type judgment on individual words in S2; if the entity type of a single word is obtained, the word is used as a candidate entity, and is masked and sent to the semantic mask model for judgment.

[0100] 4. Store the words in S2 into S1. The current state can be expressed as

[0101] {'buffer1':[],'stack1':[x0],'stack2':[],'buffer2':[x1,...,x n ]}.

[0102] 5. Put the first word in B2 into S2. The current state can be expressed as {'buffer1':[],'stack1':[x0],'stack2':[x1],'buffer2':[x2,...,x n ]}.

[0103] 6. Perform multiple word type judgments on the words in S1 and S2. If the judgment result is an entity type, the words in S1 and S2 are used as candidate entities, and after masking, they are sent to the semantic mask model for judgment. If the type is 'correlation', the words in S2 are placed in S1, and the first word of B2 is placed in S2, that is, {'buffer1':[],'stack1':[x0,x1],'stack2':[x2],'buffer2':[x3,...,x n ]}, and perform type judgment on multiple words. This process is repeated until the output type is 'not' or B2 is empty.

[0104] 7. Adjust the state to {'buffer1':[],'stack1':[],'stack2':[x0],'buffer2':[x1,...,x n ]}, store the elements of S2 into B1, that is, {'buffer1':[x0],'stack1':[],'stack2':[],'buffer2':[x2,...,x n ]}.

[0105] 8. Repeat the process from 2 to 7 until each word in the sentence is traversed by S2 to obtain a set of candidate entities. Mask the entities in the set, and send the original sentence and the masked sentence to the semantic mask model for judgment. If it meets the context semantics, it is the final recognized entity. If it does not meet the context semantics, it is discarded.

[0106] The above operations are performed on all sentences in the test set, and the entity that the semantic mask model determines to be the correct entity is the last recognized entity. The model working process can be referred to Figure 3 .

[0107] The present invention proposes a biological nested named entity recognition method based on the attention state transfer model. The candidate entities in the sentence are extracted through the attention state transfer model. The attention mechanism enhances the correlation between words. The semantic mask model is used to obtain the global context information of the sentence, thereby identifying the masked candidate entities and finally obtaining the recognized entities.

[0108] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The scope of protection of the present invention should not be regarded as limited to the specific forms described in the embodiments. The scope of protection of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.

Claims

1. A method for biological nested named entity recognition based on an attention state transition model, comprising the following steps: Step 1: Divide the biological domain text containing entity tags of five types, namely DNA, RNA, protein, cell line, and cell, into training data and test data; Step 2: According to the input forms of the attention state transition model and the semantic masking model, adjust the training data to meet the input form of the models; The input of the attention state transition model is the state of the model, which is defined as a tuple (B1, S1, S2, B2), where B1 and B2 represent two queues, used as buffers Buffer to store context information, and S1 and S2 represent stacks Stack, used to store the words that the model needs to judge in the current state; Among them, the S2 structure only stores one word for judging whether a single word constitutes an entity, and the S1 structure stores the words that may form an entity with S2; the words in the sentence are stored through the dictionary {'buffer1': [],'stack1': [],'stack2': [], 'buffer2': []} to represent the current state of the model; positive examples of the attention state transition model dataset are generated according to the entities and their types in the sentence. When a single word in S1 constitutes an entity or the words in S1 and S2 constitute an entity, the output label of the model is the type of the entity. When the words in S1 are related to the words in S2 but do not form a complete entity, 'correlation' is used as the output label of the model; negative examples of the attention state transition model dataset are generated by randomly extracting non-entity words, and 'not' is used to represent the label of the negative sample; The input of the semantic masking model is a sentence in which the original sentence and the masked sentence are separated by special identifiers. The masked sentence is based on the original sentence, and the type entities in the original sentence are replaced with their type identifiers; positive examples of the semantic masking model dataset are generated according to the entities and their types in the sentence, and negative examples of the attention state transition model dataset are generated by randomly extracting non-entity words; Step 3: Train the attention state transition model to learn the correlation between words, and extract candidate entities from the text and judge their types through the state output by the model; By concatenating context representations Non-context representation and character-level representation as the word vector of the current word Among them, Obtained through a pre-trained model; non-contextual representation Obtained through pre-trained Wordvecs; Each character in the word is generated by a BiLSTM model; [;] represents the concatenation operation of vectors; The states of B1 and B2 represent β 1 , β 2 are both obtained by extracting features from the word vectors in the structure through a unidirectional LSTM model; Among them, represents the d-dimensional vector representation of the i-th word in B1, represents the d-dimensional vector representation of the i-th word in B2; For the type judgment of a single word, the state representation of S1 is S 1 It is also obtained by extracting features from the word vectors in the structure through a unidirectional LSTM model; Among them, represents the d-dimensional vector representation of the i-th word in S1; The state representation of S2 is S 2 Yes It represents the d-dimensional vector representation of the words in S2; For the type judgment of multiple words, since the model needs to pay attention to the correlation between the words in the S1 and S2 structures, an attention mechanism is introduced; Among them, represents a scaling factor, which is used to optimize the defect of dot-product attention, scale the value to the region where the softmax function changes the most, and magnify the gap; At this time, the state of S1 represents S 1 is the result of paying attention to the word vectors in S1 and S2 and passing through the LSTM. The state of S2 represents S 2 is the result of paying attention to the word vectors in S2 and S1; S 1 = LSTM(Attention(S′ 1 , S′ 2 )) (6) S 2 = Attention(S′ 2 ,S′ 1 ) (7) Among them, represents the matrix formed by the word vectors of h words in S1, represents the matrix formed by the word vectors of the words in S2; The state of the entire model is represented as which is formed by concatenating the state representations of 4 structures; P k = [β 1 ; S 1 ; S 2 ; β 2 (8) Obtain the state representation P of the model k After that, classification is performed through a multi-layer perceptron (MLP), and the words with the classification result of entity type are used as candidate entities; Step 4: Train the semantic masking model to judge whether the candidate entities and their types conform to the context semantics; Use the BERT model to obtain the vector representations v of the global features of the masked statement and the original statement [cls] , v [sep] , concatenate the vectors v [cls] and v [sep] , and perform binary classification through a multi-layer perceptron MLP to determine whether the masked entity boundary and the type of the entity are contextually semantic Step 5: Input the test data into the attention state transition model to extract candidate entities, then mask the extracted entities and send them into the semantic masking model for screening, and finally confirm the entities that conform to the context; The input of the model is the word sequence x 0 , x 1 ,..., x n , and the overall running process is as follows:

1. All words are stored in B2 in the initial state, and the initial state is represented as {'buffer1': [],'stack1': [],'stack2': [], 'buffer2': [x 0 , x 1 ,..., x n}; 2. Put the first word of B2 into S2, and the current state is represented as {'buffer1': [],'stack1': [],'stack2': [x 0 , 'buffer2': [x 1 ,..., x n}; 3. Judge the type of a single word in S2; if the entity type of a single word is obtained, then use this word as a candidate entity, mask it and send it into the semantic masking model for judgment; 4. Store the words in S2 into S1, and the current state is represented as {'buffer1': [],'stack1': [x 0 ,'stack2': [], 'buffer2': [x 1 ,..., x n}; 5. Put the first word in B2 into S2, and the current state is represented as {'buffer1': [],'stack1': [x 0 ,'stack2': [x 1 , 'buffer2': [x 2 ,..., x n}; 6. Perform type judgment on multiple words in S1 and S2; if the judgment result is the entity type, use the words in S1 and S2 as candidate entities, mask them, and then send them into the semantic masking model for judgment; if the type is 'correlation', put the words in S2 into S1, and put the first word of B2 into S2, that is, {'buffer1': [],'stack1': [x 0 , x 1 ,'stack2': [x 2 , 'buffer2': [x 3 ,..., x n}, and then perform type judgment on multiple words, and loop this process until the output type is 'not' or B2 is empty; 7. Adjust the status to {'buffer1': [],'stack1': [],'stack2': [x 0 , 'buffer2': [x 1 ,..., x n}, store the elements of S2 into B1, that is, {'buffer1': [x 0 ,'stack1': [],'stack2': [], 'buffer2': [x 2 ,..., x n}; 8. Loop through processes 2 to 7 until each word in the sentence has been traversed by S2 to obtain a candidate entity set. Mask the entities in the set, and send the original sentence and the masked sentence into the semantic masking model for judgment. If it conforms to the context semantics, it is the finally recognized entity; if not, it is discarded. Perform the above operations on all sentences in the test set to obtain all recognized entities, including nested entities and long entities.

2. A method for biological nested named entity recognition based on an attention state transition model as described in claim 1, characterized in that: The pre-trained model packages in step 3 include BERT, BioBERT, etc.; the pre-trained Wordvecs in step 3 include BioWord2Vec and GloVe.

3. A method for biological nested named entity recognition based on an attention state transition model as described in claim 1, characterized in that: The multi-layer perceptron MLP in steps 3 and 4 includes a linear layer and a GLUE activation function.

Citation Information

Patent Citations

  • Chinese text key information extraction method based on pre-trained language model

    CN111444721A

  • BERT-FLAT-based Chinese named entity recognition method

    CN112270193A