A named entity recognition method
By performing character segmentation and masking strategy processing on the training data, a synthetic span word list is generated, and the named entity recognition model is trained based on the mask language task and the span boundary task, the problem of low recognition accuracy of Chinese long words in the existing technology is solved, and a higher recognition accuracy of long words is achieved.
Patent Information
- Application Number
- CN202510211434.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The existing BERT-based naming entity recognition method has low recognition accuracy when facing long Chinese words, especially in the fields of biomedical and legal texts, which lack effective recognition of long words composed of multiple characters and multiple words.
By performing character slicing and masking strategy processing on the training data, a synthetic span word list is generated, and the named entity recognition model is trained based on the mask language task and the span boundary task, and vector features are extracted to achieve long word recognition.
It improves the accuracy of Chinese naming entity recognition in long word recognition, and is suitable for identifying professional terms and complex names in fields such as biomedical and legal texts.
Smart Images

Figure CN119692352B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing, and more specifically, to a named entity recognition method. Background Art
[0002] Named entity recognition refers to a method of efficiently and conveniently obtaining entities containing important semantic information from a large amount of data, mainly including names of people, places, institutions, proper nouns, etc. Named entity recognition is an important basic part of information extraction. Through named entity recognition, the recognized entities play an important role in natural language processing tasks such as building question-answering systems, constructing knowledge graphs, and generating text summaries.
[0003] Short-length characters account for a large proportion of Chinese entity words, but with the development of the Chinese language system, new words composed of multiple characters are constantly added, and the number of long words in modern Chinese continues to increase and their proportion has also increased. The BERT-based named entity recognition method for Chinese only uses character granularity as the prediction and recognition target, lacks attention and prediction and recognition of long words composed of multiple characters and words, and has low recognition precision and accuracy for long words such as specific terms and professional terms in subdivided fields. For example, there are often a large number of professional terms and complex names in biomedicine, legal texts, government documents, etc. These entities are often composed of multiple characters or short words. Summary of the invention
[0004] The present invention aims at the technical problems existing in the prior art and provides a named entity recognition method for solving the problem of lack of long word recognition in a specific field of Chinese named entity recognition.
[0005] The present invention provides a method for named entity recognition, comprising:
[0006] Step S1, segmenting each text data in the training data set into characters, and extracting the character features and character position features of each character;
[0007] Step S2, masking the text data according to the masking strategy to obtain multiple masked segment sequences, generating synthetic span words within each masked segment sequence, and forming a synthetic span word list corresponding to the text data;
[0008] Step S3, extracting vector features of each synthetic span word in the synthetic span word list based on the character features and character position features of each character in the synthetic span word;
[0009] Step S4, training a named entity recognition model based on a masked language task and a span boundary task according to the vector features of all synthetic span words;
[0010] Step S5: recognizing named entities in the text data to be recognized based on the trained named entity recognition model.
[0011] The present invention provides a method for named entity recognition, which masks text data according to a masking strategy to obtain multiple masked fragment sequences and generate a list of synthetic span words; extracts vector features of each synthetic span word based on the character features and character position features of each character in the synthetic span word; trains a named entity recognition model based on the vector features of all synthetic span words and based on masked language tasks and span boundary tasks; and recognizes named entities in the text data to be recognized based on the trained named entity recognition model. In view of the fact that the existing masked language model training performs predictions at the Chinese character granularity, and the prediction training lacks word-level granularity, the present invention introduces the generation and extraction of synthetic words within the span, a position marking method, and length information embedding, realizes prediction at the Chinese word granularity level, extracts span words, and obtains long word results of named entity recognition, which is suitable for long word recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A flowchart of a named entity recognition method provided by the present invention;
[0013] Figure 2 A schematic diagram of the first pre-trained vector feature when training based on a masked language task;
[0014] Figure 3 Schematic diagram of the second pre-trained vector feature when training based on the span boundary task. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. In addition, the technical features in the various embodiments or single embodiments provided by the present invention can be arbitrarily combined with each other to form a feasible technical solution. This combination is not subject to the constraints of the sequence of steps and / or the structural composition mode, but must be based on the ability of ordinary technicians in this field to achieve. When the combination of technical solutions is contradictory or cannot be achieved, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0016] Figure 1 A flowchart of a named entity recognition method provided by the present invention is as follows: Figure 1 As shown, the method includes:
[0017] Step S1, segment each text data in the training data set into characters, and extract the character features and character position features of each character.
[0018] It is understandable that for each text data in the training data set, preprocessing is first performed, and the preprocessing mainly includes: converting traditional Chinese characters in each text data into simplified Chinese characters, and unifying them into simplified Chinese characters. Special symbols in the text data are removed, including special symbols that do not contain semantic information such as numerical serial numbers, Chinese radical characters, emoticons, and shape symbols, such as ①, ⺁, (*^▽^*), ●, etc.
[0019] The preprocessed text data is segmented into characters and the position and length information of each character is marked. The main steps include:
[0020] The input text data is segmented by the word segmenter according to the character level granularity, and the length is The text sequence , Indicates characters, .
[0021] Add special markers at the beginning of a text sequence , add special markers between sentences ;
[0022] Perform character index conversion on the text sequence, and map each character to a unique index in the vocabulary according to the vocabulary used by BERT to obtain the character identifier ; The character identifier and character embedding matrix are used to calculate the character encoding vector .
[0023] Mark absolute positions of characters within a text sequence .
[0024] Mark the characters in the text sequence with a length of 1.
[0025] Step S2, masking the text data according to the masking strategy to obtain multiple masked segment sequences, generating synthetic span words within each masked segment sequence, and forming a synthetic span word list corresponding to the text data.
[0026] The text data is masked according to the masking strategy to obtain multiple masked segment sequences, including:
[0027] Step S21, based on the Chinese word length distribution statistics, calculate the masked span length m of the random sampling of the text sequence, the masked span length The sampling probability of is calculated according to the geometric distribution formula.
[0028] Among them, set the maximum cover length ,set up Take the best value of the experimental effect, is a hyperparameter that adjusts the trend of mask span length sampling length. The part is truncated without masking.
[0029] According to the maximum mask length set and hyperparameters , according to the geometric distribution formula to calculate the mask span length The sampling probability of , where the mask span length The value is The sampling probability calculation formula is:
[0030] , ;
[0031] in It follows a geometric distribution, :
[0032] ;
[0033] The approximate calculation formula for sampling probability of mask span length m is:
[0034] .
[0035] Calculate the mean of the sampled mask length m:
[0036] .
[0037] For example, using the spanBERT model setting value, take is 10, 0.2, masking the span length The probability of being 1 is: At this setting, the average sampled mask length is about 3.8.
[0038] Among them, the maximum mask length can be adjusted according to the average value of the mask span length and hyperparameters ; Then adjust the mask span length m of random sampling.
[0039] Step S22: randomly sampling the text sequence according to the calculated mask span length m and sampling probability to obtain a mask segment sequence with a mask span length of m.
[0040] Based on the mask span length m and sampling probability, the starting point of the text sequence is randomly selected for masking to obtain the mask span sequence , Indicates masking the starting character. Indicates masking the end character.
[0041] Step S23, through multiple rounds of sampling, the coverage rate of the sampled masked segment sequence reaches the preset coverage rate of the text sequence.
[0042] Similar to the original BERT model, the masking rate is set to 15%, and the text sequence is sampled multiple times as masked fragment sequences. In each iterative sampling process, a new text segment is selected for masking until the masking rate reaches the set value, and all masked fragment sequences are obtained.
[0043] Step S24, identifying stop words in each masked segment sequence, and based on the stop words, dividing the masked segment sequence into multiple masked segment sequences to obtain all masked segment sequences corresponding to the text sequence.
[0044] According to the stop word list, the stop words in the span are identified, the masked span containing the stop words is split into subspans without stop words, and the masked span is updated.
[0045] example: , This is the first stop word in the identified span. Multiple stop words may be identified in the span. Here we only use one stop word as an example. The subspans after segmentation based on stop words are: , , since the span length is greater than 1.
[0046] Remove stop words to avoid forming redundant invalid words in subsequent span synthesis and introduce noise into subsequent model training. When extracting the span synthesis word list, segment and reorganize the span words based on the stop word list to achieve data enhancement.
[0047] As an embodiment, in step S2, generating a synthetic span word within each masked segment sequence to form a synthetic span word list corresponding to the text data includes:
[0048] For masked fragment sequences The corresponding position embedding information is ;
[0049] Randomly select two characters in the span of the masked fragment sequence as the starting character and the ending character of the synthetic span word, and connect and synthesize the selected characters to obtain multiple synthetic span words to form a synthetic span word list.
[0050] Among them, because the BERT model is trained and predicted at the finest granularity of characters for Chinese, it tends to cover and predict short segments in the choice of masking strategy, which makes it less concerned with long words and compound words formed by splicing. Therefore, span information is introduced from the masked segment sequence list to participate in pre-training tasks and fine-tuning, which is conducive to improving the recognition accuracy of words with longer text lengths. Introducing masked span word training can enable the model to learn the character relationships between masked words.
[0051] According to the mask span selection strategy, the mask span length is randomly selected, and the mask segment sequence sampled is The corresponding position embedding information is .
[0052] Extracting masked fragment sequences , randomly select two characters in the masked fragment sequence as the starting and ending characters of the synthetic span word, connect and synthesize the selected characters, and obtain multiple synthetic span words in each masked fragment sequence. For multiple masked fragment sequences sampled from the text sequence, generate all the synthetic span words to form a synthetic span word list corresponding to the text sequence.
[0053] Step S3, extracting vector features of each synthetic span word in the synthetic span word list based on the character features and character position features of each character in the synthetic span word.
[0054] Extract the vector features of each synthetic span word in the synthetic span word list, and a vector feature matrix of the synthetic span word masking the fragment sequence includes the word embedding information of the synthetic span word , position embedding information of synthetic span words and synthetic span word length information , Indicates the starting position, Indicates the end position, To mask the length of the fragment sequence.
[0055] The information can be expressed as: , Indicates that the starting character is characters, and the ending character is Compound words, , ;
[0056] , Indicates that the starting character is characters, and the ending character is The position of the synthesized word, , ;
[0057] , Indicates that the starting character is characters, and the ending character is The length of the synthesized word, , ;
[0058] The length of the composite span is calculated by the position markers of the start and end characters:
[0059] .
[0060] Among them, the word vector of the span composite word is obtained by adding the character vectors of the constituent words on average. Indicated by The characters represented by it are concatenated, and the corresponding word vector is represented .
[0061] Position of span compound words It is formed by the relative encoding of the positions of the starting character and the ending character, and the character position vector outside the span is Equivalent to starting with itself .
[0062] The length of the span compound word Perform length encoding embedding to obtain the corresponding length vector , characters outside the span The length of is 1, and the corresponding length vector is obtained .
[0063] For each synthetic span word, its vector features are obtained, including the word vector of the synthetic span word, which is obtained by adding each character vector, the position vector of the synthetic span word, and the length vector of the synthetic span word.
[0064] Step S4, training a named entity recognition model based on the vector features of all synthetic span words and the masked language task and span boundary task.
[0065] Among them, the named entity recognition model can be trained based on the two tasks, and finally the losses of the two task trainings are combined to optimize the named entity recognition model.
[0066] Among them, the named entity recognition model is trained based on the mask language task, including:
[0067] Step S41, obtaining the vector features of each synthetic span word, masking part of the character vectors in the word vector of each synthetic span word, and obtaining the masked word vector of each synthetic span word.
[0068] in, Figure 2The vector features of each synthetic span word are shown, including the word vector of the synthetic span word, the character vector of each character of the synthetic span word, the position vector of the synthetic span word and the length vector of the synthetic span word.
[0069] For each character vector of each synthetic span word, replace 80% of the character vector with special tokens. , 10% of the masked characters are replaced with random words, and 10% of the masked characters are replaced with original words to form the masked word vector of the synthetic span word. When training the named entity recognition model later, the named entity recognition model mainly predicts special tags The content of the characters at the position replaced by the random word is used to predict the specific character content at each character position in the synthetic span word.
[0070] Step S42, obtaining the mask word vector, position vector and length vector of each synthetic span word to form a first pre-trained feature vector of each synthetic span word.
[0071] Step S43: training a named entity recognition model based on the first pre-trained feature vectors of the plurality of synthetic span words.
[0072] The process of training a named entity recognition model based on the masked language task is:
[0073] Calculate the relative position encoding of the synthetic span word based on the attention layer:
[0074] ;
[0075] in, is a learnable parameter, Indicates a connection operation. Indicated by The character or span start position is The relative distance calculated from the end position of a character or span, Indicates the starting position, Indicates the end position, and similarly indicates other relative distances;
[0076] The output of the multi-head attention layer is:
[0077] ;
[0078] ;
[0079] ;
[0080] in, , is a learnable parameter, , express , The feature vector of a character, that is, the character vector, character position vector and length vector of the character, , , is the learnable weight matrix, Encode for relative position;
[0081] The output of the multi-head attention layer is input into a 2-layer feedforward network with a normalization function for normalization to obtain the prediction results of the masked language task training. , where the prediction result Includes the character content of each character position of the synthetic span word.
[0082] The named entity recognition model is trained based on the masked language task and also based on the span boundary task. The starting point and length of the masked span are randomly selected to eliminate the need to consider boundary information. It is expected that the content within the span can be summarized at the end of the masked span boundary. The purpose of the span boundary target pre-training task is to predict the span content through the span boundary characters, introduce length information and relative position encoding, and realize the prediction of the word granularity level within the span.
[0083] Among them, the named entity recognition model is trained based on the span boundary task, including:
[0084] Step S41', calculate the relative position encoding vectors of the internal characters of the synthetic span word and the characters adjacent to the boundary position of the synthetic span word, the character vectors, character position vectors and character length vectors of the boundary characters of the synthetic span word, and the length vector of the synthetic span word, to form a second pre-trained feature vector for each synthetic span word.
[0085] in, Figure 3 The second pre-trained feature vector constructed is shown, and the relative position encoding vector of the internal characters of the synthetic span word and the characters adjacent to the boundary position of the synthetic span word is calculated, including:
[0086] ;
[0087] in, Indicates that the position in the masked fragment sequence is The character or span compound word and the starting point boundary word position The relative position of , and the same is true for other relative positions;
[0088] According to the vector features and relative position encoding vector of span boundary characters and the length vector of the synthetic span word The second pre-trained feature vector of the synthetic span word is constructed, where The vector features representing the starting boundary characters of the synthetic span word, Vector feature representing the ending boundary character of the synthesized span word.
[0089] Step S42 ′: training the named entity recognition model based on the second pre-trained feature vectors of the plurality of synthetic span words.
[0090] The second pre-trained feature vector is input into a 2-layer feed-forward network with a GeLU activation function for prediction, including:
[0091] ;
[0092] ;
[0093] ;
[0094] ;
[0095] in, It means to vertically concatenate the character vector, relative position encoding vector and length vector. The starting position is , the end position is The length of the word, Prediction results for pre-training across boundary tasks, prediction results Includes the character content of each character position of the synthetic span word.
[0096] In the process of training the named entity recognition model, the training effect of the named entity recognition model is measured by the loss function, the changing trend of the loss function is recorded, and the training is terminated when the loss function reaches the threshold.
[0097] Wherein, in the step S4, during the training of the named entity recognition model, the joint loss value of the named entity recognition model is calculated according to the first loss obtained by training the named entity recognition model with the mask language task and the second loss obtained by training the named entity recognition model with the span boundary task:
[0098] The joint loss of the named entity recognition model is calculated by weighted summation:
[0099] ;
[0100] in, For the first loss, is the second loss, the first loss and the second loss are both cross entropy losses, and the formula is expressed as:
[0101] ;
[0102] in, , is a hyperparameter;
[0103] The hyperparameters of the named entity recognition model are adjusted according to the joint loss value until the joint loss value reaches a preset value to complete the training.
[0104] Step S5: recognizing named entities in the text data to be recognized based on the trained named entity recognition model.
[0105] Wherein, the step S5, identifying the named entities in the text data to be identified based on the trained named entity recognition model, includes:
[0106] Step S51, input the text sequence to be recognized into the trained named entity recognition model, and output a predicted named entity word list.
[0107] It can be understood that the text sequence to be recognized is input into the trained named entity recognition model, the content sequence of each character position of the text sequence to be recognized is output, and each character output is labeled. The corresponding label structure is a BIO structure, where B indicates that the character is the beginning of the word, I indicates the middle to the end of the word, and O indicates other types of words.
[0108] Specifically, the calculation is based on the input sequence Under the condition of , the output sequence is The conditional probability :
[0109] ;
[0110] ;
[0111] ;
[0112] in, is the normalization factor, indicating The exponential sum of the probability scores of all possible paths, represents the state characteristic function on the node, Represents the transfer feature function between nodes, which is a learnable parameter.
[0113] The output results of the fully connected layer are converted into probability distribution through sequence joint probability distribution calculation, and finally the label with the largest probability is selected as the prediction result.
[0114] The predicted label results are post-processed to connect consecutive identical labels into named entity recognition results, and character-level connected entities and a list of predicted named entities at the span word level are obtained. The duplicated items of character-level recognized entities are removed from the predicted entity results of the span word list.
[0115] Step S52, counting the word frequency of each predicted named entity word in the predicted named entity word list, wherein the word frequency includes the number of occurrences of the predicted named entity word in the to-be-recognized text sequence and the relative proportion of the text sequence.
[0116] Among them, for the multiple named entity words identified, each named entity word is counted to obtain the recommended long words appropriate to the specific field. The frequency of span words is counted, including the absolute number of occurrences and the relative proportion of text, and a word frequency statistics table is constructed.
[0117] Step S53: taking the predicted named entity words whose frequency exceeds the set frequency threshold as the final recognized named entity.
[0118] Among them, after counting the word frequency of each named entity word, a word frequency threshold is set, and the named entity words exceeding the word frequency threshold are extracted as recommended long words, indicating that the recommended long words in the field of the text to be recognized are more relevant to the field than the character-level recognition results.
[0119] Example: For the predicted named entity word list recognition results, the final named entity recognition results are those with an absolute number of occurrences greater than or equal to 2. The predicted named entity word list recognition results appear 4 times in the full text recognition, which exceeds the set absolute number of occurrences threshold. It is recommended to replace the character-level recognition prediction results with the final named entity recognition results.
[0120] The present invention provides a method and system for named entity recognition, which has the following beneficial effects:
[0121] (1) For the masked language task training, prediction is performed at the Chinese character granularity. The prediction training lacks word-level granularity. The generation and extraction of compound words within the span, the position marking method and the length information embedding are introduced to achieve Chinese word granularity level prediction, extract span words, and obtain long word results for named entity recognition.
[0122] (2) To address the problem of ignoring the relationship between characters within a span when predicting the content within the span, a method of embedding the relative position marking and length information of compound words within a span and span boundary words is proposed to enhance the semantic learning of the context within the span.
[0123] (3) By synthesizing word length information and span words to obtain word granularity information, when performing named entity recognition in a specific field, such as fields containing a large number of professional terms and long words composed of multiple words: biomedicine, legal texts, government documents and other fields, it is possible to predict, identify and extract long Chinese words, which is beneficial to the construction of knowledge graphs, decision-making systems, etc. in subsequent specific fields.
[0124] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0125] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0126] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0127] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0129] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0130] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for named entity recognition, characterized in that: include: Step S1, segmenting each text data in the training data set into characters, and extracting the character features and character position features of each character; Step S2, masking the text data according to the masking strategy to obtain multiple masked segment sequences, generating synthetic span words within each masked segment sequence, and forming a synthetic span word list corresponding to each text data; Step S3, extracting vector features of each synthetic span word in the synthetic span word list based on the character features and character position features of each character in the synthetic span word; Step S4, training a named entity recognition model based on a masked language task and a span boundary task according to the vector features of all synthetic span words; Step S5: recognizing named entities in the text data to be recognized based on the trained named entity recognition model.
2. The method for named entity recognition according to claim 1, characterized in that: The step S1, segmenting each text data in the training data set into characters, extracting the character features and character position features of each character, includes: Convert traditional Chinese characters in text data into simplified Chinese characters, and remove special characters in text data; The text data is segmented by the word segmenter according to the character level granularity to obtain a text sequence T = {c1, c2, ..., c n }, c i Represents the i-th character, 1≤i≤n; Adding a special marker [CLS] at the beginning of a sentence and a special marker [SEP] between sentences in the text sequence; The text sequence is converted into a character index, and each character is mapped to a unique index in the vocabulary according to the vocabulary used by BERT, and the character identifier X = [x1, x2, ... x n ], the character representation and the character embedding matrix are calculated to obtain the character encoding vector E C =[e1,e2,…,e n ]; Mark each character in the text sequence with an absolute position P = [p1, p2, ..., p n ]; Each character in the text sequence is marked with a length of 1.
3. The method for named entity recognition according to claim 1, characterized in that: The step S2, masking the text data according to the masking strategy to obtain a plurality of masked segment sequences, includes: Based on the statistics of Chinese word length distribution, a masked span length m of random sampling of the text sequence is calculated, and the sampling probability of the masked span length m is calculated according to a geometric distribution formula; Among them, according to the set maximum mask length m max And hyperparameter P, calculate the sampling probability of mask span length m taking value k: Where m' follows a geometric distribution, m'~G(p): P(m'=k)=(1-p) k-1 p The approximate calculation formula for sampling probability of mask span length m is: Compute the mean of the sampled masked span lengths: Adjust the maximum mask length m based on the mean of the mask span length max and hyperparameter P; According to the calculated mask span length m and sampling probability, the text sequence is randomly sampled to obtain a mask segment sequence with a mask span length of m; Through multiple rounds of sampling, the coverage rate of the sampled masked fragment sequence reaches the preset coverage rate of the text sequence; The stop words in each masked segment sequence are identified, and based on the stop words, the masked segment sequence is divided into multiple masked segment sequences to obtain all masked segment sequences corresponding to the text sequence.
4. The method for named entity recognition according to claim 1, characterized in that: In step S2, a synthetic span word is generated within each masked segment sequence to form a synthetic span word list corresponding to the text data, including: For the masked fragment sequence X M =[x Ms ,…,x Me ] The corresponding position embedding information is P M =[p Ms ,…,p Me ]; Randomly select two characters in the span of the masked fragment sequence as the starting character and the ending character of the synthetic span word, and connect and synthesize the selected characters to obtain multiple synthetic span words to form a synthetic span word list.
5. The method for named entity recognition according to claim 4, characterized in that: The step S3, based on the character feature and the character position feature of each character in the synthetic span word, extracts the vector feature of each synthetic span word in the synthetic span word list, including: Extract the vector feature matrix of all synthetic span words in the synthetic span word list, including the word embedding information X of the synthetic span words CB , the position embedding information P of the synthesized span word CB and the length information of the synthetic span word L CB , s represents the starting character position, e represents the ending character position, and m is the length of the masked fragment sequence; Among them, x CB(i,j) represents the word vector of the synthetic span word with the starting character being the i-th character and the ending character being the j-th character synthesis, s ≤ i < j ≤ e, e - s + 1 = m, where x CB(i,j) is represented by x i x i+1 …x j concatenated by the represented characters, and its corresponding word vector represents e CB(i,j) = e i + e i+1 + … + e j , e i is the character vector of the i-th character; p CB(i,j) represents the position vector corresponding to the synthetic span word with the starting character being the i-th character and the ending character being the j-th character synthesis, s ≤ i < j ≤ e, e - s + 1 = m, where the position p of the span synthesis word CB(i,j) is encoded by the relative positions of the starting character and the ending character, and the position vector of the characters outside the span is p n is equivalent to the starting characters all being themselves p CB(n,n) ; l CB(i,j) represents the length vector corresponding to the synthetic span word with the starting character being the i-th character and the ending character being the j-th character, s ≤ i < j ≤ e, e - s + 1 = m, where the length l of the span synthetic word CB(i,j) is subjected to length encoding embedding to obtain the corresponding length vector e L(i,j) , the character x outside the span n has a length of 1, and the corresponding length vector e is obtained L(n,n) ; The length of the synthetic span word is calculated by the position markers of the starting character i and the ending character j: l CB(i,j) =j-i+1。 6. The method for named entity recognition according to claim 5, characterized in that: The step S4, training the named entity recognition model based on the masked language task and the span boundary task according to the vector features of all synthetic span words, includes: Train a named entity recognition model based on a masked language task, including: Obtain the vector features of each synthetic span word, mask part of the character vectors in the word vector of each synthetic span word, and obtain the masked word vector of each synthetic span word; Obtaining a mask word vector, a position vector, and a length vector of each synthetic span word to form a first pre-trained feature vector of each synthetic span word; Training a named entity recognition model based on first pre-trained feature vectors of the plurality of synthetic span words; The named entity recognition model is trained based on the boundary spanning task, including: Calculate the relative position encoding vectors of the internal characters of the synthetic span word and the characters adjacent to the boundary position of the synthetic span word, the character vectors, the character position vectors and the character length vectors of the boundary characters of the synthetic span word, and the length vector of the synthetic span word to form a second pre-trained feature vector for each synthetic span word; The named entity recognition model is trained based on the second pre-trained feature vectors of the plurality of synthetic span words.
7. The method for named entity recognition according to claim 6, characterized in that: The training of the named entity recognition model based on the first pre-trained feature vectors of the plurality of synthetic span words comprises: Calculate the relative position encoding of the synthetic span word based on the attention layer: Among them, W r is a learnable parameter, Indicates a connection operation. Represents the relative distance calculated from the starting position of the i character or span and the ending position of the j character or span, s represents the starting position, e represents the ending position, and the same is true for other relative distances; The output of the multi-head attention layer is: Att(A,V)=softmax(A)V V. E x W v Among them, u, v are learnable parameters, The feature vector representing the i and j characters, i.e., the character vector, character position vector, and length vector of the character, W q , W v is the learnable weight matrix, R ij Encode for relative position; The output of the multi-head attention layer is input into a 2-layer feedforward network with a normalization function for normalization to obtain the prediction result y of the masked language task training M , where the predicted result y M Includes the character content of each character position of the synthetic span word.
8. The method for named entity recognition according to claim 6, characterized in that: The step of calculating the relative position encoding vectors of the internal characters of the composite span word and the characters adjacent to the boundary position of the composite span word includes: in, Indicates the relative position of the character or span compound word at position (i, j) in the masked segment sequence and the starting boundary word position (s-1, s-1). The same is true for other relative positions. According to the vector features and relative position encoding vector R of the span boundary characters B(ij,s-1) and the length vector l of the synthetic span word CB(i,j) The second pre-trained feature vector of the synthetic span word is constructed, where E s-1 The vector feature representing the starting boundary character of the synthetic span word, E e+1 The vector feature representing the ending boundary character of the synthetic span word; The second pre-trained feature vector is input into a 2-layer feed-forward network with GeLU activation function for prediction, including: L ij =j-i+1 LB'=LayerNorm(GeLU(W1LB)) y B =LayerNorm(GeLU(W1LB')) Among them, LB means vertical concatenation of character vector, relative position encoding vector and length vector, L ij It is represented by the length of the word starting at position i and ending at position j, y B The prediction result y is the pre-trained prediction result of the span boundary task. B Includes the character content of each character position of the synthetic span word.
9. The method for named entity recognition according to claim 1, characterized in that: In step S4, during the training of the named entity recognition model, the joint loss value of the named entity recognition model is calculated according to the first loss obtained by training the named entity recognition model with the mask language task and the second loss obtained by training the named entity recognition model with the span boundary task: The joint loss of the named entity recognition model is calculated by weighted summation: LOSS=αLOSS M +βLOSS B Among them, LOSS M For the first loss, LOSS B is the second loss, the first loss and the second loss are both cross entropy losses, and the formula is expressed as: LOSS=αLOSS M +βLOSS B =-αlog(x M |y M )-βlog(x B |y B ) Among them, α and β are hyperparameters; The hyperparameters of the named entity recognition model are adjusted according to the joint loss value until the joint loss value reaches a preset value to complete the training.
10. The method for named entity recognition according to claim 1, characterized in that: The step S5, identifying the named entities in the text data to be identified based on the trained named entity recognition model, includes: Input the text sequence to be recognized into the trained named entity recognition model and output a list of predicted named entity words; Counting the word frequency of each predicted named entity word in the predicted named entity word list, wherein the word frequency includes the number of occurrences of the predicted named entity word in the to-be-recognized text sequence and the relative proportion of the text sequence; The predicted named entity words whose frequency exceeds the set frequency threshold are taken as the final recognized named entities.
Citation Information
Patent Citations
Address named entity recognition tuning method based on deep learning model
CN114169332A
Small sample named entity recognition method based on multiple tasks and prompt learning
CN116151256A